Multi-center medical image data collaborative training method and system based on federated learning

CN122597958APending Publication Date: 2026-08-18FUJIAN RUISIKE MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611079850.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

中央聚合节点在对此类嵌入向量进行综合处理时,受限于现有聚合逻辑对空间连续分布结构的解析能力,所生成的全局特征信号对于白质-灰质交界带处信号连续过渡的梯度信息反映不够充分

Benefits of technology

[0023] By constructing a dynamic constraint field and performing step-by-step scanning along the field guidance direction, the scan trajectory adaptively tracks the continuous manifold of features in the shared embedding space, avoiding the fragmentation of embedding vectors in transition areas. The resulting global pathological feature description more accurately reflects the continuous anatomical change trend of cross-center images, improving the consistency of output and anatomical reconstruction in tissue boundary localization and image segmentation tasks for each center. The standard deviation of the embedding vector distribution in the neighborhood of each reference position is used as the initial value of the scan radius, and it is dynamically adjusted according to the variation amplitude of coupling strength in the dynamic constraint field, effectively reducing structural errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597958A_ABST
    Figure CN122597958A_ABST
Patent Text Reader

Abstract

The application provides a multi-center medical image data collaborative training method and system based on federated learning, and relates to the technical field of data processing. The method comprises the following steps: each participating center trains a local coding network with local multi-modal MRI images, maps samples to a shared embedding space, obtains embedding vectors, and uploads the embedding vectors to a central aggregation node; the central aggregation node determines a first reference position, a second reference position and a third reference position according to the spatial distribution of the embedding vectors. The application relies on federated learning to realize local storage of medical original images, transmission of embedding vectors to guarantee data privacy, construction of a dynamic constraint field to eliminate multi-center feature differences, adaptive scanning and collection of lesion features, fusion of path features to generate unified global features, synchronous optimization of each center network, and consideration of data privacy and cross-center lesion recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for collaborative training of multi-center medical image data based on federated learning. Background Technology

[0002] In a multi-center medical image collaborative processing system based on federated learning, each participating center typically needs to map its locally acquired multimodal MRI image data to a shared high-dimensional feature space. The central aggregation node then performs comprehensive processing on the received embedding vectors to obtain global feedback signals that can be used to update the local coding networks of each center, thereby improving the consistency of feature representations in cross-center image processing tasks.

[0003] However, in actual system operation, the MRI scanning equipment (e.g., MRI machines with different field strengths) and imaging parameter settings (e.g., repetition time, echo time, number of diffusion-sensitive gradient directions, etc.) used by different centers have objective differences, resulting in a non-uniform clustering of the embedded vectors uploaded from each center to the central node in the feature space. Existing aggregation processing modules, after receiving the embedded vectors from each center, typically use merging logic based on vector arithmetic average or fixed weights. This data processing method, when faced with multiple clusters in the feature space with significantly different distribution densities and gradual transitions between clusters, sometimes results in a certain degree of deviation between the output feature description vector and the actual continuous transition regions (e.g., the signal gradient band at the junction of white matter and gray matter) in the anatomical structure of the image.

[0004] Some systems have introduced feature filtering or clustering preprocessing steps based on spatial distance thresholds to optimize data alignment before aggregation. However, since the distance criterion relied upon by such processing logic is a preset fixed value, this fixed threshold is difficult to automatically adjust with the local data distribution when there are significant differences in sample density and scattering range in different regions of the embedding space. This can lead to the system incorrectly classifying sample points that are actually on continuous transition paths as discrete clusters when processing embedding vectors in certain regions, or conversely, excessively merging samples that actually belong to different anatomical structures. This causes structural errors in the information integration process, which in turn affects the output stability of subsequent image segmentation or tissue category labeling tasks.

[0005] For example, when multiple medical institutions jointly process multimodal MRI images of the brain, each center provides image data from T2-FLAIR and diffusion-weighted imaging (DWI) sequences. Because the b-values ​​(e.g., 800 s / mm² vs. 1500 s / mm²) and slice thickness settings used by each center in their scanning protocols are not entirely consistent, the image features of the white matter-gray matter junction and adjacent sulci and gyri form several spatially overlapping but centrally offset subgroups within the shared embedding space. When the central aggregation node synthesizes these embedding vectors, it is limited by the existing aggregation logic's ability to resolve spatially continuous distribution structures. The resulting global feature signal does not adequately reflect the gradient information of the continuous signal transition at the white matter-gray matter junction. After this signal is distributed back to each center, the local encoding network struggles to achieve a consistent optimization direction during parameter adjustment, leading to certain deviations in the tissue boundary localization outputs when processing images of the same brain region. Summary of the Invention

[0006] This invention provides a method and system for collaborative training of multi-center medical image data based on federated learning, which alleviates the problem of divergence in the direction of local model updates caused by differences in data distribution.

[0007] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0008] Firstly, a multi-center medical image data collaborative training method based on federated learning, the method comprising:

[0009] Step 1: Each participating center trains a local coding network using local multimodal MRI images, maps the samples to a shared embedding space, obtains embedding vectors, and uploads the embedding vectors to the central aggregation node;

[0010] Step 2: The central aggregation node determines the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector;

[0011] Step 3: Calculate the spatial deviation between each reference pair, and determine the coupling strength between each reference pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference pair, so as to construct a dynamic constraint field;

[0012] Step 4: Using the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, the amplitude and orientation of the changing coupling strength at each reference position are calculated in the dynamic constraint field, and the scanning radius is adjusted according to the amplitude.

[0013] Step 5: Starting from the first reference position, perform a step-by-step scan along the changing orientation with a scanning radius. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, establish a first coupling path between the second reference position and the first reference position. When the third reference position falls into the scanning radius, establish a second coupling path between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position in space.

[0014] Step 6: Along the first coupling path and the second coupling path, the local features of each reference position are merged step by step to obtain a global pathological feature description that is invariant across the central domain. This description is then distributed to each participating center to drive the parameter directionality update of the local coding network.

[0015] Secondly, a multi-center medical image data collaborative training system based on federated learning includes:

[0016] The first module is used by each participating center to train a local coding network with local multimodal MRI images, map the samples to a shared embedding space, obtain embedding vectors, and upload the embedding vectors to the central aggregation node.

[0017] The second module is used by the central aggregation node to determine the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector;

[0018] The third module is used to calculate the spatial deviation between each reference bit pair and determine the coupling strength between each reference bit pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference bit, so as to construct a dynamic constraint field.

[0019] The fourth module is used to take the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, calculate the variation amplitude and variation orientation of the coupling strength at each reference position in the dynamic constraint field, and adjust the generated scanning radius according to the variation amplitude.

[0020] The fifth module is used to perform a step-by-step scan with a scanning radius along a changing orientation, starting from the first reference position. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, a first coupling path is established between the second reference position and the first reference position. When the third reference position falls into the scanning radius, a second coupling path is established between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position spatial position.

[0021] The sixth module is used to merge the local features of each reference position along the first and second coupling paths step by step to obtain a global pathological feature description that is invariant across the central domain, and distribute it to each participating center to drive the parameter directionality update of the local coding network.

[0022] The above-described solution of the present invention has at least the following beneficial effects:

[0023] By constructing a dynamic constraint field and performing step-by-step scanning along the field guidance direction, the scan trajectory adaptively tracks the continuous manifold of features in the shared embedding space, avoiding the fragmentation of embedding vectors in transition areas. The resulting global pathological feature description more accurately reflects the continuous anatomical change trend of cross-center images, improving the consistency of output and anatomical reconstruction in tissue boundary localization and image segmentation tasks for each center. The standard deviation of the embedding vector distribution in the neighborhood of each reference position is used as the initial value of the scan radius, and it is dynamically adjusted according to the variation amplitude of coupling strength in the dynamic constraint field, effectively reducing structural errors.

[0024] The coupling strength type is determined by comparing the spatial deviation between reference positions with the equilibrium spacing, and the coupling magnitude is determined by the degree of aggregation of neighborhood embedding vectors. A dynamic constraint field is constructed by superimposing these values. Attraction constraints are applied to discrete regions to enhance cohesion, and repulsion constraints are applied to overly dense regions to prevent pattern collapse. This maintains the consistency of cross-center distribution in the embedding space throughout multiple rounds of federated training. Local feature vectors at each step position are extracted sequentially along the coupling path, and a merging coefficient is determined by the scan radius ratio for progressive merging. This ensures that the global pathological features contain the complete context along the continuous manifold from the first reference position to the third reference position. The global pathological feature description is obtained by manifold tracing and coupling path merging modulated by the dynamic constraint field. It integrates the spatial distribution structure of multi-center embedding vectors and maintains the gradient features of the transition region. After being distributed to each center, it provides a unified reference for parameter updates, alleviating the problem of divergence in local model update direction caused by differences in data distribution, improving the overall convergence effect of collaborative training and the generalization ability of models deployed independently by each center. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating a multi-center medical image data collaborative training method based on federated learning, provided in an embodiment of the present invention.

[0026] Figure 2 This is a schematic diagram of a multi-center medical image data collaborative training system based on federated learning provided in an embodiment of the present invention. Detailed Implementation

[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0028] like Figure 1As shown, embodiments of the present invention propose a multi-center medical image data collaborative training method based on federated learning, the method comprising the following steps:

[0029] Step 1: Each participating center trains a local coding network using local multimodal MRI images, maps the samples to a shared embedding space, obtains embedding vectors, and uploads the embedding vectors to the central aggregation node;

[0030] Step 2: The central aggregation node determines the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector;

[0031] Step 3: Calculate the spatial deviation between each reference pair, and determine the coupling strength between each reference pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference pair, so as to construct a dynamic constraint field;

[0032] Step 4: Using the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, the amplitude and orientation of the changing coupling strength at each reference position are calculated in the dynamic constraint field, and the scanning radius is adjusted according to the amplitude.

[0033] Step 5: Starting from the first reference position, perform a step-by-step scan along the changing orientation with a scanning radius. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, establish a first coupling path between the second reference position and the first reference position. When the third reference position falls into the scanning radius, establish a second coupling path between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position in space.

[0034] Step 6: Along the first coupling path and the second coupling path, the local features of each reference position are merged step by step to obtain a global pathological feature description that is invariant across the central domain. This description is then distributed to each participating center to drive the parameter directionality update of the local coding network.

[0035] In this embodiment of the invention, by constructing a dynamic constraint field and performing step-by-step scanning along the field guidance direction, the scanning trajectory adaptively tracks the continuous manifold of features in the shared embedding space, avoiding the fragmentation of embedding vectors in transition areas. The resulting global pathological feature description more accurately reflects the continuous anatomical change trend of cross-center images, improving the consistency of output and anatomical reconstruction in tissue boundary localization and image segmentation tasks for each center. The standard deviation of the embedding vector distribution within the neighborhood of each reference position is used as the initial value of the scanning radius, and dynamically adjusted according to the changing amplitude of coupling strength in the dynamic constraint field, effectively reducing structural errors.

[0036] The coupling strength type is determined by comparing the spatial deviation between reference positions with the equilibrium spacing, and the coupling magnitude is determined by the degree of aggregation of neighborhood embedding vectors. A dynamic constraint field is constructed by superimposing these values. Attraction constraints are applied to discrete regions to enhance cohesion, and repulsion constraints are applied to overly dense regions to prevent pattern collapse. This maintains the consistency of cross-center distribution in the embedding space throughout multiple rounds of federated training. Local feature vectors at each step position are extracted sequentially along the coupling path, and a merging coefficient is determined by the scan radius ratio for progressive merging. This ensures that the global pathological features contain the complete context along the continuous manifold from the first reference position to the third reference position. The global pathological feature description is obtained by manifold tracing and coupling path merging modulated by the dynamic constraint field. It integrates the spatial distribution structure of multi-center embedding vectors and maintains the gradient features of the transition region. After being distributed to each center, it provides a unified reference for parameter updates, alleviating the problem of divergence in local model update direction caused by differences in data distribution, improving the overall convergence effect of collaborative training and the generalization ability of models deployed independently by each center.

[0037] In a preferred embodiment of the present invention, step 1 includes:

[0038] Step 100: Each participating center performs inter-slice spatial registration on local multimodal MRI images to obtain registered multi-sequence voxel data; based on the registered multi-sequence voxel data, a local training batch containing multiple samples is constructed; based on the local training batch, forward and backward propagation are performed on the local coding network to update the parameters of the local coding network, resulting in the updated local coding network after this round of training, specifically including:

[0039] The original MRI images acquired independently by each participating center include multiple sequence modalities such as T1-weighted, T2-weighted, and FLAIR. Due to differences in equipment parameters, scan slice thickness, and acquisition resolution among different medical institutions, the original images suffer from spatial misalignment, scale inconsistency, and heterogeneous grayscale distribution, making them unsuitable for direct input into the network for training. Therefore, global spatial correction and inter-slice alignment processing are performed on the multi-sequence MRI images acquired by each center. First, rigid spatial alignment of the overall images is achieved, unifying the global spatial coordinate system and standard voxel size of all images. Then, fine-tuning and correction are performed on local minor deformations and inter-slice misalignment areas to correct local spatial deviations caused by scanning with different equipment. Ultimately, this achieves complete inter-slice alignment and spatial scale uniformity of MRI images from all centers, eliminating image spatial shift and positional heterogeneity errors caused by differences in equipment parameters across multiple centers. After registration, all image samples are screened one by one to remove invalid and abnormal samples with image defects, severe blur artifacts, or registration failures. Finally, a multi-sequence voxel data tensor with uniform dimensions, compliant image quality, and spatial alignment is obtained. The input dimension of a single sample is fixed as [4,128,128,128], which corresponds to the three-dimensional voxel scales of the four MRI modalities, length, width, and height, respectively.

[0040] After completing the inter-slice spatial registration of multimodal MRI images and obtaining the registered multi-sequence voxel data, the structured construction of training batches was carried out. Using single-round local iterative training as the basic unit, all compliant registration samples were randomly and uniformly divided without bias. The training sample size per batch was fixed at 16. The division process ensured that the MRI modality ratio, the proportion of lesion-positive samples, and the proportion of normal samples were evenly distributed within each batch, avoiding network training bias caused by batch data bias. Finally, a local training batch with uniform sample distribution, uniform data specifications, and no abnormal interference was constructed.

[0041] To adapt to the need for sparse pathological feature extraction from multimodal MRI images and to solve the problem of heterogeneous features in multicenters, a hierarchical stacked local coding network is constructed. The network consists of modality-independent convolutional layers, nonlinear activation layers, dimension-mapping pooling layers, modality feature fusion layers, and fully connected coding layers stacked in an orderly manner.

[0042] Specifically, the modality-independent convolutional layer uses four sets of parallel independent 3D convolutional units for the four MRI modalities to achieve individual feature extraction for each modality, avoiding feature interference between modalities. The convolutional kernels are uniformly 3×3×3 in size, with a stride of 1 and edge padding of 1, ensuring that the voxel spatial dimensions remain unchanged before and after convolution. The weights of the convolutional layers are initialized using the Kaiming normal distribution, with an initialization standard deviation of [value missing]. In the formula =3 represents the kernel space size. =4 represents the number of single-mode input channels. Substituting this into the calculation yields the fixed initialization standard deviation. The initial bias of the convolutional layer is uniformly fixed at 0.01 to avoid the problem of neuron output returning to zero and gradient failure; the number of output channels of each group of convolutional units is uniformly set to 32 to ensure that the primary feature dimension of a single modality is consistent.

[0043] To overcome the limitations of linear convolutional features and uncover subtle pathological features hidden in MRI images, a ReLU nonlinear activation layer is connected in series after the output of each modal convolutional layer to complete the mapping and transformation of linear features into nonlinear high-dimensional features, thereby enhancing the network's ability to represent subtle features such as lesion edges and diffuse lesions.

[0044] After each activation layer output, a 3D max-pooling downsampling operation is performed to filter core pathological features, remove redundant background voxel information, and complete the dimensionality compression and semantic mapping of 3D voxel features. A fixed 2×2×2 3D local sampling window is used, with a sliding step of 2 along the length, width, and depth dimensions of the feature tensor. A non-overlapping uniform block traversal method is used to complete global downsampling, progressively halving the 3D voxel space dimension to achieve the conversion of high-dimensional raw voxel data into low-dimensional refined semantic features. Specifically, the overall 3D feature tensor output by activation is divided into several non-overlapping 2×2×2 voxel sub-regions. Feature values ​​of all voxels within each sub-region are extracted, and the maximum value is selected as the output feature of that sub-region, discarding the remaining low-value redundant features. The 3D max-pooling calculation formula is: ;

[0045] In the formula These are the original three-dimensional feature tensor values ​​after single-modal activation. This represents the dimensional offset of the 3D sampling window; The coordinates of the original feature tensor corresponding to the downsampling are shown below. This represents the output feature value at the corresponding coordinate position after pooling downsampling.

[0046] To integrate complementary features from multimodal MRI and compensate for the deficiencies in single-modal feature representation, pixel-wise weighted fusion was performed on four groups of independent single-modal pooling feature tensors. Initially, the fusion weights were equally distributed and fixed. To ensure balanced contributions from all modal features and no modality bias in the initial training phase, the fusion calculation formula is as follows: In the formula For the first The single-modal feature tensor extracted from the modalities after convolution, activation, and pooling. For the corresponding modal fusion weights, This is the comprehensive feature tensor after multimodal fusion.

[0047] The 3D feature tensor obtained from multimodal fusion is flattened into a one-dimensional feature vector and fed into a two-layer cascaded fully connected coding network to complete dimensionality reduction and standardized coding of high-dimensional features. The first fully connected layer has an input dimension matching the flattened dimension of the fused features, and a fixed output dimension of 512. The weights are initialized with a standard deviation of 0.0625, and the initial bias is fixed at 0.01. The second fully connected layer has an input dimension of 512 and a fixed output dimension of 128, ultimately generating a feature vector of uniform dimension. The weights in this layer are initialized with a standard deviation of 0.08, and the initial bias is fixed at 0.01. Redundant feature filtering and core pathological feature condensation are achieved through these two coding layers. After completing the above complete network structure construction and global parameter initialization, an initial local coding network with a fixed structure, complete parameters, and unified multi-centers is obtained.

[0048] The constructed local training batch sample tensors are input batch by batch into the initialized local encoding network. Specifically, the four-class single-modality voxel tensors of each sample in the batch are input into the corresponding parallel convolutional units to complete the primary pathological feature extraction. The calculation formula is as follows: In the formula For the first Input voxel tensors of various modalities , These are the initial weights and bias parameters for the corresponding modal convolutional layer, where the weights are... The fixed standard deviation and bias values ​​are obtained through Kaiming normal initialization calculation. The value is uniformly fixed at 0.01. For the first The primary features of the convolution output of each modality.

[0049] The primary features output from each modality of convolution are successively input into the ReLU activation function to complete the nonlinear transformation of linear features, thereby improving the feature representation capability. The calculation formula is as follows: In the formula These are primary features of single-modal convolution. This represents the nonlinear feature after activation.

[0050] 3D max pooling downsampling is performed on the nonlinear features after activation of each modality to compress the spatial dimension and refine the core features. The calculation formula is as follows: In the formula These are the pooled and compressed unimodal mapping features. Equally weighted fusion is performed on the four sets of unimodal mapping features with uniform dimensions to integrate complementary multimodal pathological information. The calculation formula is as follows: In the formula For multimodal fusion and comprehensive features, the initial fusion weights .

[0051] The fused features are encoded and reduced in dimensionality layer by layer through a two-layer fully connected network, ultimately outputting a standardized high-dimensional feature vector. The calculation formula is as follows: In the formula , These are the first and second fully connected coding layers, respectively. This is the final output of the forward propagation in this round, a 128-dimensional high-dimensional feature vector for a single sample. After extracting features from all samples in a single batch via forward propagation, the batch training loss is calculated based on the ground truth labels of the actual pathological data, quantifying the network feature extraction error. A mean squared error loss function is used to compare the deviation between the predicted high-dimensional features and the ground truth features sample by sample and then average them. Based on the batch loss values, global backpropagation is completed using the chain rule, solving for the gradient information of the weights and biases of each layer of the network layer by layer, while fixing the global training learning rate. =0.001, all network parameters are iteratively updated using the gradient descent algorithm. The parameter update formula is: ;

[0052] In the formula , The weights and bias parameters before the update. , For the parameters after iterative updates, , These are the weight gradient and the bias gradient, respectively. After completing the single-batch parameter iteration update, all optimized network weights and bias parameters are retained, and the current round of local iteration training is terminated, resulting in a local coding network that has been updated and adapted to multi-center pathological feature extraction after this round of training.

[0053] Step 101: Input each sample from the local training batch into the local encoding network updated in this round of training to obtain the embedding vector corresponding to each sample; summarize the embedding vectors corresponding to each sample and send them to the central aggregation node, specifically including:

[0054] Reusing the fixed forward propagation logic and calculation formula from step 100, all samples in the local training batch are input into the frozen local encoding network, and a 128-dimensional original embedding vector is obtained through sample-by-sample computation, realizing the standardized mapping from the original voxel data to the high-dimensional semantic embedding vector. Anomaly filtering and standardization are performed on the batch-output embedding vectors to remove invalid anomaly vectors such as null values ​​and all zeros; L2 normalization is performed on valid vectors to unify the amplitude scale, resulting in standardized embedding vectors. All standardized embedding vectors are aggregated, adhering to the core principle of federated learning that original data does not leave the local area and privacy data does not cross-domain transmission. Only the aggregated embedding vector set is encrypted and securely sent to the central aggregation node via an encrypted privacy communication link. The original MRI voxel image data remains stored in each local center throughout the process and is not transmitted externally.

[0055] This embodiment achieves local image spatial registration to eliminate positional shifts in multimodal MRI sequences and unifies the voxel spatial benchmark; it batches local training batches to achieve sample standardization management; it locally completes forward and backward propagation parameter updates for the encoding network, ensuring that the original medical images remain within the local center throughout the entire process, balancing model training convergence with multi-center data privacy and security, and avoiding feature extraction bias caused by cross-center image spatial misalignment from the source. The trained local encoding network uniformly generates standardized embedding vectors, uploading only low-dimensional feature vectors to the central node without transmitting the original images, reducing transmission bandwidth and privacy leakage risks; and it batches and aggregates sample embedding vectors.

[0056] In a preferred embodiment of the present invention, step 2 includes:

[0057] Step 200: Using each embedding vector as the center, calculate the number of remaining embedding vectors covered within a preset neighboring distance range to obtain the neighborhood coverage of each embedding vector; based on the neighborhood coverage of each embedding vector, mark the embedding vectors whose neighborhood coverage reaches the minimum neighborhood coverage threshold as high-density embedding vectors; using each high-density embedding vector as a seed, group the high-density embedding vectors with various subspace spacings smaller than the preset neighboring distance into the same embedding vector cluster, obtaining an embedding vector cluster composed of high-density embedding vectors; based on the embedding vector cluster composed of high-density embedding vectors, group the remaining embedding vectors not yet grouped into any embedding vector cluster into the already formed cluster with the smallest spatial spacing to the remaining embedding vectors not yet grouped into any embedding vector cluster, obtaining multiple embedding vector clusters corresponding to all embedding vectors, specifically including:

[0058] The central aggregation node gathers standardized embedding vectors from all participating centers' encrypted uploads, integrates multi-center distributed feature data, eliminates cross-center feature dimensionality bias and distribution differences, and constructs a globally shared embedding space with unified dimensions, aligned distributions, and full domain coverage, forming a global embedding vector set containing all training sample features. A preset proximity distance is determined. Specifically, the method involves traversing all embedding vectors within the shared embedding space, calculating the high-dimensional Euclidean distance pairwise, and constructing a global distance distribution set. Considering the characteristics of multi-center MRI pathological features, which are often locally densely clustered and noise features are discretely distributed, the 80th percentile of the global distance distribution is selected as a unified threshold to distinguish between homologous dense pathological features and discrete noise features. It is assumed that a pre-defined proximity distance suitable for the feature distribution of this network is ultimately determined. =0.8.

[0059] Determine the preset proximity distance After setting the threshold to 0.8, each embedding vector within the shared embedding space is used as an independent center. Based on this fixed neighborhood threshold, the number of other embedding vectors contained within the effective neighborhood of each vector is counted, quantifying the local clustering and enrichment degree of a single feature point, thus obtaining the neighborhood coverage number corresponding to each embedding vector. The formula for calculating the neighborhood coverage number is as follows: ;

[0060] in For the first The neighborhood cover number of each embedded vector. This represents the total number of globally embedded vectors. For the first A number of embedding vectors to be detected. For any other embedding vector in the global domain, Let Euclidean distance be the distance between two sets of high-dimensional embedding vectors. It is a binary indicator function, which takes the value of 1 when the vector spacing is less than or equal to the preset neighbor distance, and takes the value of 0 otherwise.

[0061] Set a minimum neighborhood coverage threshold. This is used to distinguish between valid pathological feature vectors and invalid background or noise vectors. Threshold comparisons are performed on all embedded vectors across the entire domain. Embedded vectors with a neighborhood coverage of 12 or more are marked as high-density embedded vectors. These vectors exhibit high clustering of surrounding features, rich homologous pathological information, and high semantic reliability, representing core pathological features across the entire domain. Embedded vectors with a neighborhood coverage of less than 12 are marked as low-density embedded vectors. These vectors often correspond to weak features in image background, scanning random noise, and pathological transition areas, exhibiting strong discreteness and a low proportion of valid pathological semantics.

[0062] To ensure that the clustering results closely reflect the distribution of real pathological features and avoid the problems of random clustering bias and chaotic cluster structure, a hierarchical progressive clustering strategy is adopted, which prioritizes high-density seed clustering and adaptively clusters low-density vectors. Specifically, all selected high-density embedding vectors are used as reliable clustering seeds. These seed vectors are semantically pure, concentrated in distribution, and highly reliable, serving as the core benchmark for embedding clusters and effectively ensuring the accuracy of the clustering subjects. The high-dimensional Euclidean distance between all high-density embedding vectors is calculated pairwise to determine the spatial association and semantic homology between different seed vectors.

[0063] With adaptively determined preset neighbor distance As a threshold for determining intra-cluster homology association, any two high-density seed vectors with a spatial distance less than or equal to 0.8 are classified as homologous pathological features and grouped into the same embedding vector cluster; high-density vectors with a spatial distance greater than 0.8 are classified as heterologous independent features and grouped into separate clusters. For conflicts in the classification of multiple high-density seed vectors (i.e., seed A has a spatial distance less than or equal to both seed B and seed C), the threshold is set. However, the distance between B and C is greater than The transitive closure strategy is used. If A and B are in the same cluster, and A and C are in the same cluster, then B and C are automatically merged into the same cluster, regardless of whether the direct distance between B and C is less than 1 / 3. This ensures the consistency of cluster relationship propagation. By iteratively traversing and classifying all high-density embedding vectors, clustering is completed, ultimately resulting in several initial embedding vector clusters that consist solely of high-density core pathological features, are semantically pure, have a concentrated distribution, and clear boundaries.

[0064] After high-density core vector clustering, all remaining vectors not assigned to any initial cluster are low-density embedding vectors. Based on this, all unassigned low-density embedding vectors are individually selected as the feature sample set to be clustered. For each low-density embedding vector to be clustered, its high-dimensional Euclidean distance to the cluster centers of all high-density initial clusters is calculated. The cluster center of each initial cluster is calculated from the spatial mean of all high-density embedding vectors within that cluster. The spatial distance between each individual low-density vector and each initial cluster is compared, and the low-density vector is assigned to the high-density cluster with the smallest spatial distance. After completing the classification operation for all low-density embedding vectors, a complete set of embedding vector clusters is finally obtained, covering all embedding vectors in the entire domain, semantically continuous, hierarchically distinct, and closely reflecting the distribution of real MRI pathological features.

[0065] Step 201: For each embedded vector cluster, calculate the spatial compactness between the embedded vectors within the corresponding cluster to obtain the cluster compactness metric value. A higher cluster compactness metric value indicates that the embedded vectors within the corresponding cluster are more closely spaced. Specifically, this includes:

[0066] Using a single embedding vector cluster as an independent measurement unit, the spatial clustering density and pathological homology of features within the corresponding cluster are quantitatively measured based on the spatial dispersion of all embedding vectors within the cluster and their corresponding cluster centers, yielding the clustering density metric for each embedding vector cluster. The formula for calculating the clustering density metric is as follows: ;

[0067] in For the first The compactness measure of a cluster. For the first The total number of embedding vectors within each cluster. For the first The set of all embedding vectors of each cluster For the first The cluster center vector of a cluster. is the Euclidean distance between a single vector within the cluster and the cluster center.

[0068] A higher cluster compactness metric indicates a denser distribution of embedded vector space within the corresponding cluster, stronger homology of pathological features, and less interference from heterogeneous noise within the cluster; a metric value closer to 0 indicates a more discrete distribution of features within the corresponding cluster, higher heterogeneity of pathological semantics, and worse feature consistency.

[0069] Step 202: Based on the cluster compactness metric of each embedded vector cluster, extract the embedded vector clusters with cluster compactness metrics higher than the compactness threshold as compact clusters; merge compact clusters that are spatially adjacent to each other in the shared embedding space to obtain several candidate aggregation regions, specifically including:

[0070] A compactness screening threshold is set, which is directly obtained from the statistical calculation of the compactness of all clusters. Specifically, the compactness values ​​of all clusters are statistically analyzed, and the global mean and global standard deviation of the compactness of all clusters are calculated. The optimal screening threshold is obtained by adding the global mean and global standard deviation. Let's assume a compactness screening threshold... This threshold can directly eliminate invalid noise clusters with low compactness and loose dispersion, while stably retaining valid pathological clusters with concentrated distribution and consistent pathological semantics, thus achieving accurate differentiation between valid clusters and noise clusters.

[0071] Based on this threshold, high-quality compact clusters with compactness metrics exceeding the threshold are selected, while all invalid loose cluster structures are eliminated. For each of the selected high-compactness clusters, spatial association and semantic homology are quantitatively determined. Specifically, the spatial association between any two high-compactness clusters is determined by calculating the Euclidean distance between their cluster centers. If the distance between their cluster centers is less than the global average cluster distance, the two clusters are considered spatially adjacent and have continuous feature distributions. Semantic homology is determined using a fixed quantitative standard, first by calculating the mean deviation rate of the features between the two clusters. Difference between firmness and tightness The calculation formula is: ;

[0072] in , These are the feature vectors of the two cluster centers. , This represents the compactness value between the two clusters. When the mean values ​​of the two clusters' characteristics deviate... And the difference in tightness When determining whether two pathological clusters are semantically homologous, fragmented clusters that simultaneously meet the conditions of spatial adjacency and semantic homology are merged and integrated to eliminate discrete, fragmented, and invalid clusters, resulting in a unified pathological feature region with complete boundaries, continuous features, and consistent semantics. After cluster merging, several high-quality pathological feature regions with spatial independence, clear boundaries, complete semantics, high feature enrichment, and low noise interference are generated.

[0073] Step 203: Based on several candidate pooling regions, count the number of embedded vectors covered in each candidate pooling region, and determine the three candidate pooling regions with the largest number of embedded vectors as the initial pooling regions; if the spatial distance between any two initial pooling regions is less than the pooling region distance threshold, then remove the one with the smaller number of embedded vectors; determine the three final retained initial pooling regions as pooling reference regions, and take the average position of all embedded vectors in each pooling reference region in the shared embedding space as the center position of the corresponding pooling reference region; determine the center positions of the three pooling reference regions as the first reference position, the second reference position, and the third reference position in sequence, specifically including:

[0074] For all merged candidate pooling regions, the total number of embedded vectors within each region is counted. The total enrichment of vectors within a region is used as a quantitative indicator to objectively characterize the enrichment degree of pathological features, semantic weight, and pathological activity confidence of each region. The more vectors, the higher the significance and confidence of the corresponding region's pathological features. All candidate pooling regions are sorted in descending order according to the total number of embedded vectors within the region. The top three candidate pooling regions are selected as the initial core pathological regions for the entire region, ensuring that the initial selected regions have the best representativeness of pathological features.

[0075] Calculate the Euclidean distance between the cluster centers of the three initially selected core aggregation domains pairwise, and set a fixed region redundancy determination threshold. When the cluster center spacing is less than 1.2, it indicates that the two regions have high spatial overlap and redundant pathological semantics, belonging to the repetitive representation of homologous features; when the cluster center spacing is greater than or equal to 1.2, it indicates that the two regions are spatially independent, have significant semantic differences, and have no redundant interference. If the cluster center spacing of any two initially selected pooling domains is less than... If the two regions are found to be redundant regions with the same origin, then the total vector enrichment of the two redundant regions is compared. The high-quality core regions with more vectors, higher feature enrichment, and greater semantic weight are retained, while the weak redundant regions are eliminated.

[0076] After removing redundant regions, three spatially independent, non-overlapping, semantically differentiated, and feature-reliable high-quality pathological regions were retained and identified as the effective aggregated reference regions for the entire domain. The semantic center position of each reference region was calculated using the mean of all embedded vectors within each effective aggregated reference region, representing the benchmark position of each core pathological region. Based on the total vector enrichment of the three effective aggregated reference regions from largest to smallest, a hierarchical sort was performed, and the cluster center positions corresponding to the three were successively labeled as the first reference position, the second reference position, and the third reference position, completing the hierarchical division of the core pathological benchmark positions for the entire domain and distinguishing between core and secondary pathological regions. The reference region with the highest total vector enrichment and the strongest pathological feature significance corresponds to the core pathological region, while the remaining two reference regions with progressively decreasing vector enrichment and relatively low feature weights correspond to secondary pathological regions.

[0077] This embodiment distinguishes high-density feature points based on the number of neighboring coverage points, and then completes full embedding vector clustering through seed clustering and residual point nearest-neighbor classification strategies, which can distinguish between clustered lesion features and discrete noise features. A unified quantitative clustering standard is used to avoid subjective errors caused by manual clustering, achieving automatic grouping of pathological features within a shared embedding space and identifying clustered regions of homologous pathological features. A quantitative compactness metric is calculated within each cluster to objectively characterize the density of feature aggregation within the cluster, allowing for a direct distinction between tightly clustered core lesion feature clusters and loosely dispersed noise interference clusters. High-value compact clusters are selected based on the compactness metric threshold, and adjacent compact clusters are merged to form candidate pooling regions, filtering out low-compactness and discrete noise-type embedding clusters. Only candidate pooling regions containing real pathological information are retained, reducing the amount of invalid data computation and improving the accuracy of pathological region localization. The candidate pooling regions with the highest number of embedded vectors are selected, and redundant pooling regions that overlap or are too small are eliminated by using the pooling region spacing threshold. Three standardized pooling reference regions are then fixed and output. The mean value of the position within the cluster is taken as the center of the reference position, which can stably lock the three-level core pathological regions and form a first, second, and third reference position with distinct levels.

[0078] In a preferred embodiment of the present invention, step 3 includes:

[0079] Step 300: Calculate the spatial distances between the first and second reference positions, the first and third reference positions, and the second and third reference positions respectively to obtain the spatial deviation of each of the three reference positions; take the average of the spatial deviations of the three reference positions as the balancing distance; and take the length obtained by multiplying the maximum value of the spatial deviations of the three reference positions by a preset neighbor distance as the neighborhood expansion range, specifically including:

[0080] The three reference positions, after hierarchical calibration, are paired to form three independent reference position pairs. The Euclidean spatial deviation between each reference position pair is calculated to quantify the spatial dispersion and distribution differences between the core pathological regions of each group. The formulas for calculating the spatial deviation of the three groups are as follows: ;

[0081] In the formula These represent the spatial deviations of three sets of reference position pairs; The position vectors of the first, second, and third reference positions are taken in sequence. The arithmetic mean of the three sets of spatial deviations is taken as the global equilibrium distance, which is the unified criterion for determining the coupling state of the global feature space. The maximum value of the three sets of reference position deviations is extracted and combined with the fixed preset proximity distance determined in step 200. To calculate the maximum neighborhood extension range of the global field effect, and to define the effective spatial boundary of the subsequent field effect coverage and feature statistics, the formula for calculating the maximum neighborhood extension range is as follows: ;

[0082] in It is the maximum value among the three sets of reference position deviations. This represents the minimum deviation among the three sets of reference positions. By introducing half of the minimum reference position spacing as an upper limit constraint, we ensure that the statistical neighborhood of each reference position maintains spatial independence and avoids large-area overlap of the entire neighborhood due to a single maximum deviation.

[0083] Step 301: Define a neighborhood around the spatial location of each reference bit, count the number of embedded vectors covered in each neighborhood, and obtain the degree of clustering of embedded vectors in the neighborhood of each reference bit; determine the type of coupling strength between each reference bit pair based on the comparison relationship between the spatial deviation and the equilibrium spacing of each of the three reference bit pairs, wherein the spatial deviation is less than the equilibrium spacing and the coupling strength is greater than the equilibrium spacing, specifically including:

[0084] Taking the center position of each of the three hierarchical reference positions as the center of the sphere, and the maximum neighborhood expansion range of the entire domain calculated in the previous step, respectively... Using the radius of a sphere, the effective statistical neighborhood of each core pathological reference domain is defined, thus locking in the effective spatial range of the global field effect feature statistics. The total number of embedding vectors within the effective neighborhood of each reference position is counted one by one, quantifying the level of pathological feature aggregation and semantic weight corresponding to each reference position. The formula for the number of neighborhood vectors is as follows: ;

[0085] In the formula: For the first The total number of vectors in the neighborhood of each reference bit; For a global set of embedded vectors; This is the current reference position vector; This is a binary indicator function, with a value of 1 if the condition is true and 0 if the condition is false. The spatial deviation of each set of reference pairs is compared one by one with the global equilibrium spacing to quantitatively determine the spatial coupling type of each set of core pathological regions. Specifically, when the spatial deviation of a reference pair is less than the global equilibrium spacing, the region is considered to be attractively coupled, representing that the two sets of core pathological regions are spatially adjacent, feature-continuous, and pathologically semantically homologous; when the spatial deviation of a reference pair is greater than or equal to the global equilibrium spacing, the region is considered to be repulsively coupled, representing that the two sets of core pathological regions are spatially discrete, feature-heterogeneous, and pathologically semantically independent.

[0086] Step 302: Determine the magnitude of the coupling strength between each reference bit pair based on the degree of clustering of the embedded vectors in the neighborhood of each reference bit; the higher the degree of clustering of the embedded vectors, the larger the magnitude. Based on the type and magnitude of the coupling strength between each reference bit pair, taking the spatial location of each reference bit as the field source point, deduce the change in the type and magnitude of the coupling strength at each field source point with the increase of spatial distance layer by layer along each direction of space, and obtain the field distribution of each field source point. Superimpose the field distributions of all field source points to obtain a dynamic constraint field. The dynamic constraint field is used to maintain the consistency of the cross-center distribution of the embedded vectors uploaded by each participating center in the shared embedding space, specifically including:

[0087] Based on the total enrichment of neighborhood embedding vectors of three sets of reference pairs, the degree of correlation between pairwise pathological regions is quantitatively characterized. The more effective vectors in a region, the stronger the feature correlation and the more significant the spatial coupling between the two sets of pathological regions. The coupling strength of each set of reference pairs is calculated by combining the coupling coefficient. The specific calculation formula is as follows: ;

[0088] In the formula: For reference bit pair The coupling strength; The coupling coefficient is... , These represent the total number of embedding vectors within the neighborhood of the two paired reference bits, respectively. To determine... The values ​​of are used to calculate the mean of the number of embedded vectors N1, N2, and N3 in the neighborhood of the three reference positions. Let the number of neighborhood vectors equal to When the reference coupling strength is 1.0, according to the coupling strength calculation formula, when hour, ,have to In [0.7] 0, 1.3 Within the interval [0], candidate λ values ​​are traversed with a step size of 0.01. For each candidate value, the Euclidean distance from the embedded vector in the neighborhood of each reference position to the center of the reference position is calculated, and the standard deviation σ of the distance distribution is calculated. Then, the direction vector of each embedded vector in the neighborhood relative to the center of the reference position is calculated, and the distribution ratio of the direction vector in each quadrant of the three-dimensional space is calculated. The sum of the absolute deviations δ between this ratio and 1 / 8 (the theoretical ratio of each quadrant when uniformly distributed in three-dimensional space) is calculated. S = 1 / (σ·(1+δ)) is used as the structural regularity score. The higher the S, the more tightly the neighborhood embedded vectors are clustered around the center of the reference position and the more uniform the direction distribution. The value corresponding to the largest S is selected. Assuming the final fixed coefficients, =2.2.

[0089] The first, second, and third reference positions, after hierarchical division, are uniformly used as independent source points of the global dynamic constraint field. Combined with the attraction and repulsion coupling types determined in step 301, each source is given a differentiated control function. That is, the attraction coupling source, for the core area that is spatially adjacent and has the same pathological semantic origin, plays the role of bringing homogeneous features closer and strengthening the aggregation of effective pathological features; the repulsion coupling source, for the area that is spatially discrete and has heterogeneous features, plays the role of dispersing stray features and suppressing invalid noise interference.

[0090] The field intensity at each source point follows a spatial exponential decay law; the closer to the source, the stronger the field intensity constraint, and the farther away, the weaker the effect. This adapts to the distribution characteristics of MRI pathological features, which are locally concentrated and gradually change over the entire region, effectively avoiding feature discrimination errors caused by abrupt changes in field intensity. Combined with a fixed attenuation coefficient to control the attenuation rate, a smooth, continuous, and distortion-free single-source field intensity distribution is generated. The quantitative calculation of the field intensity of a single source is as follows: ;

[0091] In the formula: The spatial distance of the field source point The electric field strength value at that location; The attenuation coefficient is... This represents the coupling strength of the corresponding field source; This is the Euclidean distance from the spatial measurement point to the field source point. The principle for determining the value is to ensure that the effective range of each reference potential field source remains relatively independent in space, that is, to satisfy... ,in This is the minimum of the pairwise spacings of the three reference bits, to avoid excessive overlap of field sources at adjacent reference bits that could cause coupling types to cancel each other out. Let... Let the total mean of the pairwise spatial deviations of the three reference positions be the global mean. Calculate using the following formula:

[0092] This formula can simultaneously satisfy the requirement that the field strength decreases with the overall scale of the characteristic distribution (i.e., the spatial distance reaches...). The time field intensity decays to the source intensity. ), and the constraint of the relative independence of field source effects (i.e. (Ever established).

[0093] To achieve unified constraint control of global features, the single-source field intensity distributions of three independent field sources are linearly superimposed and fused to generate a global dynamic constraint field that covers the entire shared embedding space, has unified parameters, and is continuous and smooth. This upgrades single-point feature control to unified global feature correction. The formula for calculating the total global field intensity superposition is as follows:

[0094] In the formula To share any spatial location vector in the embedding space, For the first Spatial position vector of reference positions For the first Each source is at a distance The field strength values ​​at the location. The final constructed global dynamic constraint field actively aggregates co-source effective pathological features and suppresses heterogeneous discrete noise features through the differentiation of positive and negative field strengths; it continuously offsets the feature space offset problem caused by differences in hardware of multi-center equipment and uneven sample distribution.

[0095] In a preferred embodiment of the present invention, step 4 includes:

[0096] Step 400: Based on the dynamic constraint field and the spatial position of the first reference position, using the spatial position of the first reference position as the calculation origin, calculate the rate of change of the coupling strength along each spatial orientation starting from the calculation origin. Determine the spatial orientation where the coupling strength changes most rapidly as the alternating orientation corresponding to the first reference position, and determine the rate of change of the coupling strength at the spatial orientation where the coupling strength changes most rapidly as the alternating amplitude corresponding to the first reference position. Specifically, this includes:

[0097] Using the highest-level and most representative first reference position as the origin for gradient calculation, and relying on the constructed global dynamic constraint field, the continuous gradient variation law of the field strength is solved to uncover the core change direction of the spatial evolution of pathological features. Traversing all directions in three-dimensional space, the instantaneous rate of change of the dynamic constraint field strength with spatial distance is calculated for each direction. The specific calculation formula is as follows: ;

[0098] In the formula Spatial orientation The corresponding field strength gradient rate is used to characterize the rate of change of field constraints and pathological features in the corresponding orientation; For any spatial orientation Spatial distance The total global field strength at the location; This is the Euclidean spatial distance between the current measurement point and the origin of the first reference point; It allows for arbitrary azimuth angles in three-dimensional space.

[0099] After traversal, the gradient rate values ​​corresponding to all spatial orientations are compared. The spatial orientation corresponding to the maximum gradient rate is determined as the optimal evolution orientation of the first reference position. This orientation corresponds to the core extension direction where the semantic changes of pathological features are most significant and the evolution of field constraint regulation is most intense across the entire domain. Simultaneously, the maximum gradient rate value across the entire domain is precisely defined as the evolution amplitude of the first reference position. This parameter is used to quantitatively quantify the degree of difference in the distribution of pathological features in the target region and the intensity of the regulatory changes in the dynamic constraint field. The larger the evolution amplitude value, the more obvious the gradient difference in the spatial distribution of pathological features under this core orientation, the higher the distinction between homologous and heterogeneous features, and the greater the intensity of the regulatory changes of the dynamic constraint field on surrounding features. Conversely, the smaller the evolution amplitude value, the more gradual the distribution of pathological features in this orientation and the more stable the field constraint regulation state.

[0100] Step 401: Based on the variation amplitude corresponding to the first reference position, compare the variation amplitude corresponding to the first reference position with the upward adjustment threshold and the downward adjustment threshold. When the variation amplitude is higher than the upward adjustment threshold, multiply the initial scan radius corresponding to the first reference position by a preset first attenuation coefficient to obtain the reduced scan radius. When the variation amplitude is lower than the downward adjustment threshold, multiply the initial scan radius corresponding to the first reference position by a second expansion coefficient to obtain the expanded scan radius. When the variation amplitude is between the upward adjustment threshold and the downward adjustment threshold, the initial scan radius corresponding to the first reference position remains unchanged, thus obtaining the scan radius adjusted by the first reference position. Specifically, this includes:

[0101] The field strength gradient rate at all spatial locations across the entire domain is statistically analyzed to obtain a global gradient dataset. The global mean and global standard deviation of the gradient are then calculated. A fixed quantile selection method for the global gradient data is used to determine the critical threshold of the gradient. This involves sorting all gradient rate values ​​in ascending order and selecting the 90th quantile as the upper critical threshold and the 50th quantile as the lower critical threshold. Assuming an upper critical threshold for the gradient is determined... Lower critical threshold It can quickly distinguish between regions with high changes in pathological feature gradient mutations and regions with low changes in gradient.

[0102] The contraction and expansion coefficients were determined using a quantitative scoring method. With all network parameters and field constraints fixed, the tests were iterated sequentially within the contraction coefficient range of 0.5 to 0.9 and the expansion coefficient range of 1.1 to 1.5. Three core evaluation indicators corresponding to each coefficient group were collected, including the detection rate of subtle pathological features. The percentage of accurate capture of features representing subtle lesions. The number of true positive samples represents the number of genuine subtle pathological features that were correctly identified. The number of false negative samples represents the number of feature points that have true pathological characteristics but were not identified and were missed; the completeness of the global feature traversal. In the formula The actual number of pathological feature points traversed. The total number of effective pathological feature points across the entire domain represents the degree of feature coverage; background noise false positive rate. In the formula The number of falsely detected background noise points, The total number of sampling points represents the degree of noise interference. An equally weighted comprehensive scoring formula is used to quantify the overall merits and demerits, resulting in a comprehensive score. In the formula , It is a positive indicator; the higher the value, the better the performance. As a negative indicator, lower values ​​indicate better performance, and higher overall scores represent better coefficient fit. Through batch sample traversal and comparison, the shrinkage coefficient is assumed to be... Expansion coefficient The combination of coefficients corresponds to the highest overall score and is therefore determined to be the optimal combination.

[0103] The shrinkage coefficient appropriately compresses the scanning radius for fine pathological regions with high gradient mutations, reduces the feature sampling neighborhood, eliminates irrelevant background feature interference, and improves the ability to capture fine features such as subtle pathological textures and lesion edge differences, thus meeting the high-precision detection requirements of pathological mutation regions. The expansion coefficient appropriately enlarges the scanning radius for continuous pathological regions with low gradient and gentle slopes, expands the feature sampling coverage, and fully encompasses a large range of continuously distributed pathological feature regions, effectively avoiding problems such as feature omissions, pathological region truncation, and missing feature connections, and ensuring the integrity of full-domain feature traversal.

[0104] An adaptive adjustment mechanism for the scanning radius is constructed to achieve an adaptive match between accuracy and efficiency. The specific adjustment formula is as follows: ;

[0105] In the formula: The initial scanning radius (with a value of 1.0) corresponding to the first reference position is used to adapt the basic sampling range of this MRI pathological feature space. The variation amplitude of the first reference position is calculated by the aforementioned step 400, and the value of the maximum field strength gradient rate corresponding to the first reference position in the entire domain. This is the final scan radius after adaptive adjustment. When the change amplitude is greater than the upper critical threshold, it is determined to be a region of abrupt change in feature gradient, and the scan radius is shrunk to achieve fine scanning and capture subtle pathological features; when the change amplitude is less than the lower critical threshold, it is determined to be a region of gentle feature gradient, and the scan radius is expanded to improve the efficiency of traversing the entire feature domain; when the change amplitude is within the range of both thresholds, the initial scan radius is kept unchanged to balance scanning accuracy and computational efficiency.

[0106] Step 402: Based on the dynamic constraint field, the spatial position of the second reference position, and the spatial position of the third reference position, sequentially obtain the alternating azimuth, alternating amplitude, and adjusted scanning radius corresponding to the second reference position, as well as the alternating azimuth, alternating amplitude, and adjusted scanning radius corresponding to the third reference position, specifically including:

[0107] Reusing the three-dimensional spatial orientation traversal rules and field intensity gradient rate calculation formula from step 400, the optimal changing orientation and changing amplitude corresponding to the second and third reference positions are calculated respectively. Simultaneously, using the fixed quantile gradient threshold, contraction coefficient, expansion coefficient, and adaptive scanning radius adjustment formula from step 401, adaptive radius correction is performed on the two secondary reference positions. This step involves no new parameters or custom judgment rules. Through standardized and unified calculations, the optimal changing orientation, changing amplitude, and adaptive optimal scanning radius corresponding to the second and third reference positions are finally obtained.

[0108] This embodiment uses the first reference position as the origin to traverse the entire spatial orientation and calculate the rate of change of the dynamic constraint field, automatically locking the most drastic changes in pathological features and quantifying the magnitude of these changes. It accurately locates the extension direction of concentrated lesion texture and edge details, providing a unique spatial scanning direction and gradient quantization basis for adaptive differentiated scanning and fine lesion sampling. Dual thresholds for upward and downward adjustment are set, and the initial scanning radius is adaptively scaled according to the magnitude of the changes: the radius is reduced in high-gradient abrupt change regions to achieve dense and fine sampling of subtle lesions; the radius is expanded in low-gradient flat regions to achieve efficient traversal over a large area; and the intermediate interval maintains a baseline radius to balance accuracy and computational power, taking into account both the ability to capture subtle pathological features and the efficiency of global scanning operations. Steps 400 and 401 are reused to unify the gradient calculation and radius adaptive adjustment logic to solve for the parameters of the second and third reference positions. All three levels of reference positions use completely consistent quantization judgment standards to avoid algorithm judgment deviations caused by multiple sets of rules; the changing orientation, magnitude of change, and adaptive scanning radius of the three sets of reference positions are uniformly output to ensure a unified standard for the construction of the global scanning path.

[0109] In a preferred embodiment of the present invention, step 5 includes:

[0110] Step 500: Based on the changing azimuth corresponding to the first reference position and the adjusted scanning radius of the first reference position, taking the spatial position of the first reference position as the starting position for scanning, perform one step movement along the changing azimuth with the corresponding adjusted scanning radius. After reaching a step position, use the adjusted scanning radius corresponding to the step position as the moving distance for the next step movement, thus obtaining each step position arranged along the changing azimuth direction and the scanning radius corresponding to each step position. Specifically, this includes:

[0111] Based on the fixed parameters quantized and solved in steps 400 and 401, a global adaptive stepping position sequence is generated. All scanning directions and step sizes have clear calculation sources and quantitative judgment criteria, avoiding ambiguity in the protection range. Specifically, the optimal progressive orientation used for scanning advancement is calculated by traversing the global field intensity gradient in step 400. That is, it traverses the field intensity gradient rate in all orientations of the 3D embedded space, selecting the spatial orientation corresponding to the maximum gradient rate as the optimal progressive orientation. This is a fixed and unchanging global optimal scanning extension direction, ensuring that the scanning path always advances along the region with the most significant changes in pathological features. The scanning iteration origin is the first reference position after parameter tuning, and the step size of the first iteration uses the initial optimal scanning radius obtained through adaptive optimization in step 401.

[0112] The scanning process employs a quantization adaptive step size mechanism, with the step size calculated in real time using a gradient adaptive matching formula. In the formula For the current iteration step to progress, The adaptive optimal scanning radius corresponding to the current spatial position is quantized and output by the double threshold determination formula in step 401. The gradient threshold is fixed according to step 401. Divide the region and determine the local gradient rate. When a region is identified as a high-gradient abrupt change, the algorithm automatically matches the shrunken, smaller scan radius to achieve dense, fine sampling of lesion edges and areas with subtle textures, preserving detailed features; when the local gradient rate... When the gradient rate is determined to be in a low-gradient, flat region, the algorithm automatically matches the expanded large scan radius to achieve efficient global traversal of large, uniform pathological areas, thus improving scanning efficiency; when the gradient rate is within the range At the same time, the baseline radius remains unchanged. The step size is updated point by point through this quantitative partitioning adaptation mechanism, and all pathological feature extension areas and transition areas are completely traversed, finally generating a global adaptive scanning step position sequence with unique direction, quantitative scale, and adapted to the gradient distribution of lesions.

[0113] Step 501: Based on each step position and the adjusted scanning radius corresponding to each step position, successively compare the spatial distance between the spatial position of the second reference position and the current step position with the adjusted scanning radius corresponding to the current step position. The step position where the spatial distance is first less than or equal to the adjusted scanning radius corresponding to the current step position is determined as the first capture position. The entire scanning trajectory from the spatial position of the first reference position to the first capture position is determined as the first coupling path, specifically including:

[0114] Based on the global adaptive stepping position sequence generated in step 500, this method achieves the association and coupling of the core pathological regions of the first and second reference positions through quantized spatial distance matching rules, constructing a hierarchically coherent first coupling path. For each temporal stepping point, the three-dimensional Euclidean spatial distance between the current stepping point and the center of the second reference position is quantitatively calculated. A unique quantized coupling determination condition is adopted: When a stepping point meets the conditions of this formula, it is determined that the adaptive sampling neighborhood of the current point can completely cover the secondary core pathological region corresponding to the second reference point, and the two regions form an effective association of homologous pathological features. Points that meet this quantitative judgment condition for the first time during the full-domain scanning process are selected and marked as the first capture position. All temporal continuous scanning trajectories from the origin of the first reference point to the first capture position are integrated to construct the first coupling path connecting the core pathological region and the secondary core pathological region, realizing the orderly association and integration of cross-level pathological features.

[0115] Step 502: Based on the first capture position and the corresponding scanning radius, taking the first capture position as the starting point for continued scanning, continue stepping along the changing orientation with the adjusted scanning radius corresponding to the first capture position. The spatial distance between the spatial position of the third reference position and the current stepping position is compared with the adjusted scanning radius corresponding to the current stepping position. The stepping position where the first spatial distance is less than or equal to the adjusted scanning radius corresponding to the current stepping position is determined as the second capture position. The entire scanning trajectory from the spatial position of the first reference position to the second capture position is determined as the second coupling path, specifically including:

[0116] Based on the scanning of the first coupling path, the iteration continues, maintaining a unified global scanning direction and adaptive step size update rule. Starting from the first capture position, the iteration extends globally, completing the global traversal coupling of the three-level reference positions. The three-dimensional Euclidean space distance between the newly added step point and the center of the third reference position is calculated sequentially. The quantization criteria for the second-level coupling are as follows: The step point that first meets the matching condition across the entire region is located and marked as the second capture position. The complete continuous scan trajectory from the initial scan origin to the second capture position is integrated to construct a second coupling path that simultaneously connects the first, second, and third reference positions. These two coupling paths are hierarchically complementary and follow unified rules, fully covering the full-level pathological features of the core lesion, transitional areas, and peripheral extension areas. This avoids the feature truncation, regional disconnection, and edge detection problems caused by traditional segmented scanning, achieving full coverage and correlation of the entire pathological region.

[0117] In this embodiment, starting from the first reference position and using the optimal gradient orientation as the scanning direction, the step distance for each step is dynamically updated with the adaptive scanning radius following the current position. This achieves automatic adaptation of the scanning step length to the local pathological gradient, with dense sampling in high-detail areas and sparse sampling in flat areas, generating a continuous sequence of stepping points that closely matches the distribution of actual lesions. The first capture position is automatically captured by the spatial distance between the stepping point and the second reference position and the size of the scanning radius, constructing a first coupling path connecting the core and secondary core pathological regions. This connects the two levels of pathological feature regions, establishing a spatial association between core lesions and transitional lesions, avoiding fragmented extraction of pathological features at both levels. The scanning trajectory is continued iteratively to capture the second capture position, generating a second coupling path that connects three levels of reference positions, completely covering the core, transitional, and edge pathological regions. The two coupling paths are complementary, eliminating feature truncation and missed detection defects in lesion edges and distal extension areas, achieving seamless association of the entire pathological region.

[0118] In a preferred embodiment of the present invention, step 6 includes:

[0119] Step 600: Extract the dynamic constraint field strength vector value at each step position along the first coupling path and the second coupling path, and extract the mean and covariance statistics of all embedded vectors covered within the neighborhood radius of the corresponding step position to form a local feature vector, specifically including:

[0120] Refined local feature extraction is performed on all temporal step points of the two coupled paths one by one. The field intensity vector is solved by quantization using a global dynamic constraint field, and a local feature sample set is constructed by combining image embedding features. Specifically, for each three-dimensional step coordinate position on the path, the global dynamic constraint field constructed in step 302 is called to quantize and solve the field intensity vector at the current position, which is calculated by the single-source field intensity attenuation formula in step 302. This formula can accurately quantify the field constraint strength and spatial correlation direction of each sampling position, and completely preserve the coupling correlation information between multi-center offset correction and pathological region.

[0121] Adaptive scanning radius at each step point To define the neighborhood truncation scale, the local sampling space is limited, and all valid MRI image embedding feature vectors within this neighborhood are collected to form the original image feature set for the current location. This enables gradient-adaptive, differentiated, and refined sampling. To accurately quantify the distribution characteristics of local pathological features, statistical parameters are solved based on the original feature set, including the local feature mean vector, which characterizes the overall central distribution and apparent characteristics of pathological features in the current sampling area; and the local feature covariance matrix, which quantifies the dispersion, texture differences, and feature correlation distribution of local features.

[0122] The field strength vector, local feature mean vector, and local feature covariance matrix flattened from the single-point solution are concatenated in a fixed-dimensional order according to spatial constraint features, texture appearance features, and feature distribution features to generate a standardized local feature vector for each single point. This feature vector integrates spatial correlation correction information of the global dynamic constraint field, texture appearance information of local lesions, and local feature distribution difference information, providing a complete representation dimension. Following the scanning step sequence of the two coupled paths, the standardized local feature vectors of all points are sequentially arranged and collected to obtain complete and refined local feature sequences corresponding to the first and second coupled paths, respectively.

[0123] Step 601: Based on the sequential arrangement of each step position in the local feature sequence of the first path along the scanning direction, the local features of the first step position are taken as the intermediate result of the first path merging; along the subsequent order of the scanning direction, the local features of the current step position are taken one by one, and the local features of the current step position are merged with the intermediate result of the first path merging, including: using the ratio of the scanning radius corresponding to the current step position to the sum of the scanning radii corresponding to all step positions on the first coupled path as the merging coefficient, and merging the local features of the current step position into the intermediate result of the first path merging according to the merging coefficient to obtain the updated intermediate result of the first path merging; the intermediate result of the first path merging when all step positions on the first coupled path have been processed is taken as the first path merging feature; the second path merging feature is obtained by processing all step positions on the second coupled path corresponding to the local feature sequence of the second path, specifically including:

[0124] For the first path local feature sequence and the second path local feature sequence output in step 600, a stepwise weighted merging strategy along the scanning time sequence is adopted to iteratively fuse the standardized local features of discrete time sequence distribution within the path, generating the first path merged feature and the second path merged feature containing multi-scale path path information.

[0125] For the first coupled path, the standardized local features corresponding to the first step position in the local feature sequence of the first path are extracted and directly initialized as the intermediate result of the first path merging, completing the initial state assignment for iterative fusion. Based on this, along the fixed temporal order of the scanning direction, all remaining step positions of the path are traversed sequentially, and the standardized local features corresponding to the current step position are extracted one by one. These are then weighted and iteratively fused with the real-time updated intermediate result of the first path merging, realizing the step-by-step merging and updating of features.

[0126] Adaptive radius proportion quantization of merging coefficients, original merging coefficients The calculation formula is: ;

[0127] In the formula The original merge coefficient corresponding to the current traversal step position; The real-time optimal scanning radius is obtained after adaptive adjustment of the current step position; For the first coupling path within the first The adaptive scan radius corresponding to each step position; This represents the total number of all step positions contained in the first coupling path.

[0128] To prevent a few extremely large scan radii in the path from dominating the entire merging result, an upper limit threshold is set for the merging coefficient of a single point. If all Then directly order If it exists Then the following recursive truncation and reallocation process will be executed:

[0129] Locate the step position with the largest original merge coefficient. ,Right now ,like If so, the truncation process will terminate, and all current processes will be terminated. That is, the final merging coefficient; if Then let Calculate the cut-off margin .Will Redistribute according to the current radius ratio of the remaining steps, for : ;use Update all Return to the previous steps to determine if the problem still exists. until all merging coefficients do not exceed .

[0130] After calculating the merging coefficients, a fixed step-by-step iterative merging formula is used to update intermediate features during successive iterations. The iterative merging calculation formula is as follows:

[0131] In the formula For the first The first path merging intermediate results obtained before the next iteration have accumulated and saved the fused pathological feature information of all step points before the current traversal position; For the current number The standardized local feature vectors corresponding to each step position; For the current number The final merging coefficient of each step position after the above truncation and redistribution process; This is the newly obtained intermediate result of the first path merging after this weighted iteration update. After each single-point feature merging is completed, the new merging intermediate result is immediately updated and retained, replacing the historical intermediate result to participate in the next round of iteration calculation, and the point-by-point traversal and feature fusion of all step positions of the first coupled path are completed in sequence.

[0132] After all step positions of the first coupled path have been processed and the iteration has fully converged, the final stable intermediate result of the first path merging is determined as the first path merging feature. This merging feature is... Its dimension is the same as that of a single-point local feature vector, but it incorporates information from each step position across the entire path.

[0133] By reusing the entire set of temporal traversal rules, merging coefficient calculation methods (including recursive truncation and redistribution processes), and step-by-step iterative fusion mechanism of the first coupling path, the second path's local feature sequences are synchronously and equally quantized at all step positions, and finally iteratively converge to obtain the merged features of the second path. The two merge features are calculated independently, and normalization is performed within each path without interference, ensuring that they do not interfere with each other. and All of them are within a uniform feature scale range before subsequent splicing.

[0134] Step 602: The first path merging feature and the second path merging feature are concatenated and fused along the feature dimension to obtain a global pathological feature description; the global pathological feature description is distributed to each participating center, and each participating center uses the global pathological feature description as an update reference to drive the directional update of the parameters of the local coding network, specifically including:

[0135] The two single-path merged features output in step 601 are subjected to lossless dimensional fusion to integrate pathological feature information across the entire domain and all levels, constructing a unified standard global pathological feature representation. Furthermore, the collaborative correction and targeted iterative update of network parameters across multiple centers are achieved using a federated learning distributed architecture. Specifically, a dual-path global feature fusion operation is performed, fusing the first-path merged features with the second-path merged features through dimensional concatenation. The global feature fusion calculation formula is as follows: ;

[0136] In the formula: The path merging feature corresponding to the first coupling path carries pathological information of the core lesion and the secondary core transition area. The path merging feature corresponding to the second coupling path carries the pathological information of the global edge extension area and the distant associated pathological information. This is a channel-dimensional splicing operation used to integrate the complementary path features of two hierarchical levels into a single high-dimensional global feature without losing individual point details, spatial correlations, and distribution statistics. To generate a unified and comprehensive pathological feature across the entire domain, the system fully covers the multi-scale pathological features and field constraint association information of the core lesion, transitional area, and peripheral area.

[0137] Employing a privacy-preserving federated collaborative training architecture, all raw MRI image data from multiple multi-center branch nodes are stored locally throughout the entire process, never leaving the local terminal. Only the path merging features calculated by each branch are uploaded to the upstream central aggregation node. The central aggregation node uniformly receives data from all branch nodes. , Features, batch execution of global splicing and fusion to obtain standardized global pathological features. Then, the global standard feature is distributed to all multi-center branch nodes as a unified training and fitting supervision benchmark for all local networks, thereby eliminating the feature representation offset problem caused by heterogeneous multi-center devices, uneven sample distribution, and imaging gain deviation from the root.

[0138] Each multi-center branch node uses the globally fused features distributed by the network as the prediction output and the real pathological labels as the supervision benchmark to construct a standardized supervision loss function. This method is adapted to MRI pathological feature classification tasks and uses the cross-entropy loss function as the sole supervision loss function. This loss function is used to quantify the degree of deviation between the network's predicted pathological features and the real pathological labels, providing gradient supervision for parameter updates.

[0139] Furthermore, multi-center local network parameter directional updates are achieved through back gradient propagation. The network parameter iterative update formula is as follows: Fixed learning rate This avoids the problems of excessively large step size causing oscillations and excessively small step size causing slow convergence. These are all trainable parameters of the multicenter local MRI feature encoding network before and after the update, specifically including convolutional layer weights, convolutional bias terms, normalization layer scaling parameters, and offset parameters. To monitor the back gradient of the loss function with respect to the network parameters, this method quantifies the direction and magnitude of parameter optimization. Each branch node completes a single parameter iteration update based on the aforementioned fixed formula and fixed hyperparameters, retaining effective feature representations adapted to the local data distribution while forcibly aligning with a globally unified pathological feature standard. Multiple rounds of federated feature fusion, loss calculation, and gradient update iterations are repeatedly executed to continuously correct cross-center feature space offsets, ultimately achieving unified logic for multi-center network pathological feature extraction and simultaneous improvement in the accuracy of subtle lesion identification and the model's cross-center generalization ability.

[0140] In this embodiment, the dynamic constraint field strength vector, the mean of the neighborhood embedding vector, and the covariance statistics of each step point are simultaneously extracted and fused into a local feature vector. This vector simultaneously carries spatial constraint correction information, lesion appearance features, and feature discrete distribution information. A single vector comprehensively represents local multi-scale pathological information, ensuring complete feature dimensionality and reducing the risk of information loss during subsequent feature fusion. Along the scanning time sequence, the scanning radius ratio is used as the merging coefficient to weight the local features of the fusion path. Lesion regions with larger coverage areas have higher weights, highlighting core lesion features while preserving subtle edge details. Two path-normalized merged features are generated independently to unify the feature representation scale of a single path, avoiding feature imbalance caused by differences in sampling scales at different points. The merged features of the two paths are concatenated to obtain a global pathological feature description containing pathological information at all levels, which is uniformly distributed to all participating centers as a global standard. Each center updates its local encoding network parameters based on the unified global features, correcting feature offsets caused by cross-center imaging and sample distribution, unifying the multi-center feature extraction logic, and improving the model's generalization ability to identify subtle lesions across centers.

[0141] like Figure 2 As shown, embodiments of the present invention also provide a multi-center medical image data collaborative training system based on federated learning, comprising:

[0142] The first module is used by each participating center to train a local coding network with local multimodal MRI images, map the samples to a shared embedding space, obtain embedding vectors, and upload the embedding vectors to the central aggregation node.

[0143] The second module is used by the central aggregation node to determine the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector;

[0144] The third module is used to calculate the spatial deviation between each reference bit pair and determine the coupling strength between each reference bit pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference bit, so as to construct a dynamic constraint field.

[0145] The fourth module is used to take the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, calculate the variation amplitude and variation orientation of the coupling strength at each reference position in the dynamic constraint field, and adjust the generated scanning radius according to the variation amplitude.

[0146] The fifth module is used to perform a step-by-step scan with a scanning radius along a changing orientation, starting from the first reference position. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, a first coupling path is established between the second reference position and the first reference position. When the third reference position falls into the scanning radius, a second coupling path is established between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position spatial position.

[0147] The sixth module is used to merge the local features of each reference position along the first and second coupling paths step by step to obtain a global pathological feature description that is invariant across the central domain, and distribute it to each participating center to drive the parameter directionality update of the local coding network.

[0148] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0149] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-center medical image data collaborative training method based on federated learning, characterized in that, The method includes: Step 1: Each participating center trains a local coding network using local multimodal MRI images, maps the samples to a shared embedding space, obtains embedding vectors, and uploads the embedding vectors to the central aggregation node; Step 2: The central aggregation node determines the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector; Step 3: Calculate the spatial deviation between each reference pair, and determine the coupling strength between each reference pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference pair, so as to construct a dynamic constraint field; Step 4: Using the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, the amplitude and orientation of the changing coupling strength at each reference position are calculated in the dynamic constraint field, and the scanning radius is adjusted according to the amplitude. Step 5: Starting from the first reference position, perform a step-by-step scan along the changing orientation with a scanning radius. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, establish a first coupling path between the second reference position and the first reference position. When the third reference position falls into the scanning radius, establish a second coupling path between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position in space. Step 6: Along the first coupling path and the second coupling path, the local features of each reference position are merged step by step to obtain a global pathological feature description that is invariant across the central domain. This description is then distributed to each participating center to drive the parameter directionality update of the local coding network.

2. The multi-center medical image data collaborative training method based on federated learning according to claim 1, characterized in that, Step 1 includes: Each participating center performs inter-slice spatial registration on local multimodal MRI images to obtain registered multi-sequence voxel data; and constructs local training batches containing multiple samples based on the registered multi-sequence voxel data. Based on the local training batch, perform forward and backward propagation on the local coding network to update the parameters of the local coding network, and obtain the local coding network updated in this round of training; Each sample in the local training batch is input into the local encoding network updated by this round of training to obtain the embedding vector corresponding to each sample; the embedding vectors corresponding to each sample are aggregated and sent to the central aggregation node.

3. The multi-center medical image data collaborative training method based on federated learning according to claim 2, characterized in that, Step 2 includes: Based on the spatial distribution of the embedding vectors in the shared embedding space, all embedding vectors in the shared embedding space are divided into multiple embedding vector clusters according to their spatial proximity. For each cluster of embedded vectors, the spatial compactness between the embedded vectors within the corresponding cluster is calculated to obtain the cluster compactness metric value of the corresponding cluster. The higher the cluster compactness metric value, the closer the embedded vectors within the corresponding cluster are to each other. Based on the compactness metric of each embedded vector cluster, embedded vector clusters with compactness metrics higher than the compactness threshold are extracted as compact clusters; compact clusters that are spatially adjacent to each other in the shared embedding space are merged to obtain several candidate aggregation domains. Based on several candidate pooling domains, the number of embedded vectors covered in each candidate pooling domain is counted, and the three candidate pooling domains with the largest number of embedded vectors are determined as the initial pooling domains. If the spatial distance between any two initial pooling domains is less than the pooling domain distance threshold, the one with the smaller number of embedded vectors is removed. The three final retained initial pooling domains are determined as pooling reference domains, and the average position of all embedded vectors in each pooling reference domain in the shared embedding space is taken as the center position of the corresponding pooling reference domain. The center positions of the three pooling reference domains are sequentially determined as the first reference position, the second reference position, and the third reference position.

4. The multi-center medical image data collaborative training method based on federated learning according to claim 3, characterized in that, Based on the spatial distribution of the embedding vectors in the shared embedding space, all embedding vectors in the shared embedding space are divided into multiple embedding vector clusters according to their spatial proximity, including: Using each embedding vector as the center, calculate the number of remaining embedding vectors covered within a preset neighborhood distance range to obtain the neighborhood coverage of each embedding vector; Based on the neighborhood coverage of each embedding vector, embedding vectors whose neighborhood coverage reaches the minimum neighborhood coverage threshold are marked as high-density embedding vectors; using each high-density embedding vector as a seed, high-density embedding vectors with various subspace spacings less than the preset neighbor distance are grouped into the same embedding vector cluster, resulting in an embedding vector cluster composed of high-density embedding vectors. Based on the embedding vector clusters composed of high-density embedding vectors, the remaining embedding vectors that have not been assigned to any embedding vector clusters are assigned one by one to the already formed clusters with the smallest spatial distance to the remaining embedding vectors that have not been assigned to any embedding vector clusters, thus obtaining multiple embedding vector clusters corresponding to all embedding vectors.

5. The multi-center medical image data collaborative training method based on federated learning according to claim 4, characterized in that, Step 3 includes: Calculate the spatial distances between the first and second reference positions, the first and third reference positions, and the second and third reference positions respectively to obtain the spatial deviation of each of the three reference positions; The average of the spatial deviations of the three reference positions is taken as the balance distance; the length obtained by multiplying the maximum value of the spatial deviations of the three reference positions with the preset neighbor distance is the neighborhood expansion range. A neighborhood is defined around the spatial location of each reference position, and the number of embedded vectors covered in each neighborhood is counted to obtain the degree of clustering of embedded vectors in the neighborhood of each reference position. Based on the comparison between the spatial deviation of each of the three reference positions and the balance distance, the type of coupling strength between each reference position pair is determined. When the spatial deviation is less than the balance distance, the coupling strength is attractive; when the spatial deviation is greater than the balance distance, the coupling strength is repulsive. The magnitude of the coupling strength between each reference bit pair is determined based on the degree of clustering of the embedding vectors in the neighborhood of each reference bit. The higher the degree of clustering of the embedding vectors, the greater the magnitude. Based on the type and magnitude of the coupling strength between each reference pair, taking the spatial position of each reference as the field source point, the changes in the type and magnitude of the coupling strength at each field source point with the increase of spatial distance are deduced layer by layer along each direction of space, so as to obtain the field distribution of each field source point. The field effect distributions of all field source points are superimposed to obtain a dynamic constraint field, which is used to maintain the consistency of cross-center distribution of the embedding vectors uploaded by each participating center in the shared embedding space.

6. The multi-center medical image data collaborative training method based on federated learning according to claim 5, characterized in that, Step 4 includes: Based on the dynamic constraint field and the spatial position of the first reference position, the spatial position of the first reference position is used as the calculation origin. The rate of change of the coupling strength along each spatial orientation starting from the calculation origin is calculated one by one. The spatial orientation with the fastest change in coupling strength is determined as the incremental orientation corresponding to the first reference position. The rate of change of coupling strength at the spatial orientation with the fastest change in coupling strength is determined as the incremental amplitude corresponding to the first reference position. Based on the variation amplitude corresponding to the first reference position, the variation amplitude corresponding to the first reference position is compared with the upward adjustment threshold and the downward adjustment threshold. When the variation amplitude is higher than the upward adjustment threshold, the initial scanning radius corresponding to the first reference position is multiplied by the preset first attenuation coefficient to obtain the reduced scanning radius. When the variation amplitude is lower than the downward adjustment threshold, the initial scanning radius corresponding to the first reference position is multiplied by the second expansion coefficient to obtain the expanded scanning radius. When the variation amplitude is between the upward adjustment threshold and the downward adjustment threshold, the initial scanning radius corresponding to the first reference position remains unchanged, thus obtaining the scanning radius after the first reference position is adjusted. Based on the dynamic constraint field, the spatial position of the second reference position, and the spatial position of the third reference position, the alternating orientation, alternating amplitude, and adjusted scanning radius corresponding to the second reference position are obtained in sequence, as well as the alternating orientation, alternating amplitude, and adjusted scanning radius corresponding to the third reference position.

7. The multi-center medical image data collaborative training method based on federated learning according to claim 6, characterized in that, Step 5 includes: Based on the changing orientation corresponding to the first reference position and the scan radius adjusted by the first reference position, the spatial position of the first reference position is taken as the starting position of the scan. A step movement is performed along the changing orientation with the corresponding adjusted scan radius. After reaching a step position, the adjusted scan radius corresponding to the step position is taken as the movement distance of the next step movement, so as to obtain each step position arranged along the changing orientation direction and the scan radius corresponding to each step position. Based on each step position and the adjusted scanning radius corresponding to each step position, the spatial distance between the spatial position of the second reference position and the current step position is compared with the adjusted scanning radius corresponding to the current step position. The step position where the spatial distance is less than or equal to the adjusted scanning radius corresponding to the current step position is determined as the first capture position. The entire scanning trajectory from the spatial position of the first reference position to the first capture position is determined as the first coupling path. Based on the first capture position and the corresponding scanning radius, taking the first capture position as the starting point for continued scanning, the system continues to move in a stepping motion along the changing orientation with the adjusted scanning radius corresponding to the first capture position. The spatial distance between the spatial position of the third reference position and the current stepping position is compared with the adjusted scanning radius corresponding to the current stepping position. The stepping position where the spatial distance is less than or equal to the adjusted scanning radius corresponding to the current stepping position is determined as the second capture position. The entire scanning trajectory from the spatial position of the first reference position to the second capture position is determined as the second coupling path.

8. The multi-center medical image data collaborative training method based on federated learning according to claim 7, characterized in that, Step 6 includes: The dynamic constraint field strength vector value at each step position and the mean and covariance statistics of all embedded vectors covered within the neighborhood radius of the corresponding step position are extracted one by one along the first coupling path and the second coupling path respectively to form a local feature vector. Based on the local feature sequence of the first path and the local feature sequence of the second path, along the scanning direction from the scanning start end to the scanning end along the first coupled path or the second coupled path, the local feature vectors at each step position are used as processing units to perform merging processing one by one to obtain the first path merging feature and the second path merging feature of the first coupled path and the second coupled path. The first path merged features and the second path merged features are spliced ​​and fused along the feature dimension to obtain a global pathological feature description; The global pathological feature description is distributed to each participating center, and each participating center uses the global pathological feature description as an update reference to drive the directional update of the parameters of the local coding network.

9. The multi-center medical image data collaborative training method based on federated learning according to claim 8, characterized in that, Based on the local feature sequences of the first and second paths, along the scanning direction from the start to the end of the scan along the first or second coupled path, the local feature vectors at each step position are used as processing units to perform merging processing one by one, resulting in the first path merging features and the second path merging features of the first and second coupled paths, including: Based on the order of each step position in the local feature sequence of the first path from first to last along the scanning direction, the local features of the first step position are used as the intermediate result of the first path merging. Following the subsequent sequence along the scanning direction, local features at the current step position are taken one by one, and the local features at the current step position are merged with the intermediate results of the first path merging. This includes: using the ratio of the scanning radius corresponding to the current step position to the sum of the scanning radii corresponding to all step positions on the first coupled path as the merging coefficient, and merging the local features at the current step position into the intermediate results of the first path merging according to the merging coefficient to obtain the updated intermediate results of the first path merging. The intermediate result of the first path merging after all step positions on the first coupled path have been processed is used as the first path merging feature; the second path merging feature is obtained by processing all step positions on the second coupled path corresponding to the local feature sequence of the second path.

10. A multi-center medical image data collaborative training system based on federated learning, wherein the system implements the method as described in any one of claims 1 to 9, characterized in that, include: The first module is used by each participating center to train a local coding network with local multimodal MRI images, map the samples to a shared embedding space, obtain embedding vectors, and upload the embedding vectors to the central aggregation node. The second module is used by the central aggregation node to determine the first reference position, the second reference position, and the third reference position based on the spatial distribution of the embedded vector; The third module is used to calculate the spatial deviation between each reference bit pair and determine the coupling strength between each reference bit pair by combining the degree of aggregation of embedded vectors in the neighborhood of each reference bit, so as to construct a dynamic constraint field. The fourth module is used to take the standard deviation of the embedded vector distribution in the neighborhood of each reference position as the initial value of the scanning radius of each reference position, calculate the variation amplitude and variation orientation of the coupling strength at each reference position in the dynamic constraint field, and adjust the generated scanning radius according to the variation amplitude. The fifth module is used to perform a step-by-step scan with a scanning radius along a changing orientation, starting from the first reference position. The scanning trajectory is modulated by a dynamic constraint field to track the feature continuous manifold in the shared embedded space. During the scan, a first coupling path is established between the second reference position and the first reference position. When the third reference position falls into the scanning radius, a second coupling path is established between the third reference position and the first reference position. The endpoint of the coupling path is the step position when the scanning radius covers the corresponding reference position spatial position. The sixth module is used to merge the local features of each reference position along the first and second coupling paths step by step to obtain a global pathological feature description that is invariant across the central domain, and distribute it to each participating center to drive the parameter directionality update of the local coding network.