An elderly gynecological ultrasound intelligent diagnosis system fusing multi-modal data
By employing multimodal data fusion techniques involving cross-modal coding, hierarchical decomposition, and dynamic feature routing, the system addresses the lack of dynamism and hierarchy in geriatric gynecological ultrasound intelligent diagnostic systems, enabling high-precision identification and reliable diagnosis of complex lesions in elderly patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MEI HOSPITAL UNIV OF CHINESE ACAD OF SCI
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-16
AI Technical Summary
Existing intelligent ultrasound diagnostic systems for geriatric gynecology lack dynamism and hierarchy in multimodal data fusion, and cannot effectively integrate imaging and clinical physiological parameters, resulting in limited ability to identify complex lesions in elderly patients, and failing to fully consider the overall physiological state information under the theory of traditional Chinese medicine.
A cross-modal coding module is used to transform ultrasound images and clinical physiological parameters into feature maps and semantic embedding vectors. A hierarchical decomposition module separates global topology, regional tissue and local microtexture features. A dynamic feature routing network adaptively reconstructs features and combines a multi-branch collaborative diagnostic network to perform confidence weighting and contextual correlation analysis to generate a structured diagnostic report.
It improves the accuracy and reliability of identifying complex pathological changes in elderly gynecological patients, and can adapt to the characteristics of different cases to generate more consistent and interpretable diagnostic results.
Smart Images

Figure CN122224461A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent medical diagnostic technology, specifically to an intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data. Background Technology
[0002] In the clinical diagnosis of gynecological diseases in the elderly, ultrasound imaging has become a primary method due to its non-invasive and convenient advantages. Current intelligent assisted diagnostic systems typically employ single-modal image analysis or simple multimodal data co-processing methods. Single-modal methods rely solely on ultrasound images for pattern recognition, failing to integrate key clinical and physiological parameters of the patient, resulting in insufficient information utilization. Multimodal co-processing methods often involve post-processing or linearly weighting of image features and clinical parameters; this superficial fusion mechanism struggles to uncover the complex nonlinear relationships and deep semantic interactions between different modalities.
[0003] Existing technical solutions have shortcomings. Their feature fusion process is static and fixed, unable to dynamically adjust the weights and interaction methods of different information sources according to the differences in specific cases. Conventional models process features in a flat way, lacking hierarchical encoding of medical prior knowledge, and cannot effectively separate and coordinate the macroscopic anatomical structure, mesoscopic regional characteristics, and microscopic texture information of lesions. This results in limited ability of the model to identify complex lesions that are common in elderly patients and have atypical imaging manifestations, and insufficient discriminative power of feature representation.
[0004] Furthermore, existing intelligent diagnostic systems, when collecting and analyzing clinical physiological parameters, fail to fully consider the overall physiological state information closely related to gynecological diseases under the traditional Chinese medicine theoretical system, such as the impact of long-term sleep patterns on endocrine and bodily functions. For elderly gynecological patients, changes in sleep patterns are often an important indicator of their overall health status and specific disease risks. Current technologies lack mechanisms to incorporate such information reflecting the patient's long-term physiological rhythms and overall state into multimodal fusion analysis, potentially overlooking key individualized factors influencing disease occurrence and development, thus limiting the diagnostic system's in-depth understanding and differentiation of elderly gynecological diseases accompanied by complex systemic symptoms.
[0005] This invention aims to address the problems of rigid multimodal data fusion mechanisms and a lack of hierarchical and dynamic feature representation in intelligent ultrasound diagnosis of elderly gynecological conditions. It seeks to establish a technical solution capable of adaptively recombining multi-level features and achieving deep semantic alignment between imaging and clinical data to improve the accuracy and reliability of identifying specific pathological changes in the elderly. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides an intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data, the system comprising: The data integration module acquires ultrasound scan images of elderly gynecological patients and extracts accompanying clinical and physiological parameters, integrating them into an initial dataset. The cross-modal coding module inputs the initial data set into the cross-modal codec. The point multi-cross-modal codec converts the ultrasound image data into feature maps and maps the clinical physiological parameters into semantic embedding vectors. The hierarchical decomposition module performs cross-attention fusion of the feature map and the semantic embedding vector in the feature interaction subspace to generate a multimodal interaction feature tensor, and performs hierarchical structured decomposition to separate global topological layer features, regional organization layer features and local microtexture layer features. The route reorganization module constructs a learnable dynamic feature routing network and adaptively allocates path weights based on the correlation strength between the global topology layer features and the regional organization layer features, for dynamically reorganizing the features of each layer obtained from the hierarchical structured decomposition. The feature refinement module aggregates the features recombined by the dynamic feature routing network with the local microtexture layer features through a gated residual connection, and outputs a refined feature tensor. The diagnostic decision module inputs the refined feature tensor into the multi-branch collaborative diagnostic network, aggregates the outputs of each branch of the multi-branch collaborative diagnostic network, and performs confidence weighting and contextual correlation analysis through the decision fusion layer to generate a structured diagnostic report.
[0008] Preferably, the acquisition of ultrasound scan images of elderly gynecological patients and the extraction of accompanying clinical physiological parameters, integrated into an initial dataset, includes: Multi-sectional scanning of the target area was performed on elderly gynecological patients using ultrasound diagnostic equipment to acquire dynamic ultrasound image sequences containing blood flow Doppler signals. Simultaneously, the patient's clinical physiological parameters corresponding to this ultrasound examination are extracted from the hospital information system. These clinical physiological parameters include hormone level indicators, a summary of past medical history, and a description of the clinical indications for this examination. The dynamic ultrasound image sequence is preprocessed, including temporal frame alignment to eliminate the influence of respiratory motion, multi-plane spatial registration to construct a three-dimensional data volume, and attenuation compensation correction for tissues of elderly patients. The clinical physiological parameters are standardized and encoded, transforming unstructured text descriptions into standardized vector representations; The preprocessed dynamic ultrasound image sequence is associated and bound with the standardized coded clinical physiological parameter vector by timestamp and examination identifier, and encapsulated into an initial data set containing images, parameters and metadata.
[0009] Preferably, the step of inputting the initial data set into a cross-modal codec, whereby the point-multi-modal codec converts the ultrasound image data into a feature map and maps the clinical physiological parameters into semantic embedding vectors, includes: The cross-modal codec includes an image coding branch and a parametric coding branch; The image coding branch uses a three-dimensional convolutional neural network to extract deep features from the dynamic ultrasound image sequence and outputs a four-dimensional feature map. The dimensions of the four-dimensional feature map include spatial height, spatial width, slice depth, and feature channels. The parameter encoding branch uses a recurrent neural network and a self-attention mechanism to perform sequence modeling on the standardized clinical physiological parameter vector, capture the temporal dependence and semantic association between parameters, and finally output a fixed-dimensional semantic embedding vector.
[0010] Preferably, the step of performing cross-attention fusion between the feature map and the semantic embedding vector in the feature interaction subspace to generate a multimodal interaction feature tensor includes: Construct a shared feature interaction subspace, flatten the spatial dimensions of the four-dimensional feature map, and form a series of spatial location feature vectors; The semantic embedding vector is used as the query vector, and the flattened spatial location feature vector is used as the key vector and value vector, and cross-attention calculation is performed. Through the cross-attention calculation, the semantic embedding vector applies different attention weights to different spatial location feature vectors, thereby achieving targeted injection of clinical semantic information into image spatial features; The weighted aggregated spatial location feature vectors are reconstructed into a new feature map with the same spatial size as the original four-dimensional feature map, which serves as the multimodal interactive feature tensor.
[0011] Preferably, the hierarchical structured decomposition, separating global topological layer features, regional organization layer features, and local micro-texture layer features, includes: The multimodal interaction feature tensor is convolved with a convolution kernel with a large receptive field to extract features describing the overall connectivity between the entire lesion region and the surrounding tissues, which are then used as the global topology layer features. The multimodal interactive feature tensor is convolved using a medium-sized convolution kernel to extract features describing the division of different functional regions within the lesion and their interrelationships, which are then used as the regional tissue layer features. The multimodal interactive feature tensor is convolved using small-scale convolution kernels to extract fine texture features describing tissue boundaries, internal micro-calcifications, or cystic regions, which are then used as the local microtexture layer features.
[0012] Preferably, the construction of a learnable dynamic feature routing network, and the adaptive allocation of path weights based on the correlation strength between the global topology layer features and the regional organization layer features, for dynamically reorganizing the features of each layer obtained from the hierarchical structured decomposition, includes: The dynamic feature routing network takes the global topology layer features and regional organization layer features as input; Calculate the feature similarity matrix between the global topology layer features and the regional organization layer features at each spatial location; The feature similarity matrix is input into a lightweight multilayer perceptron, which outputs a dynamic routing weight matrix for each spatial location and feature channel. The dynamic routing weight matrix is used to perform channel weighting and spatial modulation on the regional organization layer features, and the modulated features are added element-wise to the global topology layer features to complete feature recombination.
[0013] Preferably, the feature aggregation process, which involves combining the features reconstructed by the dynamic feature routing network with the local microtexture layer features through a gated residual connection to output a refined feature tensor, includes: Design a gating mechanism that takes the recombined features as input and generates a gating coefficient map ranging from zero to one. Each value in the gating coefficient map represents the importance of the local microtexture layer features at the corresponding spatial location; The gating coefficient map is multiplied element-wise with the local microtexture layer features to achieve soft selection of the local microtexture layer features; The soft-selected local microtexture layer features are residually joined with the recombined features, i.e., added element by element, to output the final refined feature tensor.
[0014] Preferably, inputting the refined feature tensor into the multi-branch collaborative diagnostic network includes: The multi-branch collaborative diagnostic network includes parallel branches for lesion identification, boundary delineation, and pathological grading. The lesion identification branch adopts a fully convolutional network structure to process the refined feature tensor and output a probability map of whether a lesion exists and a preliminary heatmap of lesion categories. The boundary delineation branch adopts an encoder-decoder structure and integrates an attention mechanism in the decoding stage to focus on extracting the boundary features between lesions and normal tissues from the refined feature tensor, and outputting a high-precision lesion contour segmentation map. The pathological grading branch adopts a structure combining global pooling and fully connected layers to perform overall feature summarization and mapping on the refined feature tensor, and outputs a grading probability distribution representing different pathological severity levels.
[0015] Preferably, the outputs of each branch of the multi-branch collaborative diagnostic network are aggregated, and a confidence-weighted and context-related analysis is performed through a decision fusion layer to generate a structured diagnostic report, including: The decision fusion layer receives a lesion category heatmap from the lesion identification branch, a lesion contour segmentation map from the boundary delineation branch, and a graded probability distribution from the pathological grading branch. For each candidate lesion region, a fusion weight is calculated based on its confidence score in the output of different branches, wherein the confidence score is generated by the internal logic of each branch; The aforementioned fusion weights are used to weight and integrate the regional, boundary, and hierarchical information about the lesions from different branches; The integrated information is filled in according to the preset report template, including lesion location, size, morphological characteristics, boundary clarity, internal echo characteristics, and pathological grading suggestions, and a structured diagnostic report containing text descriptions and key image annotations is automatically generated.
[0016] Preferably, the system operation process further includes: Establish a diagnostic feedback loop by comparing the structured diagnostic report generated by the system with the subsequent pathologically confirmed diagnostic results. Based on the comparison differences, the network parameters in the cross-modal codec, the dynamic feature routing network, and the multi-branch collaborative diagnostic network are adjusted using the backpropagation algorithm; Through continuous iterative updates, the intelligent ultrasound diagnostic system for geriatric gynecology, which integrates multimodal data, adapts to the data characteristics and diagnostic standards of different medical institutions, thereby improving the system's generalization ability and diagnostic consistency.
[0017] Compared with the prior art, the beneficial effects of the present invention are: The multimodal interactive feature tensor is structurally separated into three layers—global topology, regional organization, and local microtexture—through a hierarchical decomposition module, clearly defining the pathological information corresponding to different scales. Based on this, the dynamic feature routing network adaptively calculates path weights according to the correlation strength between global and regional features, real-time reorganizing the contribution ratio of each layer of features. This mechanism enables the system to automatically strengthen the discrimination weight of the local microtexture layer based on the heterogeneity of specific cases, while weakening the global structural interference caused by degradation. This overcomes the limitations of fixed weights or simple splicing fusion, giving the model the ability to dynamically adjust feature representation and improving feature discrimination accuracy in complex pathological scenarios.
[0018] The cross-modal coding module establishes deep semantic associations between ultrasound image feature maps and clinical physiological parameter semantic vectors in the feature interaction subspace through a cross-attention mechanism, rather than superficial splicing. The multi-branch collaborative diagnostic network in the decision-making stage employs confidence-weighted and context-dependent analysis to integrate diagnostic logic specific to geriatric gynecology for associative reasoning. This design allows the diagnostic process to incorporate domain prior knowledge, systematically evaluating the supportive or contradictory relationships between multimodal evidence. This results in more consistent and interpretable diagnostic judgments when faced with atypical imaging presentations in elderly patients, particularly improving the reliability of differentiating between multiple etiologies. Attached Figure Description
[0019] Figure 1 This is a timing diagram of the intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data, as described in this invention. Figure 2 A flowchart for integrating the initial dataset; Figure 3 A flowchart for generating multimodal interaction feature tensors; Figure 4 Trend chart of contribution of global / regional / local features; Figure 5 Heat map for comprehensive evaluation of the entire process of the intelligent ultrasound diagnostic system for geriatric gynecology. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Please see Figure 1This invention provides an intelligent ultrasound diagnostic system for elderly gynecological patients that integrates multimodal data. The system includes: acquiring ultrasound scan images of elderly gynecological patients and their accompanying clinical and physiological parameters, and integrating them into an initial dataset. This initial dataset is then processed by a cross-modal coding module; the cross-modal encoder-decoder transforms the ultrasound image data into a feature map rich in semantic information, while mapping the clinical and physiological parameters into compact semantic embedding vectors. A hierarchical decomposition module, within a shared feature interaction subspace, uses a cross-attention mechanism to deeply fuse the feature map and semantic embedding vectors, generating a multimodal interactive feature tensor; and performs hierarchical structured decomposition on this tensor, separating global topological layer features, regional tissue layer features, and local microtexture layer features. A route reorganization module constructs a learnable dynamic feature routing network, which adaptively allocates path weights based on the correlation strength between global topological layer features and regional tissue layer features, thereby dynamically reorganizing the features at each layer. A feature refinement module aggregates the route-reorganized features and local microtexture layer features through gated residual connections, outputting a highly integrated and refined feature tensor. The diagnostic decision module inputs the refined feature tensor into a multi-branch collaborative diagnostic network, which performs lesion identification, boundary delineation and pathological grading tasks in parallel. The decision fusion layer performs confidence weighting and contextual analysis on the output of each branch, and automatically generates a structured diagnostic report accordingly.
[0022] In one embodiment of the present invention, see [reference] Figure 2 A complete examination begins with a multi-slice scan of the target pelvic region of an elderly female patient. During the scan, the ultrasound diagnostic equipment simultaneously acquires dynamic ultrasound image sequences containing blood flow Doppler signals, which are stored digitally. At the same time, the system automatically extracts the patient's clinical physiological parameters associated with a unique identifier for this examination from the hospital information system. These extracted parameters explicitly include serum estradiol levels, a summary of the archived medical history of uterine fibroids, and a description of the clinical indications for postmenopausal bleeding recorded on the examination request form.
[0023] Optionally, the clinical physiological parameters may also include recent sleep quality assessment data collected by standardized scales, such as quantitative indicators like sleep duration, difficulty falling asleep, and number of nighttime awakenings. These indicators reflect the patient's autonomic nervous system function and long-term health background, and can serve as auxiliary reference information for judging the balance of Qi, blood, Yin, and Yang in the body within the framework of traditional Chinese medicine's syndrome differentiation and treatment system. During the data integration phase, this sleep quality assessment data will be standardized and encoded along with other clinical physiological parameters, transforming it into a machine-readable vector representation for subsequent cross-modal fusion analysis with ultrasound imaging features.
[0024] Preprocessing was performed on the acquired dynamic ultrasound image sequences. This preprocessing included temporal frame alignment using a mutual information-based registration algorithm to eliminate image shifts caused by respiratory motion; spatial registration of multi-section images acquired from sagittal, transverse, and coronal planes to construct a three-dimensional data volume in a unified coordinate system; and applying a deep learning-based attenuation model to the three-dimensional data volume to compensate for attenuation changes in the acoustic properties of tissues in elderly patients. Extracted clinical physiological parameters were standardized and encoded. Natural language processing techniques were used to convert the summary text of past medical history and the description text of clinical indications into word vector sequences, which, along with numerical hormone level indicators, were input into a standardized mapping layer, outputting a fixed-dimensional standardized clinical physiological parameter vector. The preprocessed dynamic ultrasound image sequences and the standardized clinical physiological parameter vectors were associated and bound based on their shared examination timestamps and examination identifiers, encapsulating them into an initial data set containing image data blocks, parameter vectors, and metadata describing the data source and version.
[0025] The cross-modal coding module receives the encapsulated initial data set. The module's built-in cross-modal codec initiates the image coding branch and the parametric coding branch. The image coding branch employs a 3D convolutional neural network with multi-level 3D convolution and 3D pooling operations to extract deep features from the dynamic ultrasound image sequence. The 3D convolutional neural network ultimately outputs a four-dimensional feature map, with its four dimensions corresponding to spatial height, spatial width, slice depth, and feature channels, respectively. The parametric coding branch first uses a bidirectional gated recurrent unit network to perform sequence modeling on standardized clinical physiological parameter vectors, capturing the temporal dependencies between parameters. Then, the output of the recurrent network is fed into a self-attention mechanism layer for semantic association enhancement. The self-attention mechanism layer calculates the association weights between the query vector, key vector, and value vector. The calculation process of the association weights can be expressed by the following formula: Where: symbol This represents the normalized attention weight of the m-th query vector to the n-th key vector, denoted by the symbol. Represents the m-th query vector, symbol Represents the nth key vector, symbol and These represent the linear transformation functions applied to the query vector and key vector, respectively, with symbols... This represents the total length of the input sequence. The parameter encoding branch ultimately outputs a fixed-dimensional semantic embedding vector.
[0026] In one embodiment of the present invention, see [reference] Figure 3The four-dimensional feature map and semantic embedding vector output by the cross-modal coding module are fed into the hierarchical decomposition module. The hierarchical decomposition module constructs a shared feature interaction subspace. Within this subspace, the spatial height, width, and slice depth dimensions of the four-dimensional feature map are flattened, arranging the feature vectors at each spatial location into a sequence. The semantic embedding vector output by the parametric coding branch is used as the query vector, and the flattened sequence of spatial location feature vectors is used as the key and value vectors, respectively, for cross-attention calculation. Cross-attention calculation allows the semantic embedding vector to apply different attention weights to feature vectors at different spatial locations, thereby achieving targeted injection and fusion of the semantic information contained in clinical physiological parameters into the spatial features of ultrasound images. In some embodiments, the calculation process of the attention weights is expressed by the following formula: Where: symbol The symbol represents the normalized attention weight of the feature vector at the q-th spatial location on the p-th semantic dimension when the semantic embedding vector is used as a query. The p-th dimension component of the semantic embedding vector is represented by the symbol. Represents the feature vector of the q-th spatial location, with the symbol... and These represent the linear projection functions applied to the query component and the key vector, respectively, with symbols... This represents the total number of spatial location feature vectors after flattening. The weighted and aggregated sequence of spatial location feature vectors is reconstructed into a new feature map with the same spatial height, spatial width, and slice depth as the original four-dimensional feature map. The new feature map is defined as a multimodal interaction feature tensor.
[0027] A hierarchical structured decomposition operation is performed on the generated multimodal interaction feature tensor. A convolutional kernel with a large receptive field is used to convolve the multimodal interaction feature tensor. This large receptive field convolutional kernel is designed to cover the entire target anatomical region. The features extracted by its convolution operation describe the overall connectivity, relative position, and layout of the entire lesion region and surrounding tissues. These features are separated and defined as global topological layer features. A medium-sized convolutional kernel is then used to convolve the multimodal interaction feature tensor. In some embodiments, the size of the medium-sized convolutional kernel corresponds to the size of a typical lesion substructure. The features extracted by its convolution operation describe the division of different functional or density regions within the lesion and the relationships between these regions. These features are separated and defined as regional tissue layer features. Finally, a small-scale convolutional kernel is used to convolve the multimodal interaction feature tensor. Optionally, the size of the small-scale convolution kernel is designed to capture subtle texture changes. The features extracted by its convolution operation describe the high-frequency information of tissue boundaries, the fine texture patterns of small calcifications or small cystic regions present inside, and these features are separated and defined as local microtexture layer features.
[0028] In one embodiment of the invention, global topology layer features, regional organization layer features, and local microtexture layer features are passed to a route reorganization module. The route reorganization module constructs a learnable dynamic feature routing network. This network takes the global topology layer features and regional organization layer features as input and calculates a feature similarity matrix between the global topology layer features and the regional organization layer features at each spatial location. In some embodiments, calculating the feature similarity matrix involves measuring the cosine similarity of the feature vectors of the two feature layers at corresponding spatial locations. The calculated feature similarity matrix is input into a lightweight multilayer perceptron, which consists of two fully connected layers and a nonlinear activation function. The lightweight multilayer perceptron outputs a dynamic routing weight matrix for each spatial location and each feature channel. It can be understood that the values of the dynamic routing weight matrix are adaptively learned by the network based on the correlation strength between the input features. The dynamic routing weight matrix is used to perform channel weighting and spatial modulation operations on the regional organization layer features. The channel weighting and spatial modulation operations are manifested by multiplying the dynamic routing weight matrix with the regional organization layer features element by element, and then adding the modulated regional organization layer features with the global topology layer features element by element, thereby completing the dynamic recombination of features.
[0029] The feature refinement module receives features reconstructed by the dynamic feature routing network and local micro-texture layer features. The module employs a gating mechanism. This mechanism takes the reconstructed features as input, processes them through a convolutional layer and a sigmoid activation function, and generates a gating coefficient map with values ranging from zero to one. The value at each coordinate point in the gating coefficient map represents the importance of the local micro-texture layer features at that spatial location. The generation process of the values in the gating coefficient map can be represented by the following formula: Where: symbol Indicates spatial coordinates ( The gating coefficient value at () is represented by the symbol. Represents the Sigmoid activation function, symbol The convolutional layer representing the gating mechanism corresponds to the first... The convolution kernel weights for each input channel, with symbols... The features represented by the reassembled features of the dynamic feature routing network are in spatial coordinates ( ) The characteristic values of each channel, symbol The bias parameter representing the gated convolutional layer, symbol This represents a convolutional summation operation. The gated coefficient map is multiplied element-wise with the local micro-texture layer features, achieving soft selection of these features. The soft-selected local micro-texture layer features are then joined with the features reconstructed by the dynamic feature routing network using a residual connection, i.e., element-wise addition. Optionally, the residual connection can include an additional 1x1 convolutional layer to adjust the channel dimensions, ultimately outputting a refined feature tensor with aggregated information.
[0030] In one embodiment of the present invention, the refined feature tensor is input to a multi-branch collaborative diagnostic network. The multi-branch collaborative diagnostic network includes a lesion identification branch, a boundary delineation branch, and a pathological grading branch that operate in parallel. The lesion identification branch processes the refined feature tensor using a fully convolutional network structure. The fully convolutional network structure consists of multiple consecutive convolutional layers and upsampling layers, but does not contain fully connected layers. The lesion identification branch outputs a two-dimensional probability map with the same spatial size as the input tensor. Each pixel value in the two-dimensional probability map represents the probability of a lesion existing at that spatial location. Simultaneously, the lesion identification branch outputs a preliminary lesion category heatmap, where different colored areas indicate the probability of the distribution of different types of lesions, such as suspected endometrial cancer, ovarian cysts, or uterine fibroids.
[0031] The boundary delineation branch employs an encoder-decoder structure to process the refined feature tensor. The encoder extracts deep features through downsampling, while the decoder restores spatial resolution through upsampling. An attention mechanism is integrated into the decoder stage, enabling the network to focus on the boundary features between lesions and normal tissue. The boundary delineation branch ultimately outputs a high-precision lesion contour segmentation map, where each pixel is classified as either a lesion region or background region. The pathological grading branch uses a structure combining global pooling and fully connected layers. The global pooling layer aggregates the refined feature tensor spatially to obtain a global feature vector. This global feature vector is then fed into a series of fully connected layers for nonlinear mapping. The pathological grading branch outputs a grading probability distribution representing different pathological severity levels. For example, for endometrial lesions, the grading probability distribution can correspond to multiple levels of probability values, such as normal, benign hyperplasia, dysplasia, and cancer.
[0032] The decision fusion layer receives a lesion category heatmap from the lesion identification branch, a lesion contour segmentation map from the boundary delineation branch, and a grading probability distribution from the pathology grading branch. For each candidate lesion region identified by the boundary delineation branch, the decision fusion layer calculates a fusion weight based on its confidence score in the output of each branch. The confidence score of the lesion identification branch is derived from the maximum activation value of the corresponding category in the lesion category heatmap; the confidence score of the boundary delineation branch is derived from the boundary sharpness score of the segmented contour; and the confidence score of the pathology grading branch is derived from the entropy of the maximum probability value in the grading probability distribution. In essence, the confidence score of each branch reflects the degree of certainty with respect to the current region. The fusion weights are used to weight and integrate the information from different branches. This weighted integration process can be represented by the following formula: Where: symbol Indicates the first The fused feature vector of each candidate lesion region, symbol The symbol represents the total number of branches (3 in this case). Representing the The region in the first The fusion weights on each branch, with symbols Representing the The branch is the first Each region provides a feature information vector, symbol This represents scalar multiplication.
[0033] The decision fusion layer populates the integrated information according to a preset structured report template, automatically generating a structured diagnostic report containing text descriptions and key image annotations. The report template requires fields to be filled in, including lesion location, lesion size, lesion morphological characteristics, boundary clarity, internal echo characteristics, and pathological grading recommendations. In some embodiments, when processing a specific polycystic ovary syndrome (PCOS) case, the decision fusion layer refers to Table 1 to extract key quantitative and descriptive information from each branch for report generation.
[0034] Table 1: Output Information of Example Lesions from Multi-Branch Collaborative Diagnostic Network Optionally, when generating the report, the decision fusion layer will perform correlation verification between the lesion size calculated by the boundary delineation branch, the category determined by the lesion identification branch, and the suggested level of the pathological grading branch to ensure the logical consistency of the report content. Finally, the structured diagnostic report is presented in a combination of text and images. The text part systematically describes the above fields, while the image part automatically marks the lesion outline and adds category and grading labels.
[0035] See Figure 4 In the hierarchical decomposition stage of the intelligent diagnostic system for geriatric gynecological ultrasound integrating multimodal data, the contributions of global topological layer features, regional tissue layer features, and local microtexture layer features exhibit differentiated dynamic trends. Specifically, as the system operation extends from the data integration stage to the diagnostic decision-making stage, the contribution of each layer's features gradually changes with the stage's progression: local microtexture layer features (orange) significantly lead in contribution after the hierarchical decomposition stage, maintaining a high level; regional tissue layer features (purple) maintain a moderate contribution after the hierarchical decomposition stage; and global topological layer features (blue) have a relatively low and stable contribution. This trend reflects the system's increased reliance on local fine texture information after the hierarchical decomposition stage, while simultaneously considering the feature support of regional tissue and global topology, consistent with the feature synergy logic of "local details - regional structure - global correlation" in multimodal ultrasound diagnosis.
[0036] In one embodiment of the present invention, a geriatric gynecological ultrasound intelligent diagnostic system integrating multimodal data establishes a diagnostic feedback closed loop. The diagnostic feedback closed loop compares the structured diagnostic report generated by the system with the final diagnostic result confirmed by subsequent pathological examination. This comparison is performed at the case level, involving checking the differences between the lesion presence, lesion type, and pathological grading recommendations identified in the report and the final diagnostic result confirmed by pathological examination. In a specific implementation, a structured diagnostic report might classify a lesion in a certain uterine region as "high risk of endometrial cancer (G2)," while the subsequent pathological examination confirms a final diagnostic result of "complex endometrial hyperplasia with localized atypical hyperplasia." The system records the specific information regarding the deviations in lesion type judgment and pathological grading.
[0037] Based on the differences identified during the comparison, the network parameters in the cross-modal encoder / decoder, dynamic feature routing network, and multi-branch collaborative diagnostic network are adjusted using the backpropagation algorithm. The backpropagation algorithm calculates the gradient of the loss function between the structured diagnostic report output and the final diagnostic result confirmed by pathological examination. This gradient propagates back along the paths of the multi-branch collaborative diagnostic network, feature refinement module, routing reconstruction module, hierarchical decomposition module, and cross-modal coding module. The update process of the network parameters can be expressed by the following formula: Where: symbol Indicates the iteration step The network parameter set at time, symbol Indicates the iteration step The network parameter set at time, symbol The learning rate hyperparameter is represented by the symbol. Represents the loss function Relative to network parameters gradient, sign A smart ultrasound diagnostic system for geriatric gynecology, representing the fusion of multimodal data, uses input data... and current parameters The calculated output, symbol This represents the final diagnostic result confirmed by pathological examination. It can be understood that the backpropagation algorithm updates the network parameters in a direction that reduces the discrepancy between the diagnostic output and the final diagnostic result confirmed by pathological examination.
[0038] Through a continuous iterative update process, the geriatric gynecological ultrasound intelligent diagnostic system, which integrates multimodal data, can adapt to the data characteristics and diagnostic standards of different medical institutions. In some embodiments, when the system is deployed in a new medical institution, the imaging characteristics of the ultrasound equipment, the recording format of clinical physiological parameters, and the pathological diagnostic standards of that institution may differ from the original training data. The diagnostic feedback loop continuously collects newly generated, pathologically confirmed final diagnostic results locally as supervisory signals and drives fine-tuning of network parameters. In specific implementations, after the system undergoes iterative updates based on hundreds of local cases, the parameter encoding branch and image encoding branch in the cross-modal encoder / decoder will adaptively adjust the feature extraction strategy for the unique ultrasound image noise patterns and slightly different hormone level reporting ranges of that medical institution. The dynamic feature routing network will learn to adapt to the feature association patterns under the new data, and the decision threshold in the multi-branch collaborative diagnostic network will also evolve accordingly, thereby improving the system's generalization ability and diagnostic consistency under the specific institutional data characteristics. Optionally, the iterative update process can be set to be triggered after a safe number of accumulated cases or executed periodically based on a sliding time window to ensure the stability of system optimization.
[0039] See Figure 5 The heatmap uses different modules (data integration, cross-modal coding, hierarchical decomposition, etc.) as the row dimension and core performance indicators (data processing efficiency, feature extraction accuracy, etc.) as the column dimension, visually presenting the overall performance of each module under different indicators through color gradients (corresponding to a 10-point scoring system). Specifically, the data integration module scored 9.2 in the "data processing efficiency" indicator (red, representing high performance), but only 5.5 in the "institutional adaptability" indicator (dark blue, relatively weak performance); the feedback loop module achieved a high score of 9.5 in the "diagnostic consistency" indicator, reflecting its outstanding performance in the stability of diagnostic results; the feature refinement and multi-branch diagnosis modules achieved scores of 9.2 and 9.0 respectively in the "feature extraction accuracy" indicator, reflecting the accuracy of the system in the multimodal feature fusion stage; the scores of each module in indicators such as "error control level" and "institutional adaptability" are generally distributed in the low to medium range, indicating that there is still room for optimization in the error control and cross-institutional adaptability dimensions of the system. This heatmap clearly depicts the strengths and weaknesses of each process module in the system across different performance dimensions, providing data support for future optimization directions.
[0040] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0041] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A smart ultrasound diagnostic system for geriatric gynecology that integrates multimodal data, characterized in that, include: The data integration module acquires ultrasound scan images of elderly gynecological patients and extracts accompanying clinical and physiological parameters, integrating them into an initial dataset. The cross-modal coding module inputs the initial data set into the cross-modal codec. The point multi-cross-modal codec converts the ultrasound image data into feature maps and maps the clinical physiological parameters into semantic embedding vectors. The hierarchical decomposition module performs cross-attention fusion of the feature map and the semantic embedding vector in the feature interaction subspace to generate a multimodal interaction feature tensor, and performs hierarchical structured decomposition to separate global topological layer features, regional organization layer features and local microtexture layer features. The route reorganization module constructs a learnable dynamic feature routing network and adaptively allocates path weights based on the correlation strength between the global topology layer features and the regional organization layer features, for dynamically reorganizing the features of each layer obtained from the hierarchical structured decomposition. The feature refinement module aggregates the features recombined by the dynamic feature routing network with the local microtexture layer features through a gated residual connection, and outputs a refined feature tensor. The diagnostic decision module inputs the refined feature tensor into the multi-branch collaborative diagnostic network, aggregates the outputs of each branch of the multi-branch collaborative diagnostic network, and performs confidence weighting and contextual correlation analysis through the decision fusion layer to generate a structured diagnostic report.
2. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The process of acquiring ultrasound scan images of elderly gynecological patients and extracting accompanying clinical and physiological parameters, integrating them into an initial dataset, includes: Multi-sectional scanning of the target area was performed on elderly gynecological patients using ultrasound diagnostic equipment to acquire dynamic ultrasound image sequences containing blood flow Doppler signals. Simultaneously, the patient's clinical physiological parameters corresponding to this ultrasound examination are extracted from the hospital information system. These clinical physiological parameters include hormone level indicators, a summary of past medical history, and a description of the clinical indications for this examination. The dynamic ultrasound image sequence is preprocessed, including temporal frame alignment to eliminate the influence of respiratory motion, multi-plane spatial registration to construct a three-dimensional data volume, and attenuation compensation correction for tissues of elderly patients. The clinical physiological parameters are standardized and encoded, transforming unstructured text descriptions into standardized vector representations; The preprocessed dynamic ultrasound image sequence is associated and bound with the standardized coded clinical physiological parameter vector by timestamp and examination identifier, and encapsulated into an initial data set containing images, parameters and metadata.
3. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The initial data set is input into a cross-modal codec. The point-multi-modal codec converts the ultrasound image data into a feature map and maps the clinical physiological parameters into semantic embedding vectors, including: The cross-modal codec includes an image coding branch and a parametric coding branch; The image coding branch uses a three-dimensional convolutional neural network to extract deep features from the dynamic ultrasound image sequence and outputs a four-dimensional feature map. The dimensions of the four-dimensional feature map include spatial height, spatial width, slice depth, and feature channels. The parameter encoding branch uses a recurrent neural network and a self-attention mechanism to perform sequence modeling on the standardized clinical physiological parameter vector, capture the temporal dependence and semantic association between parameters, and finally output a fixed-dimensional semantic embedding vector.
4. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 3, characterized in that, In the feature interaction subspace, the feature map and the semantic embedding vector are cross-attention fused to generate a multimodal interaction feature tensor, including: Construct a shared feature interaction subspace, flatten the spatial dimensions of the four-dimensional feature map, and form a series of spatial location feature vectors; The semantic embedding vector is used as the query vector, and the flattened spatial location feature vector is used as the key vector and value vector, and cross-attention calculation is performed. Through the cross-attention calculation, the semantic embedding vector applies different attention weights to different spatial location feature vectors, thereby achieving targeted injection of clinical semantic information into image spatial features; The weighted aggregated spatial location feature vectors are reconstructed into a new feature map with the same spatial size as the original four-dimensional feature map, which serves as the multimodal interactive feature tensor.
5. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The execution hierarchical structured decomposition separates global topological layer features, regional organization layer features, and local micro-texture layer features, including: The multimodal interaction feature tensor is convolved with a convolution kernel with a large receptive field to extract features describing the overall connectivity between the entire lesion region and the surrounding tissues, which are then used as the global topology layer features. The multimodal interactive feature tensor is convolved using a medium-sized convolution kernel to extract features describing the division of different functional regions within the lesion and their interrelationships, which are then used as the regional tissue layer features. The multimodal interactive feature tensor is convolved using small-scale convolution kernels to extract fine texture features describing tissue boundaries, internal micro-calcifications, or cystic regions, which are then used as the local microtexture layer features.
6. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The construction of a learnable dynamic feature routing network, and the adaptive allocation of path weights based on the correlation strength between the global topology layer features and the regional organization layer features, for dynamically reorganizing the features of each layer obtained from the hierarchical structured decomposition, includes: The dynamic feature routing network takes the global topology layer features and regional organization layer features as input; Calculate the feature similarity matrix between the global topology layer features and the regional organization layer features at each spatial location; The feature similarity matrix is input into a lightweight multilayer perceptron, which outputs a dynamic routing weight matrix for each spatial location and feature channel. The dynamic routing weight matrix is used to perform channel weighting and spatial modulation on the regional organization layer features, and the modulated features are added element-wise to the global topology layer features to complete feature recombination.
7. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 6, characterized in that, The process of combining the features reconstructed by the dynamic feature routing network with the local micro-texture layer features through a gated residual connection to aggregate features and output a refined feature tensor includes: Design a gating mechanism that takes the recombined features as input and generates a gating coefficient map ranging from zero to one. Each value in the gating coefficient map represents the importance of the local microtexture layer features at the corresponding spatial location; The gating coefficient map is multiplied element-wise with the local microtexture layer features to achieve soft selection of the local microtexture layer features; The soft-selected local microtexture layer features are residually joined with the recombined features, i.e., added element by element, to output the final refined feature tensor.
8. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The step of inputting the refined feature tensor into the multi-branch collaborative diagnostic network includes: The multi-branch collaborative diagnostic network includes parallel branches for lesion identification, boundary delineation, and pathological grading. The lesion identification branch adopts a fully convolutional network structure to process the refined feature tensor and output a probability map of whether a lesion exists and a preliminary heatmap of lesion categories. The boundary delineation branch adopts an encoder-decoder structure and integrates an attention mechanism in the decoding stage to focus on extracting the boundary features between lesions and normal tissues from the refined feature tensor, and outputting a high-precision lesion contour segmentation map. The pathological grading branch adopts a structure combining global pooling and fully connected layers to perform overall feature summarization and mapping on the refined feature tensor, and outputs a grading probability distribution representing different pathological severity levels.
9. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 8, characterized in that, The outputs of each branch of the multi-branch collaborative diagnostic network are aggregated and subjected to confidence weighting and contextual correlation analysis through a decision fusion layer to generate a structured diagnostic report, including: The decision fusion layer receives a lesion category heatmap from the lesion identification branch, a lesion contour segmentation map from the boundary delineation branch, and a graded probability distribution from the pathological grading branch. For each candidate lesion region, a fusion weight is calculated based on its confidence score in the output of different branches, wherein the confidence score is generated by the internal logic of each branch; The aforementioned fusion weights are used to weight and integrate the regional, boundary, and hierarchical information about the lesions from different branches; The integrated information is filled in according to the preset report template, including lesion location, size, morphological characteristics, boundary clarity, internal echo characteristics, and pathological grading suggestions, and a structured diagnostic report containing text descriptions and key image annotations is automatically generated.
10. The intelligent ultrasound diagnostic system for geriatric gynecology that integrates multimodal data according to claim 1, characterized in that, The system's operation process also includes: Establish a diagnostic feedback loop by comparing the structured diagnostic report generated by the system with the subsequent pathologically confirmed diagnostic results. Based on the comparison differences, the network parameters in the cross-modal codec, the dynamic feature routing network, and the multi-branch collaborative diagnostic network are adjusted using the backpropagation algorithm; Through continuous iterative updates, the intelligent ultrasound diagnostic system for geriatric gynecology, which integrates multimodal data, adapts to the data characteristics and diagnostic standards of different medical institutions, thereby improving the system's generalization ability and diagnostic consistency.