A multi-modal medical image data processing method and device
By employing multimodal medical image data processing methods that combine medical images and text data, multi-scale feature extraction and cross-modal fusion are performed, solving the problem that single-dimensional analysis is insufficient to capture health status and achieving accurate analysis and improved reliability of lesion features.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2025-12-29
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, single-dimensional analysis of medical image data is insufficient to fully capture a patient's health status, resulting in low reliability of lesion feature analysis results and difficulty in adapting to lesion features of different sizes and shapes.
A multimodal medical image data processing method is adopted, which combines medical image data and medical text data through a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network to perform multi-scale basic visual feature extraction, spatial structure capture, semantic feature extraction, and prior weighted fusion, thereby obtaining the matching probability between multimodal data and target lesions.
It improves the accuracy and reliability of lesion analysis results, can comprehensively capture the characteristics of lesions of different sizes and shapes, and enhances the adaptability of the model.
Smart Images

Figure CN121789004B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing technology, and in particular to a multimodal medical image data processing method and apparatus. Background Technology
[0002] Currently, intelligent processing of medical data has become an important means to improve diagnostic and treatment efficiency and assist clinical decision-making. Existing technologies analyze medical image data (such as CT, MRI, and X-rays) to assess a patient's health status. However, analyzing only a single dimension of medical image data often fails to comprehensively capture a patient's health condition. Furthermore, existing image analysis methods are ill-suited to the diverse lesion characteristics of medical images, resulting in low reliability of the analysis results.
[0003] Therefore, a new technical solution is urgently needed to solve the above-mentioned technical problems. Summary of the Invention
[0004] This invention provides a multimodal medical image data processing method and apparatus, which can improve the reliability of lesion analysis results.
[0005] In a first aspect, the present invention provides a multimodal medical image data processing method, comprising: Multimodal data corresponding to the target lesion are input into a trained multimodal medical data processing model. The multimodal data includes medical image data and medical text data. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The medical image data is subjected to multi-scale basic visual feature extraction through the basic visual feature extraction network to obtain a multi-scale basic visual feature set. The structured image feature extraction network captures the spatial structure of the multi-scale basic visual feature set to obtain structured image features including lesion shape, boundary and relative position. The text feature extraction network is used to extract semantic features from the medical text data to obtain text semantic features. The cross-modal feature fusion network is used to perform prior weighted fusion of the structured image features and the text semantic features to obtain the matching probability between the multimodal data and the target lesion.
[0006] Secondly, the present invention provides a multimodal medical image data processing apparatus, comprising: The multimodal data processing module inputs the multimodal data corresponding to the target lesion into the trained multimodal medical data processing model. The multimodal data includes medical image data and medical text data. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The basic feature extraction module is connected to the multimodal data processing module. It extracts multi-scale basic visual features from the medical image data through the basic visual feature extraction network to obtain a multi-scale basic visual feature set. The structured feature extraction module, connected to the basic feature extraction module, captures the spatial structure of the multi-scale basic visual feature set through the structured image feature extraction network to obtain structured image features including lesion shape, boundary and relative position; The semantic feature extraction module is connected to the multimodal data processing module. Through the text feature extraction network, it extracts semantic features from the medical text data to obtain text semantic features. The feature weighted fusion module, connected to the structured feature extraction module and the semantic feature extraction module, performs prior weighted fusion of the structured image features and the text semantic features through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion.
[0007] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method described in the first aspect of the present invention.
[0008] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in the first aspect of the present invention.
[0009] This invention provides a multimodal medical image data processing method and apparatus. It can improve the accuracy of lesion analysis results by combining semantic understanding and rich image analysis results through multimodal data analysis and feature fusion of medical image data and medical text data. Multi-scale parallel convolutional kernels are constructed to comprehensively capture lesion features of different sizes in medical images, improving the adaptability of the model. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Picture 1 This is a flowchart of a multimodal medical image data processing method provided in an embodiment of the present invention; Picture 2 This is a schematic diagram of a multimodal medical data processing model provided in an embodiment of the present invention; Picture 3 This is a schematic diagram of the structure of a multimodal medical image data processing device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0013] Please refer to Picture 1 This invention provides a multimodal medical image data processing method, which includes: Step 100: Input the multimodal data of the corresponding target lesion into the trained multimodal medical data processing model; Multimodal data includes medical image data and medical text data. Multimodal medical data processing models include basic visual feature extraction networks, structured image feature extraction networks, text feature extraction networks, and cross-modal feature fusion networks. Step 102: Extract multi-scale basic visual features from medical image data using a basic visual feature extraction network to obtain a multi-scale basic visual feature set; Step 104: Capture the spatial structure of the multi-scale basic visual feature set through a structured image feature extraction network to obtain structured image features including lesion shape, boundary and relative position; Step 106: Extract semantic features from medical text data using a text feature extraction network to obtain text semantic features; Step 108: Through a cross-modal feature fusion network, the structured image features and text semantic features are fused with prior weights to obtain the matching probability between multimodal data and target lesions.
[0014] In this embodiment of the invention, medical imaging data includes CT images, MRI images, ultrasound images, etc., and medical text data is typically the patient's electronic medical record. Medical imaging data and medical text data describing the same lesion are input into a trained multimodal medical data processing model. The matching probability between the multimodal data and the target lesion is determined based on the output of the multimodal medical data processing model. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The basic visual feature extraction network is preferably Inception-ResNet-V2, used to extract multi-scale basic visual features from the medical imaging data, obtaining a multi-scale basic visual feature set. The structured image feature extraction network is equipped with multi-scale parallel convolutional kernels, used to process the multi-scale basic visual feature set separately, capturing the geometric structure and spatial relationships (spatial structure information) in the image, obtaining structured image features including lesion shape, boundary, and relative position. The medical text features are input into the text feature extraction network (preferably a BERT network) to extract semantic features from the medical text data, obtaining text semantic features. The preferred cross-modal feature fusion network is a cross-attention mechanism fusion network, which integrates image and text features to form a unified feature representation. Then, the model is trained and optimized, and finally the matching probability between multimodal data and target lesions is output.
[0015] In one embodiment of the present invention, the structured image feature extraction network includes: a first convolutional layer, a normalization layer, a second convolutional layer, a max pooling layer, a dynamic weight aggregation layer, and a fully connected layer. The first convolutional layer consists of four layers, each with 128 channels, and the layers are connected in series. The second convolutional layer comprises multiple parallel convolutional layers, each with a different scale of convolutional kernels. The structured image feature extraction network captures the spatial structure of a multi-scale set of basic visual features, obtaining structured image features including lesion shape, boundaries, and relative positions, including: By sequentially performing convolution operations on the multi-scale basic visual feature set through four cascaded first convolutional layers, local texture features and edge features are extracted to obtain convolutional feature maps. The convolutional feature map is normalized by a normalization layer to obtain a normalized feature map. By using each second convolutional layer that includes convolutional kernels of different scales, the local spatial features of lesions of different sizes in the normalized feature map are captured respectively, and the multi-scale local feature tensor output by each second convolutional layer is obtained. The size of multiple multi-scale local feature tensors is unified by using a max pooling layer, and the size-unified multi-scale local feature tensors are then spliced together to obtain a multi-scale local feature set that integrates the local spatial features of lesions of all sizes. The connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer are calculated based on a preset dynamic weight optimization algorithm through a dynamic weight aggregation layer. The multi-scale local feature set is then weighted and summed based on the connection weights to obtain an aggregated structured feature tensor, which includes two dimensions: direction and length. By performing dimensionality mapping on the aggregated structured feature tensor through a fully connected layer, structured image features including lesion shape, boundary, and relative position are obtained.
[0016] In this embodiment, the number of layers in the first convolutional layer is increased from two to four, improving the network's filtering and feature extraction capabilities by increasing network depth. Each convolutional layer has 128 channels and uses a 3×3 scale convolutional filter to reduce the number of parameters in the convolutional layer, allowing the network to achieve more nonlinearity. A batch normalization layer is added after the four first convolutional layers to normalize the convolutional feature maps obtained from the first convolutional layers, resulting in a normalized feature map, which improves convergence speed. To address the issue of large differences in lesion size in medical images, the second convolutional layer consists of multiple parallel convolutional layers, each with convolutional kernels of different scales, arranged in parallel. Each second convolutional layer simultaneously captures local spatial features from the normalized feature map, obtaining multi-scale local feature tensors of local spatial features for lesions of different sizes based on the convolutional kernels within each second convolutional layer. The max-pooling layer unifies the size of multiple multi-scale local feature tensors and then concatenates these unified tensors to obtain a multi-scale local feature set that integrates the local spatial features of lesions of all sizes. The dynamic weight aggregation layer calculates the connection weights between the multi-scale local feature tensors and the aggregated structured feature tensor using a pre-defined dynamic weight optimization algorithm. Based on these connection weights, the multi-scale local feature set is weighted and summed to obtain the aggregated structured feature tensor. Three fully connected layers (first, second, and third) sequentially perform dimension mapping on the aggregated structured feature tensor to obtain structured image features including lesion shape, boundaries, and relative positions. This aggregated structured feature tensor includes two dimensions: direction and length, used for subsequent similarity calculations between the length scalar and the direction vector.
[0017] In a preferred embodiment, to address the issue that lung nodules in lung images may range in size from 3mm to 30mm, and that single-scale convolutional kernels are insufficient to simultaneously capture both small and large lesions, the number of second convolutional layers is set to three. The convolutional kernel sizes within each second convolutional layer are 1x1, 3x3, and 5x5, respectively. The convolutional layer corresponding to the smaller kernel captures micronodules, while the convolutional layer corresponding to the larger kernel captures large tumors. The normalized feature maps calculated in batches are simultaneously input into the parallel-configured 1×1, 3×3, and 5×5 second convolutional layers. Each second convolutional layer spontaneously forms convolutional kernels sensitive to different spatial frequencies based on the differences in gradient magnitudes during backpropagation (the optimizer adaptively updates parameters based on the gradient norm, causing the kernel weights of the 1×1 branch to converge primarily along the high-frequency direction. Similarly, for large masses with d>20mm, the gradient is maximum in the 5×5 branch, causing the parameters of this branch to converge along the low-frequency direction), thus automatically completing the feature allocation for small, medium, and large lesions. Output three tensors S1, S2, and S3. Then max pool S1 and S2 to the same size as S3. Finally, concatenate the three tensors along the cap_type dimension and feed them into a unified dynamic weight optimization algorithm.
[0018] In one embodiment of the present invention, a dynamic weight aggregation layer calculates the connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer based on a preset dynamic weight optimization algorithm. Based on these connection weights, a weighted summation of the multi-scale local feature set is performed to obtain an aggregated structured feature tensor, including: Determine the initial correlation between each branch in the multi-scale feature extraction branch layer and the dynamic weight aggregation layer; The initial correlation degree is converted into dynamic weights with values between 0 and 1 by normalization calculation, where the sum of the dynamic weights of all multi-scale local feature tensors for the same dynamic weight aggregation layer is 1. The dynamic weights are iteratively calculated based on a preset dynamic weight optimization algorithm, and the final dynamic weights are obtained after a preset number of iterations. Based on the final dynamic weights and the preset prediction matrix, all multi-scale local feature tensors are weighted and summed to generate the input of the aggregation layer; pass Squash The function activates the input to the aggregation layer, resulting in an activated aggregated structured feature tensor.
[0019] In this embodiment, the initial correlation degree between each branch in the multi-scale feature extraction layer and the dynamic weight aggregation layer is set to 0 to eliminate initial weight bias and ensure that subsequent iterative optimization of dynamic weights is dynamically generated based on feature matching degree. The initial correlation degree is converted into dynamic weights through normalization calculation, and the sum of the dynamic weights of all multi-scale local feature tensors for the same dynamic weight aggregation layer is limited to 1. The dynamic weights are iteratively calculated based on a preset dynamic weight optimization algorithm, and the final dynamic weights are obtained after a preset number of iterations. Specifically, the matching similarity (e.g., cosine similarity) between each multi-scale local feature tensor and the aggregated feature is calculated. The higher the similarity, the better the feature tensor can characterize the core structural information of the lesion, such as shape and boundary. Based on the similarity results, the dynamic weights of the initial or previous iteration are updated; that is, the weights of feature tensors with high similarity are increased, and the weights of those with low similarity are decreased. The updated weights are normalized to ensure again that the sum of the dynamic weights of all multi-scale local feature tensors for the same aggregated feature is 1, ensuring the rationality of weight allocation. The updated weights are used to perform a weighted summation of the multi-scale local feature tensors, yielding temporary aggregated features for this iteration, providing a reference for the next similarity calculation. This process is repeated until the preset number of iterations is reached, at which point iteration stops, and the normalized dynamic weights from the last iteration are used as the final dynamic weights. If the weight change during iteration is less than a preset threshold, the iteration can be terminated early to ensure a balance between efficiency and accuracy. Based on the final dynamic weights and the preset prediction matrix, all multi-scale local feature tensors are weighted and summed. The prediction matrix is used to achieve dimensional mapping and semantic association from multi-scale local feature tensors to aggregated structured feature tensors. The weighted summation process aggregates the effective features of high-weight multi-scale local feature tensors into a unified aggregated layer input, achieving the fusion of multi-scale lesion features. Squash The function activates the input of the aggregation layer. This function can compress the magnitude of the vector to the range of 0 to 1, while preserving the directional information of the vector (the closer the magnitude is to 1, the more accurately the corresponding feature can represent the lesion structure; the closer the magnitude is to 0, the more invalid the feature is). This achieves the effect of strengthening the effective feature and suppressing the invalid feature, and finally outputs a structured aggregated structured feature tensor.
[0020] In one embodiment of the present invention, the trained multimodal medical data processing model further includes: a first linear layer and a second linear layer; The text semantic features are transformed using the first linear layer to obtain the text. token gather; The structured image features are transformed using a second linear layer to obtain the image. token gather.
[0021] In this embodiment, as Picture 2As shown, the multimodal medical data processing model also includes a first linear layer and a second linear layer. The first and second linear layers are used to perform dimensionality transformation (linear projection) on text semantic features and structured image features, respectively, unifying their feature dimensions to the same standard to adapt to the input requirements of the cross-modal feature fusion network. Simultaneously, they compress high-dimensional features into compact feature vectors, ultimately forming the text... token and images token Among them, the text token It is generated based on the semantic features of the text, but the text has a contextual order (such as the semantic difference between "cough" and "ground-glass nodules"). Therefore, after linear projection of the text, it will become the text. token Adding positional encoding allows subsequent networks to recognize the sequential relationships of semantic units, preventing the loss of sequence logic due to dimensionality transformation and ensuring the integrity of text semantics. Structured image features already contain spatial structural information (such as lesion location and boundaries), and after dimensionality transformation is completed through linear projection in the second linear layer, the image is directly generated. token For sets, no additional position encoding is required.
[0022] In one embodiment of the present invention, a cross-modal feature fusion network is used to perform prior weighted fusion of structured image features and text semantic features to obtain the matching probability between multimodal data and target lesions, including: Determine image token Each in the set token The first validity score, and those with a first validity score greater than the first validity threshold. token As a valid image token ; Determine text token The second validity score of each token in the set, and those with a second validity score greater than the second validity threshold. token As valid text token ; Based on valid images token and valid text token This yields the matching probability between multimodal data and the target lesion.
[0023] In this embodiment, because medical imaging data often contains artifacts, metal implants, black areas at the scanning edges, and other interference, the generated image may be affected. token The set misleads attention. Therefore, use a lightweight CNN convolutional neural network head (3 depthwise separable convolutions) for each token Scoring is performed to predict "whether it is an effective organization," and a first effectiveness score between 0 and 1 is obtained. Based on a learnable threshold (first effectiveness threshold) with an initial value of 0.3, those scores below the first effectiveness threshold are... token Masking, preserving the valid image token However, the text may contain vague descriptions (such as "maybe", "?", "not clearly seen"), these token Unreliable. Detected via an uncertainty detection head (linear layer + ... sigmoid For each token Scoring is performed to predict whether it is a deterministic description, resulting in a second validity score between 0 and 1. Based on a second validity threshold (preferably 0.5), scores below the second validity threshold are... token Masking yields the valid text. token .
[0024] In one embodiment of the present invention, based on effective images token and valid text token The matching probability between multimodal data and target lesions is obtained, including: valid image token Decompose it into an image orientation unit vector and an image length scalar; valid text token Decompose it into a unit vector of text direction and a scalar of text length; The first similarity between the image orientation unit vector and the text orientation unit vector is calculated based on cosine similarity. sdir ; The second similarity between the image length scalar and the text length scalar is calculated using a Gaussian kernel function. slen ; Based on preset weighting coefficients First similarity sdir Second similarity slen The similarity matrix is obtained using the following formula. Scaps0 :
[0025] Medical image data is input into a trained lesion probability judgment model to obtain the output of the lesion probability judgment model and the effective image. token The probability of lesions aligned with the grid Hles ; Based on lesion probability Hles and similarity matrix Scaps0 The first matrix is obtained. Scaps1 = Scaps0 · Hles ; Based on valid images token and valid text token Suppressing texture conflicts Ptexture and the first matrix Scaps1 The second matrix is obtained through the following formula. Scaps2 :
[0026] use SoftGround Function and Second Matrix Scaps2 The attention weight matrix is obtained using the following formula. A :
[0027] Based on attention weight matrix A Valid images token and valid text token The cross-modal fusion features are obtained through the following formula. F :
[0028] Based on cross-modal features F This yields the matching probability between multimodal data and the target lesion.
[0029] In this embodiment, because traditional cross-attention mechanisms only perform dot products, they confuse direction and length, which can easily lead to misjudgments in medicine. Therefore, each token A vector is decomposed into a unit vector of direction and a scalar of length. In the valid text... token In this model, the text direction unit vector represents the semantic direction, and the text length scalar represents the semantic strength. Effective images are calculated based on cosine similarity. token With valid text token The directional unit vector similarity, ranging from [-1, 1], is used to measure the consistency of semantic direction between the two vectors. The length scalar matching degree is calculated using a Gaussian kernel function, with a value ranging from [0, 1]. This is based on preset weight coefficients. The normalized directional similarity (first similarity) sdir ) and length matching degree (second similarity) slen Weighted fusion is performed to obtain a similarity matrix. The original medical images are then input into the pre-trained, frozen lightweight UNet model, and the output is compared with... token Grid-aligned lesion probability heatmap (lesion probability) Hles, The value range is [0,1]), and the similarity matrix is analyzed using this heatmap. Scaps Perform pixel-level weighted updates to strengthen the similarity weights corresponding to lesion regions, resulting in the first matrix. Scaps1 = Scaps0 · Hles Texture collision penalty term is applied using an exponential function. Ptexture The processing is performed, and the result is multiplied by the updated first matrix to suppress semantic conflicts (such as the correspondence between frosted glass and calcified textures). token (Similarity), to obtain the second matrix. Using... SoftGround The function is applied to the second matrix refined by two levels of medical a priori knowledge. Scaps2 Normalization is performed to obtain the attention weight matrix. A For valid images tokenWith valid text token Cross-modal feature weighted aggregation.
[0030] In one embodiment of the present invention, it further includes: The medical image data is subjected to grayscale normalization processing to normalize the grayscale values of the medical image data to the range of [0,1]. The medical image data after grayscale normalization is filtered for noise using a Gaussian filter to obtain preprocessed medical image data.
[0031] In this embodiment, the medical image data needs to be preprocessed before being input into the trained multimodal medical data processing model. The grayscale of the medical image data is normalized to the [0,1] range, and the input image size is normalized to 299×299×3 (width, height, and number of channels). A Gaussian filter is used to remove noise from the image. The image is then formatted to suit Inception-ResNet-V2 feature extraction.
[0032] Similarly, medical text data also needs to be preprocessed before being input into a trained multimodal medical data processing model. The medical text data undergoes cleaning and preprocessing steps such as word segmentation, stop word removal, noise and special character removal to ensure the data quality input to the model. Then, the WordPiece word segmentation method is used to cut the text into words or sub-words to better handle unknown words. Finally, the segmented text is converted into the input format required by the BERT network, including input ID, attention mask, and segment embedding. When extracting semantic features from the preprocessed medical text data using the BERT network, the pre-trained BERT word segmenter and model are first loaded. The preprocessed input ID and attention mask are input into the BERT network. The input embedding part of the BERT network adds the input ID, segment embedding, and position embedding to obtain the final embedding vector. The input embedding is then processed through the multi-layer self-attention mechanism and feedforward neural network of the transformer architecture. Finally, the hidden state of the last layer is compressed into a fixed-length vector through pooling layers, outputting the semantic text features.
[0033] According to another embodiment, the present invention provides a multimodal medical image data processing apparatus. Picture 3 A schematic block diagram of a multimodal medical image data processing device is shown. It is understood that this device can be implemented using any computing or processing power device, equipment, platform, or cluster of devices. Picture 3As shown, the device includes: a multimodal data processing module 300, a basic feature extraction module 302, a structured feature extraction module 304, a semantic feature extraction module 306, and a feature weighted fusion module 308. The main functions of each component are as follows: The multimodal data processing module inputs the multimodal data corresponding to the target lesion into the trained multimodal medical data processing model. The multimodal data includes medical image data and medical text data. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The basic feature extraction module is connected to the multimodal data processing module. It extracts multi-scale basic visual features from the medical image data through the basic visual feature extraction network to obtain a multi-scale basic visual feature set. The structured feature extraction module, connected to the basic feature extraction module, captures the spatial structure of the multi-scale basic visual feature set through the structured image feature extraction network to obtain structured image features including lesion shape, boundary and relative position; The semantic feature extraction module is connected to the multimodal data processing module. Through the text feature extraction network, it extracts semantic features from the medical text data to obtain text semantic features. The feature weighted fusion module, connected to the structured feature extraction module and the semantic feature extraction module, performs prior weighted fusion of the structured image features and the text semantic features through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion.
[0034] In a preferred embodiment, the structured image feature extraction network includes: a first convolutional layer, a normalization layer, a second convolutional layer, a max pooling layer, a dynamic weight aggregation layer, and a fully connected layer. The first convolutional layer consists of four layers, each with 128 channels, and the layers are connected in series. The second convolutional layer comprises multiple parallel layers, each with a different kernel scale. The structured image feature extraction network captures the spatial structure of the multi-scale basic visual feature set to obtain structured image features including lesion shape, boundaries, and relative positions, including: The multi-scale basic visual feature set is convolved sequentially by four cascaded first convolutional layers to extract local texture features and edge features, thereby obtaining a convolutional feature map. The normalization layer performs normalization calculations on the convolutional feature map to obtain a normalized feature map; By capturing the local spatial features of lesions of different sizes in the normalized feature map through each second convolutional layer including convolutional kernels of different scales, the multi-scale local feature tensor output by each second convolutional layer is obtained. The maximum pooling layer is used to unify the size of multiple multi-scale local feature tensors, and the size-unified multiple multi-scale local feature tensors are spliced together to obtain a multi-scale local feature set that integrates the local spatial features of lesions of all sizes. The dynamic weight aggregation layer calculates the connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer based on a preset dynamic weight optimization algorithm. The multi-scale local feature set is then weighted and summed based on the connection weights to obtain an aggregated structured feature tensor, which includes two dimensions: direction and length. The aggregated structured feature tensor is subjected to dimensionality mapping processing through the fully connected layer to obtain structured image features including lesion shape, boundary and relative position.
[0035] In a preferred embodiment, the step of calculating the connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer based on a preset dynamic weight optimization algorithm through the dynamic weight aggregation layer, and performing a weighted summation process on the multi-scale local feature set based on the connection weights to obtain an aggregated structured feature tensor, includes: Determine the initial correlation degree between each branch in the multi-scale feature extraction branch layer and the dynamic weight aggregation layer; The initial correlation degree is converted into dynamic weights with values between 0 and 1 by normalization calculation, wherein the sum of the dynamic weights of all multi-scale local feature tensors for the same dynamic weight aggregation layer is 1. The dynamic weights are iteratively calculated based on a preset dynamic weight optimization algorithm, and the final dynamic weights are obtained after a preset number of iterations. Based on the final dynamic weights and the preset prediction matrix, all multi-scale local feature tensors are weighted and summed to generate the input of the aggregation layer; pass Squash The function activates the input of the aggregation layer to obtain the activated aggregated structured feature tensor.
[0036] In a preferred embodiment, the trained multimodal medical data processing model further includes: a first linear layer and a second linear layer; The text semantic features are transformed using the first linear layer to obtain the text. token gather; The structured image features are subjected to dimensionality transformation processing through the second linear layer to obtain the image. tokengather.
[0037] As a preferred embodiment, the step of performing prior weighted fusion of the structured image features and the text semantic features through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion includes: Determine the image token Each in the set token The first validity score, and those with a first validity score greater than the first validity threshold. token As a valid image token ; Determine the text token Each in the set token The second validity score, which is greater than the second validity threshold. token As valid text token ; Based on the effective image token and the valid text token The matching probability between the multimodal data and the target lesion is obtained.
[0038] As a preferred embodiment, the step based on the effective image token and the valid text token Obtaining the matching probability between the multimodal data and the target lesion includes: The valid image token Decompose it into an image orientation unit vector and an image length scalar; The valid text token Decompose it into a unit vector of text direction and a scalar of text length; The first similarity between the image orientation unit vector and the text orientation unit vector is calculated based on cosine similarity. sdir ; The second similarity between the image length scalar and the text length scalar is calculated using a Gaussian kernel function. slen ; Based on preset weighting coefficients The first similarity sdir and the second similarity slen The similarity matrix is obtained using the following formula. Scaps0 :
[0039] The medical image data is input into a trained lesion probability judgment model to obtain the output of the lesion probability judgment model and the effective image. token The probability of lesions aligned with the grid Hles ; Based on the probability of the lesion Hles and the similarity matrix Scaps0 The first matrix is obtained. Scaps1= Scaps0·Hles ; Based on the effective image token and the valid text token Suppressing texture conflicts Ptexture and the first matrix Scaps1 The second matrix is obtained through the following formula. Scaps2 :
[0040] use SoftGround function and the second matrix Scaps2 The attention weight matrix is obtained using the following formula. A :
[0041] Based on the attention weight matrix A The effective image token and the valid text token The cross-modal fusion features are obtained through the following formula. F :
[0042] Based on the cross-modal features F The matching probability between the multimodal data and the target lesion is obtained.
[0043] As a preferred embodiment, it also includes: The medical image data is subjected to grayscale normalization processing to normalize the grayscale values of the medical image data to the range of [0,1]. The medical image data after grayscale normalization is subjected to noise filtering using a Gaussian filter to obtain preprocessed medical image data.
[0044] According to another embodiment, an electronic device is also provided, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, it implements a combination... Picture 1 The method.
[0045] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0046] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0047] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for processing multimodal medical image data, characterized in that, include: Multimodal data corresponding to the target lesion are input into a trained multimodal medical data processing model. The multimodal data includes medical image data and medical text data. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The medical image data is subjected to multi-scale basic visual feature extraction through the basic visual feature extraction network to obtain a multi-scale basic visual feature set. The structured image feature extraction network captures the spatial structure of the multi-scale basic visual feature set to obtain structured image features including lesion shape, boundary and relative position. The text feature extraction network is used to extract semantic features from the medical text data to obtain text semantic features. The structured image features and the text semantic features are fused using a priori weighted method through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion. The trained multimodal medical data processing model further includes: a first linear layer and a second linear layer; The text semantic features are transformed using the first linear layer to obtain the text. token gather; The structured image features are subjected to dimensionality transformation processing through the second linear layer to obtain the image. token gather; The step of performing prior weighted fusion of the structured image features and the text semantic features through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion includes: Determine the image token Each in the set token The first validity score, and those with a first validity score greater than the first validity threshold. token As a valid image token ; Determine the text token Each in the set token The second validity score, which is greater than the second validity threshold. token As valid text token ; Based on the effective image token and the valid text token The matching probability between the multimodal data and the target lesion is obtained. In a valid text token, the text direction unit vector is used to represent the semantic direction, and the text length scalar is used to represent the semantic strength. Based on the valid image token and the valid text token Obtaining the matching probability between the multimodal data and the target lesion includes: The valid image token Decompose it into an image orientation unit vector and an image length scalar; The valid text token Decompose it into a unit vector of text direction and a scalar of text length; The first similarity between the image orientation unit vector and the text orientation unit vector is calculated based on cosine similarity. sdir , sdir The value range is [-1, 1], which is used to measure the consistency of the semantic direction between the two. The second similarity between the image length scalar and the text length scalar is calculated using a Gaussian kernel function. slen , slen The value range is [0,1]; Based on preset weighting coefficients λ The first similarity sdir and the second similarity slen The similarity matrix is obtained using the following formula. Scaps0 : Scaps0=λ·(sdir+1) / 2+(1-λ)·slen ; The medical image data is input into a trained lesion probability judgment model to obtain the output of the lesion probability judgment model and the effective image. token The probability of lesions aligned with the grid Hles , Hles The value range is [0,1]; Based on the probability of the lesion Hles and the similarity matrix Scaps0 The first matrix is obtained. Scaps1=Scaps0· Hles To perform pixel-level weighted updates on the similarity matrix and strengthen the similarity weights corresponding to lesion regions; Based on the effective image token and the valid text token Suppressing texture conflicts Ptexture and the first matrix Scaps1 The second matrix is obtained through the following formula. Scaps2 To suppress texture conflicts through an exponential function Ptexture The process involves multiplying the result of the processing with the updated first matrix to suppress semantic conflicts. Scaps2 = Scaps1 · exp(-Ptexture) ; use softmax The function is applied to the second matrix refined by two levels of medical a priori knowledge. Scaps2 After normalization, the attention weight matrix is obtained using the following formula. A : A = softmax(Scaps2) ; Based on the attention weight matrix A The effective image token and the valid text token The cross-modal fusion features are obtained through the following formula. F : ; Based on the cross-modal features F The matching probability between the multimodal data and the target lesion is obtained.
2. The method according to claim 1, characterized in that, The structured image feature extraction network includes: a first convolutional layer, a normalization layer, a second convolutional layer, a max pooling layer, a dynamic weight aggregation layer, and a fully connected layer. The first convolutional layer consists of four layers, each with 128 channels, and the layers are connected in series. The second convolutional layer comprises multiple parallel layers, each with a different kernel scale. The structured image feature extraction network captures the spatial structure of the multi-scale basic visual feature set to obtain structured image features including lesion shape, boundaries, and relative positions, including: The multi-scale basic visual feature set is convolved sequentially by four cascaded first convolutional layers to extract local texture features and edge features, thereby obtaining a convolutional feature map. The normalization layer performs normalization calculations on the convolutional feature map to obtain a normalized feature map; By capturing the local spatial features of lesions of different sizes in the normalized feature map through each second convolutional layer including convolutional kernels of different scales, the multi-scale local feature tensor output by each second convolutional layer is obtained. The maximum pooling layer is used to unify the size of multiple multi-scale local feature tensors, and the size-unified multiple multi-scale local feature tensors are spliced together to obtain a multi-scale local feature set that integrates the local spatial features of lesions of all sizes. The dynamic weight aggregation layer calculates the connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer based on a preset dynamic weight optimization algorithm. The multi-scale local feature set is then weighted and summed based on the connection weights to obtain an aggregated structured feature tensor, which includes two dimensions: direction and length. The aggregated structured feature tensor is subjected to dimensionality mapping processing through the fully connected layer to obtain structured image features including lesion shape, boundary and relative position.
3. The method according to claim 2, characterized in that, The process involves calculating the connection weights between the multi-scale feature extraction branch layer and the dynamic weight aggregation layer based on a preset dynamic weight optimization algorithm, and then performing a weighted summation on the multi-scale local feature set based on these connection weights to obtain an aggregated structured feature tensor, including: Determine the initial correlation degree between each branch in the multi-scale feature extraction branch layer and the dynamic weight aggregation layer; The initial correlation degree is converted into dynamic weights with values between 0 and 1 by normalization calculation, wherein the sum of the dynamic weights of all multi-scale local feature tensors for the same dynamic weight aggregation layer is 1. The dynamic weights are iteratively calculated based on a preset dynamic weight optimization algorithm, and the final dynamic weights are obtained after a preset number of iterations. Based on the final dynamic weights and the preset prediction matrix, all multi-scale local feature tensors are weighted and summed to generate the input of the aggregation layer; pass Squash The function activates the input of the aggregation layer to obtain the activated aggregated structured feature tensor.
4. The method according to claim 1, characterized in that, Also includes: The medical image data is subjected to grayscale normalization processing to normalize the grayscale values of the medical image data to the range of [0,1]. The medical image data after grayscale normalization is subjected to noise filtering using a Gaussian filter to obtain preprocessed medical image data.
5. A multimodal medical image data processing device, characterized in that, For performing the method as described in any one of claims 1-4, comprising: The multimodal data processing module inputs the multimodal data corresponding to the target lesion into the trained multimodal medical data processing model. The multimodal data includes medical image data and medical text data. The multimodal medical data processing model includes a basic visual feature extraction network, a structured image feature extraction network, a text feature extraction network, and a cross-modal feature fusion network. The basic feature extraction module is connected to the multimodal data processing module. It extracts multi-scale basic visual features from the medical image data through the basic visual feature extraction network to obtain a multi-scale basic visual feature set. The structured feature extraction module, connected to the basic feature extraction module, captures the spatial structure of the multi-scale basic visual feature set through the structured image feature extraction network to obtain structured image features including lesion shape, boundary and relative position; The semantic feature extraction module is connected to the multimodal data processing module. Through the text feature extraction network, it extracts semantic features from the medical text data to obtain text semantic features. The feature weighted fusion module, connected to the structured feature extraction module and the semantic feature extraction module, performs prior weighted fusion of the structured image features and the text semantic features through the cross-modal feature fusion network to obtain the matching probability between the multimodal data and the target lesion.
6. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program, implements the method as described in any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-4.