Lung age evaluation method and system based on multi-modal data fusion

Through the lung age assessment method with multimodal data fusion, the integration of lung CT images and text information has solved the problem of inaccurate lung age prediction in the prior art, and achieved more accurate lung age assessment and the formulation of personalized treatment plans.

CN120299710APending Publication Date: 2025-07-11SHAN DONG MSUN HEALTH TECH GRP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510376824.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When predicting lung age, existing deep learning algorithms have insufficient imaging modal selection, feature extraction dynamics, multimodal fusion depth and age prediction rationality, resulting in inaccurate lung age assessment.

Method used

The multimodal data fusion method is adopted to integrate lung CT images and text information, and through multi-scale feature extraction, cross-modal deep fusion and loss function optimization, the deep interaction between images and text features is achieved, and the accuracy of lung age assessment is improved.

Benefits of technology

It improves the accuracy and clinical applicability of lung age assessment, can capture airway branches and micro-nodules more accurately, identify early signs of lung aging in advance, assist in the formulation of personalized treatment plans, and improve the primary prevention effect of the disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299710A_ABST
    Figure CN120299710A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of lung age evaluation, and discloses a lung age evaluation method and system based on multi-modal data fusion, and the method comprises the steps: obtaining data, and carrying out the preprocessing of the obtained data; the method comprises the following steps: acquiring CT image data, processing the CT image data to obtain convolution features, performing anisotropic multi-head attention processing on the convolution features to obtain Transform features, and fusing the convolution features and the Transform features to obtain final image features; performing multi-step processing on the text information to obtain final text features; performing cross-modal deep fusion on the final image features and the final text features to obtain fused features, and performing lung age evaluation and prediction by using the fused features; a loss function is defined, model parameters are optimized, and a trained lung age evaluation model is obtained. According to the method, the lung age evaluation accuracy and clinical applicability based on the CT image are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lung age assessment, and particularly to a lung age assessment method and system based on multi-modal data fusion. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Recently, a research team developed an artificial intelligence (AI) model that can estimate an individual's age using chest X-rays based on a deep learning model with 101,296 "chest X-rays", enabling it to accurately identify age-related chest X-ray image features, and confirmed that when the age estimated by the AI model is higher than the actual age, the individual is more likely to suffer from chronic diseases such as hypertension, hyperuricemia, and chronic obstructive pulmonary disease. This is a major breakthrough in the field of medical imaging. However, a chest X-ray shows all the structures of the chest on one film, and this overlapping image may cause the organizational structures to obscure each other, making it difficult to observe the details of some small lesions clearly. In contrast, CT is a tomographic imaging technology that can divide the human body into hundreds of slices to observe lesions, effectively avoiding tissue overlap and showing the details of the lungs more clearly, detecting tiny lesions that cannot be found by chest X-rays, such as early lung cancer and pulmonary nodules. These are of great significance for accurately assessing the degree of lung aging of patients. Therefore, based on lung CT images, the functional status and aging degree of the lungs can be more objectively evaluated, providing a better option for predicting the biological age of the lungs.

[0004] With the rapid development of artificial intelligence technology, especially the wide application of deep learning algorithms in the field of medical imaging, new ideas and methods have been provided for the research of lung aging biomarkers. Different from traditional manual diagnosis, deep learning algorithms can automatically process and analyze a large amount of lung CT image data and medical history reports, and then accurately predict the "lung age" based on individual lung CT images. However, existing deep learning algorithms have deficiencies in aspects such as imaging modality selection, dynamic feature extraction, depth of multi-modal fusion, and rationality of age prediction when performing "lung age" prediction. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a lung age assessment method and system based on multi-modal data fusion, which integrates lung CT images and text information, has powerful dynamic feature capture and anatomical structure perception capabilities, realizes multi-modal deep interactive lung age prediction, and significantly improves the accuracy and clinical applicability of lung age assessment based on CT images.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a lung age assessment method based on multi-modal data fusion, comprising the following steps: Obtain lung CT image data and related text information, and preprocess the obtained data; Perform multi-scale feature extraction on the CT image data, enhance the feature maps of each scale using differential features, splice the enhanced feature maps of each scale to obtain convolutional features, perform anisotropic multi-head attention processing on the convolutional features to obtain Transformer features, and fuse the convolutional features and Transformer features to obtain the final image features; Extract the features of the text information to obtain a word vector representation, extract the key entities of the text information, learn the representation of the key entities in the knowledge graph, and fuse the representation of the key entities in the knowledge graph with the word vector representation to obtain the final text features; Perform cross-modal deep fusion on the final image features and text features to obtain the fused features, and use the fused features to perform lung age assessment prediction; Define a loss function, optimize the model parameters, and obtain a trained lung age assessment model.

[0007] As an alternative implementation, enhancing the feature maps of each scale using differential features specifically includes: Use the finite difference method to calculate the differential features of each scale feature map along three directions respectively, splice the differential features of each feature map along the three directions calculated to obtain the differential features of each feature map, calculate the adaptive coefficient of each scale feature map according to the variance of each scale feature map, and fuse the differential features of each feature map with the original feature map based on the adaptive coefficient to obtain the enhanced feature maps of each scale.

[0008] As an alternative implementation, performing anisotropic multi-head attention processing on the convolutional features specifically includes: Calculate the curvature of the convolutional feature map, introduce the curvature feature into the standard position encoding of the convolutional feature map to obtain the convolutional feature map after curvature-aware position encoding, construct an anisotropic weight matrix, perform multi-head attention calculation on the convolutional feature map after curvature-aware position encoding, and incorporate the anisotropic weight matrix into the calculation of the attention score to obtain Transformer features.

[0009] As an alternative implementation, learning the representation of the key entities in the knowledge graph specifically includes: Construct a lung health knowledge graph, use named entity recognition technology to extract the key entities of the text information, and learn the representation of the key entities in the knowledge graph through a graph attention network.

[0010] As an alternative implementation, perform cross-modal deep fusion on the final image features and text features, specifically as follows: Using the final image features as queries, the final text features as keys and values, calculate the attention scores, perform weighted summation on the text features according to the attention scores to obtain the text features after interacting with the image features, using the final text features as queries, the final image features as keys and values, calculate the reverse attention scores, perform weighted summation on the image features according to the reverse attention scores to obtain the image features after interacting with the text features, fuse and update the interacted image and text features, and obtain the fused features after multiple iterations.

[0011] As an alternative implementation, preprocess the acquired data, specifically as follows: Normalize the pulmonary CT image data, perform spatial resampling using the linear interpolation algorithm, and adjust the voxel spacing; Use natural language processing tools to process the relevant text information, remove the noise in the text, perform word segmentation and part-of-speech tagging operations on the processed text, split the text into individual words, and tag the part of speech of each word; Convert the time series data into the time difference relative to the current time.

[0012] In a second aspect, the present invention provides a lung age assessment system based on multi-modal data fusion, including: A data acquisition and preprocessing module, configured to: acquire pulmonary CT image data and relevant text information, and preprocess the acquired data; An image data processing module, configured to: perform multi-scale feature extraction on the CT image data, enhance the feature maps of each scale using differential features, splice the enhanced feature maps of each scale to obtain convolutional features, perform anisotropic multi-head attention processing on the convolutional features to obtain Transformer features, and fuse the convolutional features and Transformer features to obtain the final image features; A text information processing module, configured to: extract the features of the text information to obtain word vector representations, extract the key entities of the text information, learn the representations of the key entities in the knowledge graph, and fuse the representations of the key entities in the knowledge graph with the word vector representations to obtain the final text features; A feature fusion and lung age assessment module, configured to: perform cross-modal deep fusion on the final image features and text features to obtain the fused features, and use the fused features to perform lung age assessment prediction; A model training module, configured to: define a loss function, optimize the model parameters, and obtain a trained lung age assessment model.

[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0015] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a lung age assessment method and system based on multi-modal data fusion, which realizes accurate quantification and assessment of lung age driven by multi-modal data. By integrating lung CT images and text information, the limitations of traditional single-modal assessment are broken through, and it no longer relies one-sidedly on a single image or parameter. This helps to provide more targeted individualized lung function assessments for patients with chronic obstructive pulmonary disease, and further assist in formulating graded treatment plans. At the same time, it can also identify early signs of lung aging in asymptomatic smokers in advance, contributing to the primary prevention of diseases such as lung cancer, which is of great significance for the early detection and early intervention of diseases.

[0017] The present invention proposes a lung age assessment method and system based on multi-modal data fusion, which has strong capabilities of capturing dynamic features and perceiving anatomical structures. By using dynamic sparse convolution to optimize neighborhood sampling, complex structures such as airway branches and micro-nodules can be accurately captured. Combining the discriminative ability of curvature-aware Transformer for lung tissue texture, the prediction accuracy of lung age is greatly improved.

[0018] The present invention proposes a lung age assessment method and system based on multi-modal data fusion, which realizes multi-modal deep interactive lung age prediction. The bidirectional cross-modal attention mechanism realizes deep semantic alignment of image and text features, and excavates hidden associations. The constraint of the hybrid loss function ensures that the predicted lung age is accurate and conforms to the population age distribution, improving the clinical interpretability.

[0019] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0021] Figure 1 Flowchart of a lung age assessment method based on multi-modal data fusion provided by Embodiment 1 of the present invention. Specific implementation manners

[0022] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0024] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0026] Embodiment 1 As Figure 1 shown, this embodiment provides a lung age assessment method based on multi-modal data fusion, including the following steps: S1. Obtain lung CT image data and relevant text information, and preprocess the obtained data; S2. Extract multi-scale features from the CT image data, enhance the feature maps of each scale using differential features, splice the enhanced feature maps of each scale to obtain convolutional features, perform anisotropic multi-head attention processing on the convolutional features to obtain Transformer features, and fuse the convolutional features and Transformer features to obtain the final image features; S3. Extract the features of the text information to obtain a word vector representation, extract the key entities of the text information, learn the representation of the key entities in the knowledge graph, and fuse the representation of the key entities in the knowledge graph with the word vector representation to obtain the final text features; S4. Perform cross-modal deep fusion on the final image features and text features to obtain the fused features, and use the fused features to perform lung age assessment prediction; S5. Define a loss function, optimize the model parameters, and obtain a trained lung age assessment model.

[0027] Collect a large amount of lung CT image data from different medical institutions, ensuring that the samples cover different age groups, lung health conditions, as well as text information such as gender, smoking history, and lung infection status. These text information should record in detail the patient's gender (male / female), exact age, smoking frequency and smoking duration, and the type (such as Streptococcus pneumoniae infection) and infection time of lung infection, etc.

[0028] Preprocess the collected data. Among them, the preprocessing stage of lung CT image data mainly includes normalization, spatial resampling, and size adjustment.

[0029] In the normalization process, traverse all the images in the lung CT image dataset, and count the minimum value and the maximum value of the pixel values. Then, normalize each pixel value in each image according to the formula to map the pixel values to the interval [0, 1], and output the normalized lung CT image data.

[0030] For spatial resampling and size adjustment, for lung CT image data, use the linear interpolation algorithm for spatial resampling to uniformly adjust the voxel spacing to (1mm, 1mm, 1mm). If the original image voxel spacing is different, then calculate the interpolation according to the new spacing requirements on the basis of the original pixel values to re-determine the position and value of each pixel point, ensuring spatial consistency. In the x and y dimensions, crop or pad the image to a fixed size of 224×224 to ensure that all images are of the same size in these two dimensions. At the same time, uniformly extract 25 slices in the z-axis (slice direction), and finally output the preprocessed lung CT image data in the format of 25×224×224.

[0031] Preprocess the text data: Use natural language processing tools such as NLTK and spaCy to remove the noise in the text, including typos, special symbols, redundant spaces, etc. At the same time, perform word segmentation and part-of-speech tagging operations on the processed text, split the text into individual words, and tag the part of speech of each word to provide a basis for subsequent text information extraction.

[0032] For time series data, such as infection time, smoking start time, etc., convert them into time differences relative to the current time.

[0033] Further process the preprocessed data.

[0034] Feature extraction from lung CT images: It includes two stages, namely convolutional feature extraction and Transformer feature extraction. Combining the local feature capture ability of the convolutional neural network (CNN) and the global feature modeling advantage of Transformer. First, perform preliminary feature extraction on the lung CT image through a 3D convolutional layer to obtain the feature map F CNN . Then use F CNN as the input of the Transformer to learn and fuse global features.

[0035] Among them, the specific process of convolutional feature extraction is as follows: Multi-scale dynamic sparse convolution: Initialization operation: Input the lung CT image, and its size is expressed as H × W × D , where H represents the height, W represents the width, D represents the depth. Determine the three-dimensional neighborhood points N of sparse sampling. By default, take the 27-neighborhood, and at the same time set the learnable position offset Δ p and convolution kernels of different scales, such as 3×3×3, 5×5×5, etc.

[0036] Sampling point position adjustment: Taking each voxel in the image as the center, perform sparse sampling according to the set neighborhood N . For each sampling point, according to the formula p ′ = p + Δ p adjust the position, where p is the original sampling point position, p ′ is the adjusted position. In this way, more representative sampling points are obtained at different scales.

[0037] ; α c,i,j,k Use the Softmax function to normalize the dot product result of the weight matrix W d and the global feature F global . The Softmax function ensures that α c,i,j,k values are between 0 and 1, and the sum of the coefficients at all positions is 1. These coefficients will be used to weight the features at different positions.

[0038] ; Outer layer summation represents summation over all channels. HereC is the number of channels. The inner summation represents summing over all points in the sparse sampled three-dimensional neighborhood point set .

[0039] For each neighborhood point ([[]] i , j , k ), calculate the weighted convolution result: α c,i,j,k is the coefficient calculated in the first step, used to weight the contribution of this point. W i,j,k is the convolution kernel weight, used to perform convolution operation at this point. x ( p + Δ p i,j,k ) represents the value at the adjusted position on the input feature map according to the learnable position offset Δ p i,j,k p + Δ p i,j,k .

[0040] Differential feature enhancement: Differential feature calculation: For each scale feature map obtained by multi-scale dynamic sparse convolution , use the finite difference method to calculate the differential feature. Taking the first-order derivative as an example, the formula for calculating the differential feature along the x direction is: . Similarly, the differential features along the y and z directions can be calculated. Concatenate the differential features in three directions to obtain the complete differential feature .

[0041] Adaptive coefficient calculation: Calculate the adaptive coefficient according to the variance of the feature map , and the formula is . The variance can reflect the degree of change of the feature. Transform the variance into an adaptive coefficient through the softmax function, which is used to dynamically adjust the enhancement degree of the differential feature.

[0042] Feature fusion: Fuse the differential feature with the original feature map to obtain the enhanced feature map , and the formula is . Through this fusion method, the perception ability of the feature map to details such as texture is enhanced.

[0043] Feature concatenation: Concatenate the enhanced feature maps at different scales​ ( s = 1, 2, ⋯, S , S being the number of scales) are concatenated along the channel dimension. Assume that the number of channels of each is C s , then the number of channels of the concatenated convolutional feature F CNN is , and the height, width, and depth remain unchanged. Through feature concatenation, rich feature information at different scales is integrated, providing a more comprehensive data basis for subsequent processing.

[0044] The specific process of Transformer feature extraction is as follows: Curvature-Aware Position Encoding: Curvature calculation: For the feature map F CNN obtained from the convolutional part, the Sobel-Hessian operator is used to calculate the curvature feature. The Sobel-Hessian operator calculates the gradients of the feature map in different directions to obtain second-order derivative information, and thus constructs the Hessian matrix . Calculate the curvature according to the Hessian matrix K = det( H ), where det( H ) represents the determinant of the Hessian matrix, and the curvature K reflects the degree of bending of the feature map in the local area.

[0045] Standard position encoding calculation: Calculate the standard position encoding PE standard , which is usually generated according to the size and position information of the feature map and is used to add position-related information to the features. The common standard position encoding formula is: , , where pos represents the position index, i represents the dimension index, d m odel is the feature dimension.

[0046] Curvature-Aware Position Encoding generation: Introduce the curvature feature into the standard position encoding, and obtain the Curvature-Aware Position Encoding through the formula PE curvature , where λ is the learnable curvature coefficient. By adjusting the value of λ , the influence degree of the curvature feature on the position encoding can be controlled, enabling the model to better capture the spatial position and local structure information of the features.

[0047] Anisotropic multi-head attention: Feature sequence preparation: Reshape the feature map after curvature-aware positional encoding F CNN into a sequence form X = F CNN + PE curvature , and use it as the input of the multi-head attention mechanism. At this time, X has a dimension of L × B × d m odel , where L is the sequence length, B is the batch size, d m odel is the feature dimension.

[0048] Anisotropic weight matrix construction: According to the lung anatomical plane directions, namely axial, coronal, and sagittal, assign learnable importance parameters γm ( m ∈{ axial , coronal , sagittal}), and construct the directional weight matrix Θ d . Θ d is decomposed into in the anatomical plane directions, where W m is the weight matrix corresponding to the plane.

[0049] Multi-head attention calculation: During the multi-head attention calculation process, use the formula for calculation, where Q = XW Q , K = XW K , V = XW V , W Q 、 W K 、 W V are learnable weight matrices, d k is the dimension of the key vector. When calculating the attention scores, incorporate the directional weight matrix Θ d into it, that is , enabling the model to more precisely allocate attention based on the structural features in different directions of the lungs, enhancing the adaptability to the lung structure directionality.

[0050] Heterogeneous Feature Gated Fusion: Global Average Pooling Calculation: Respectively perform global average pooling operations on the convolutional features F CNN and the Transformer features after anisotropic multi-head attention processing F attention . The global average pooling formula is , obtaining GAP ( F CNN ) and GAP ( F attention ), compressing their dimensions into one-dimensional vectors for subsequent calculation of the fusion gate.

[0051] Fusion Gate Calculation: Concatenate GAP ( F CNN ) and GAP ( F attention ) to obtain the input vector GAP ( F CNN ); GAP ( F attention )]. Calculate the fusion gate W g through a fully connected layer containing a learnable weight matrix g and the Sigmoid function, and the formula is g = Sigmoid( Wg GAP ( F CNN ); GAP ( F attention )]). The value of the fusion gate g ranges from 0 to 1 and is used to dynamically adjust the fusion ratio of the convolutional features and the Transformer features.

[0052] Feature Fusion: Weightedly fuse the convolutional features g and the Transformer features F CNN according to the fusion gate F attention , and the formula is F final = g ⊙ FCNN + (1 - g ) ⊙​F attention , where ⊙ represents element-wise multiplication. Through this fusion method, the local feature capture ability of the convolutional neural network and the global feature modeling advantage of the Transformer are fully combined to obtain a more representative final feature representation F final , enhancing the comprehensive perception ability of the lung health status.

[0053] This collaborative method enables the model to not only focus on the local detailed features of the lungs, such as tiny nodules and texture changes, but also capture the overall structural features, enhancing the comprehensive perception ability of the lung health status.

[0054] The specific steps for feature extraction of text information are as follows: Knowledge graph-guided text feature learning: Construct a lung health knowledge graph, including entities such as diseases, symptoms, risk factors, and their relationships. For the input text information, use named entity recognition (NER) technology to extract key entities, such as "smoking history" and "history of lung infection".

[0055] Learn the representation of text entities in the knowledge graph through the graph attention network (GAT). For node i in the knowledge graph, its feature representation is h i , calculate the attention coefficient α ij : ; where α is a learnable attention vector, W is a weight matrix, N i is the set of neighbor nodes of node i, and || represents the concatenation operation.

[0056] The updated feature of node i .

[0057] Fuse the feature representation of text entities in the knowledge graph with the word vector representation of the text itself, and map it to a fixed dimension through a fully connected layer to obtain a text feature vector with more semantic understanding and correlation information.

[0058] Time series analysis to capture dynamic information: For text content containing time information, such as the time of lung infection and the starting time of smoking, convert it into time series data. Use the long short-term memory network (LSTM) to model the time series and capture the changing trend of text information over time.

[0059] Let the input time series data be T = [t1, t2,..., t T , the hidden state of LSTM be h t , and the cell state be c t .

[0060] Calculate the forget gate ft Input gate i t Output gate o t And candidate cell state : ; ; ; ; Update the cell state and calculate the hidden state where σ is the Sigmoid function and tanh is the hyperbolic tangent function.

[0061] Use the last hidden state of the LSTM as the feature representation of the time series text information and fuse it with other text features to more comprehensively reflect the dynamic changes in the patient's lung health status.

[0062] Based on the above operations, obtain the final image features and the final text features, and perform cross-modal deep fusion on the final image features and text features to obtain the fused features.

[0063] The specific process of cross-modal deep fusion is as follows: Construct a cross-modal interaction network to achieve deep interaction between image features and text features. The network consists of multiple interaction layers, and each interaction layer contains a cross-modal attention mechanism and a feature update module.

[0064] In the cross-modal attention mechanism, use the image features as the query, the text features as the key and value, and calculate the attention score: where W Q and W K are learnable weight matrices and d is the feature dimension.

[0065] Weighted sum the text features according to the attention score to obtain the text features after interacting with the image features W V is a learnable weight matrix.

[0066] Similarly, use the text features as the query, the image features as the key and value, and calculate the reverse attention score A text2img to obtain the image features after interacting with the text features .

[0067] Through the feature update module, fuse and update the interacted image and text features, such as ,

[0068] When iterating through multiple interaction layers, the feature update for each iteration is calculated by the formula and finally, after multiple iterations, a fully interacted and fused feature representation F fusion is obtained. Here, W 2 is the weight matrix, b 2 is the bias term. In actual calculation, the steps of the above cross-modal attention mechanism and feature update module are continuously repeated to enable in-depth interaction between image features and text features until a stable F fusion is obtained.

[0069] This cross-modal interaction and fusion method can promote the information flow between image and text features, enhancing the model's ability to understand and utilize multi-modal data.

[0070] The fused features are used for lung age assessment and prediction.

[0071] The fused feature vector F fusion is input into a prediction head composed of multiple fully connected layers. The number of neurons in the first fully connected layer is set to 256. The input feature vector is multiplied by the weight matrix of this fully connected layer (the dimension is determined according to the input feature vector and the number of neurons), then the bias term is added, and then a non-linear transformation is performed through the ReLU activation function to output an intermediate feature vector; then the intermediate feature vector continues to be input into the next fully connected layer (the number of neurons is 128), and the same matrix multiplication, bias addition, and ReLU activation operations are performed; finally, it is input into the output layer (the number of neurons is 1) to output the predicted "lung age" value. Batch Normalization operations are added after each fully connected layer. By normalizing the input data of each layer, the training process is accelerated and the stability of the model is improved, preventing overfitting.

[0072] A hybrid loss function combining mean squared error (MSE) and age distribution consistency loss is adopted. For the mean squared error loss, the difference between the predicted "lung age" and the true "lung age" is calculated by the formula where N is the number of samples, is the predicted "lung age" of the th sample, and is its true "lung age".

[0073] By comparing the distribution of the predicted "lung age" with the age distribution in the real dataset, the Kullback-Leibler Divergence (KL divergence) is used to measure the difference between the two, ensuring that the distribution of the prediction results at different age groups conforms to the actual situation. First, the probability distribution of the predicted "lung age" and the age distribution in the real dataset are statistically analyzed, and then according to the KL divergence formula Calculate the age distribution consistency loss.

[0074] The final hybrid loss function is Let the adjustment coefficient be λ. The model updates the parameters of the model (including the learnable parameters in each previous module, such as convolution kernel weights, fully connected layer weights, etc.) by backpropagation according to the value of the loss function through an optimization algorithm based on gradient descent (such as the Adam optimizer, etc.), continuously optimizing the model to make its prediction results closer to the true "lung age", and outputting the finally predicted "lung age" label.

[0075] Example 2 This example provides a lung age assessment system based on multi-modal data fusion, including: A data acquisition and preprocessing module, configured to: acquire lung CT image data and related text information, and preprocess the acquired data; An image data processing module, configured to: perform multi-scale feature extraction on the CT image data, enhance the feature maps of each scale using differential features, splice the enhanced feature maps of each scale to obtain convolution features, perform anisotropic multi-head attention processing on the convolution features to obtain Transformer features, and fuse the convolution features and Transformer features to obtain the final image features; A text information processing module, configured to: extract the features of the text information to obtain a word vector representation, extract the key entities of the text information, learn the representation of the key entities in the knowledge graph, and fuse the representation of the key entities in the knowledge graph with the word vector representation to obtain the final text features; A feature fusion and lung age assessment module, configured to: perform cross-modal deep fusion on the final image features and text features to obtain the fused features, and use the fused features to perform lung age assessment prediction; A model training module, configured to: define a loss function, optimize the model parameters, and obtain a trained lung age assessment model.

[0076] It should be noted here that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0077] In more examples, there is also provided: An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Example 1 is completed. For the sake of brevity, it will not be elaborated here.

[0078] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0079] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0080] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0081] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0082] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented.

[0083] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules may be combined or divided as needed. The machine-executable instructions for program modules may be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.

[0084] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0085] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0086] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0087] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A lung age assessment method based on multi-modal data fusion, characterized in that, It includes the following steps: Obtain lung CT image data and related text information, and preprocess the obtained data; Extract multi-scale features from the CT image data, enhance the feature maps at each scale using differential features, splice the enhanced feature maps at each scale to obtain convolutional features, perform anisotropic multi-head attention processing on the convolutional features to obtain Transformer features, and fuse the convolutional features and Transformer features to obtain the final image features; Extract the features of the text information to obtain the word vector representation, extract the key entities of the text information, learn the representation of the key entities in the knowledge graph, and fuse the representation of the key entities in the knowledge graph with the word vector representation to obtain the final text features; Perform cross-modal deep fusion on the final image features and text features to obtain the fused features, and use the fused features for lung age assessment and prediction; Define the loss function, optimize the model parameters, and obtain the trained lung age assessment model.

2. The method for lung age assessment based on multi-modal data fusion according to claim 1, wherein Enhance the feature maps at each scale using differential features, specifically: Use the finite difference method to calculate the differential features of each scale feature map along three directions, splice the differential features of each feature map along the three directions calculated to obtain the differential features of each feature map, calculate the adaptive coefficient of each scale feature map according to the variance of each scale feature map, and fuse the differential features of each feature map with the original feature map based on the adaptive coefficient to obtain the enhanced feature maps at each scale.

3. The method for lung age assessment based on multimodal data fusion according to claim 1, wherein Perform anisotropic multi-head attention processing on the convolutional features, specifically: Calculate the curvature of the convolutional feature map, introduce the curvature feature into the standard position encoding of the convolutional feature map to obtain the convolutional feature map after curvature-aware position encoding, construct an anisotropic weight matrix, perform multi-head attention calculation on the convolutional feature map after curvature-aware position encoding, and incorporate the anisotropic weight matrix into the calculation of the attention score to obtain the Transformer features.

4. The lung age assessment method based on multimodal data fusion according to claim 1, characterized in that, Learn the representation of the key entities in the knowledge graph, specifically: Construct a lung health knowledge graph, use the named entity recognition technology to extract the key entities of the text information, and learn the representation of the key entities in the knowledge graph through the graph attention network.

5. The lung age assessment method based on multimodal data fusion according to claim 1, characterized in that, Perform cross-modal deep fusion on the final image features and text features, specifically: Use the final image features as the query, the final text features as the key and value, calculate the attention score, perform weighted summation on the text features according to the attention score to obtain the text features after interacting with the image features, use the final text features as the query, the final image features as the key and value, calculate the reverse attention score, perform weighted summation on the image features according to the reverse attention score to obtain the image features after interacting with the text features, fuse and update the interacted image and text features, and obtain the fused features after multiple iterations.

6. The method for lung age assessment based on multimodal data fusion according to claim 1, wherein Preprocess the obtained data, specifically: Normalize the lung CT image data, perform spatial resampling using the linear interpolation algorithm, and adjust the voxel spacing; Use natural language processing tools to process relevant text information, remove noise in the text, perform word segmentation and part-of-speech tagging operations on the processed text, split the text into individual words, and tag the part of speech of each word; Convert time series data into time differences relative to the current time.

7. A lung age assessment system based on multimodal data fusion, characterized in that Comprising: A data acquisition and preprocessing module, configured to: acquire lung CT image data and relevant text information, and preprocess the acquired data; An image data processing module, configured to: perform multi-scale feature extraction on CT image data, enhance the feature maps of each scale using differential features, splice the enhanced feature maps of each scale to obtain convolutional features, perform anisotropic multi-head attention processing on the convolutional features to obtain Transformer features, and fuse the convolutional features and Transformer features to obtain final image features; A text information processing module, configured to: extract the features of text information to obtain word vector representations, extract the key entities of text information, learn the representations of key entities in the knowledge graph, and fuse the representations of key entities in the knowledge graph with the word vector representations to obtain final text features; A feature fusion and lung age assessment module, configured to: perform cross-modal deep fusion on the final image features and text features to obtain fused features, and use the fused features to perform lung age assessment prediction; A model training module, configured to: define a loss function, optimize model parameters, and obtain a trained lung age assessment model.

8. An electronic device, characterized in that, Comprising a memory, a processor, and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method according to any one of claims 1-6 is completed.

9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method according to any one of claims 1-6 is completed.

10. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Cited By

  • Intelligent evaluation system for lung collapse under thoracoscope

    CN121724993A

  • A thoracoscope lung collapse intelligent evaluation system

    CN121724993B