A cardiovascular lesion prediction system and method for constructing a deep network model
The cardiovascular lesions prediction system constructed through deep network models uses clinical text data and earlobe image features, combined with statistical analysis and image processing, to solve the problems of high model cost and high computational complexity in the existing technology, and realizes accurate prediction and early screening of cardiovascular lesions.
Patent Information
- Application Number
- CN202510428591.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-08
AI Technical Summary
When the prior art deeply explores the wrinkle characteristics of earlobes to predict coronary heart disease, the model is costly and has high computational complexity, making it difficult to quickly and conveniently evaluate cardiovascular lesions.
The cardiovascular lesions prediction system constructed using deep network model is used to collect the patient's clinical text data and binaural earlobe images, and the earlobe feature map is extracted using the object detection device. The enhancement device performs affine transformation and Gaussian fuzzy processing, combining statistical intergroup difference detection and random forest screening of the feature selection device. Finally, the feature fusion is calculated through the attention mechanism of the prediction device to obtain the prediction results.
It reduces the cost of the model, simplifies the computational complexity, improves the accuracy of prediction results, and achieves the purpose of early screening of cardiovascular lesions.
Smart Images

Figure CN119943414B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and artificial intelligence, and more particularly, to a cardiovascular lesion prediction system and method for constructing a deep network model. Background Art
[0002] The impact of cardiovascular lesions on human health is self-evident. Coronary atherosclerotic heart disease (CAD) is the most representative. Since the grading of earlobe wrinkles is significantly positively correlated with the severity of CAD, and earlobe images, as a common body surface feature, have the characteristics of easy acquisition and non-invasiveness. Therefore, screening for CAD based on earlobe wrinkles has potential application prospects. Traditional technical solutions usually judge earlobe wrinkles based on experience, which is somewhat subjective and lacks standard classification definitions and objective quantification bases.
[0003] However, the prior art discloses a technical solution of "artificial intelligence + radiomics" to screen for CAD through facial features and earlobe image features, and evaluate the cardiovascular lesion conditions triggered by CAD. This technical solution applies artificial intelligence technology to the processing, analysis, and interpretation of medical image data, analyzes the image data of a large number of cases, deeply mines the unique image features of diseases, and makes result predictions on this basis.
[0004] However, in the process of deeply mining earlobe wrinkle features for CAD prediction with this "artificial intelligence + radiomics" technical solution, due to the need for dual image processing, there are technical defects such as a relatively high model cost, especially a large computational complexity, and thus it cannot quickly and conveniently evaluate cardiovascular lesions. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to overcome the technical defects of relatively high model cost and large computational complexity existing in the prior art in the process of deeply mining earlobe wrinkle features for CAD prediction. To overcome the above defects of the prior art, the present invention provides a cardiovascular lesion prediction system and method for constructing a deep network model, including a cardiovascular lesion prediction system for constructing a deep network model and a cardiovascular lesion prediction method for constructing a deep network model.
[0006] A cardiovascular lesion prediction system for constructing a deep network model provided by the present invention includes:
[0007] An acquisition device configured to acquire clinical text data and bilateral earlobe images of a patient;
[0008] A target detection device, which communicates with the acquisition device and is configured to extract the earlobe part from the binaural earlobe images in a manner of extracting an object through a bounding box to obtain an earlobe feature map;
[0009] An enhancement device, which communicates with the target detection device and is configured to perform an affine transformation on the earlobe feature map to obtain transformed images with different perspectives, scales and shapes, and then perform Gaussian blur processing on the transformed images to obtain a feature-enhanced image;
[0010] A feature selection device, which communicates with the acquisition device and is configured to sequentially perform statistical between-group difference detection, regression analysis and importance ranking of the random forest algorithm on the clinical text data to obtain a clinical data feature set;
[0011] A prediction device, which communicates with the enhancement device and the feature selection device simultaneously, and is configured to calculate the importance of the clinical data feature set for the feature-enhanced image through an attention mechanism, obtain the features after the fusion of the clinical data feature set and the feature-enhanced image based on this importance, and then obtain the cardiovascular lesion prediction result of the patient by using the fused features through a fully connected mapping.
[0012] The cardiovascular lesion prediction system for constructing a deep network model disclosed by the present invention addresses the above-listed technical deficiencies. By setting up a collection device to collect the clinical text data and bilateral earlobe images of patients, for the bilateral earlobe images, the target detection device extracts the earlobe parts in the bilateral earlobe images in the way of extracting objects with bounding boxes to obtain earlobe feature maps. Subsequently, the enhancement device performs affine transformation and Gaussian blur processing on the earlobe feature maps to obtain feature-enhanced images, and the feature-enhanced images are received by the set prediction device. For the clinical text data, by setting up a feature selection device to perform statistical inter-group difference detection, regression analysis, and random forest importance screening detection on it in sequence to obtain a clinical data feature set, and the clinical data feature set is received by the set prediction device. Finally, the prediction device calculates the importance of the clinical data feature set for the feature-enhanced images through an attention mechanism and obtains the fused features, and then uses the fused features through a fully connected mapping to obtain the cardiovascular lesion prediction result of the patient. It can be seen that the cardiovascular lesion prediction system disclosed by the present invention abandons the existing technical route of screening for coronary heart disease through facial features and earlobe image features of the face, adopts the clinical text data and bilateral earlobe images of patients, and combines the technical routes of statistical analysis and image processing, avoiding the technical deficiency of large computational complexity caused by dual image processing. Moreover, since the process of obtaining the clinical data feature set from the clinical text data is to perform statistical inter-group difference detection, regression analysis, and random forest importance screening detection in sequence, not only can important features of the clinical data be obtained, but also the model layout is easy and the model cost is low. In addition, since the prediction device calculates the importance of the clinical data feature set for the feature-enhanced images through an attention mechanism and obtains the cardiovascular lesion prediction result of the patient through a fully connected mapping using the fused features, it realizes the combination of text features and image features to obtain the cardiovascular lesion prediction result, which can ensure the accuracy of the prediction result and achieve the purpose of early screening for cardiovascular lesions.
[0013] In a possible implementation manner, the collection device includes:
[0014] A text module, communicating with the feature selection device, is configured to collect the clinical text data of the patient; the clinical text data includes the patient's demographic characteristics, medical history, risk factors, and laboratory examination information;
[0015] A photographing module, communicating with the target detection device, is configured to collect the bilateral earlobe images of the patient;
[0016] Adopting the above solution can achieve obtaining the clinical text data and bilateral earlobe images of the patient in parallel, ensuring the orderliness of separate processing of the text data and the images.
[0017] In a possible implementation, the target detection device includes, arranged in sequence along the running direction:
[0018] A mosaic module, configured to scale the binaural earlobe image to a predetermined standard size by means of adaptive image scaling to obtain a standard size image;
[0019] A backbone network module, configured to convert the standard size image into multiple feature maps with different scales by means of a fusion of a focusing operation, a convolution operation, and a feature pyramid pooling operation;
[0020] A neck network module, configured to perform feature fusion on all the feature maps output by the backbone network module by means of feature fusion to obtain multiple feature representations with different scales;
[0021] A detection head module, configured to extract the earlobe part from all the feature representations obtained by the neck network module by means of bounding box extraction of an object to obtain an earlobe feature map;
[0022] The target detection device corresponding to the structure in this solution can accurately detect the earlobe part in the image, and can also use a bounding box to identify the required object and its position in the inspection image.
[0023] In a possible implementation, the feature selection device is configured to perform the following steps:
[0024] A1: Perform an inter-group comparison on the measurement data in the clinical text data by means of an independent samples t-test or a Mann-Whitney test, and select features with a statistical difference P-value less than 0.05 to obtain a measurement comparison result;
[0025] A2: Perform an inter-group comparison on the count data in the clinical text data by means of a chi-square test, and select features with a statistical difference P-value less than 0.05 to obtain a count comparison result;
[0026] A3: Eliminate features with a correlation less than a first threshold from the measurement comparison result and the count comparison result by means of 10-fold cross-validation and the least absolute shrinkage and selection operator method to obtain a strongly correlated feature set;
[0027] A4: Rank the importance of all the features in the strongly correlated feature set by means of a random forest algorithm to select features with an importance stronger than a second threshold to obtain the clinical data feature set;
[0028] This solution extracts the measurement data and count data of the clinical text data, selects corresponding statistical analysis methods according to their characteristics, obtains the clinical data feature set, and selects features with high correlation.
[0029] In a possible implementation, the prediction device includes:
[0030] An image feature module, communicating with the enhancement device, is configured to perform image feature extraction on the feature-enhanced image through dense convolution and channel weight adjustment to obtain an image feature vector;
[0031] A text feature module, communicating with the feature selection device, is configured to adjust the clinical data feature set through a self-attention mechanism to obtain a text feature vector;
[0032] A feature fusion module, communicating with both the image feature module and the text feature module, is configured to calculate the importance of the text feature vector to the image feature vector through an attention mechanism, take this importance as the importance of the clinical data feature set to the feature-enhanced image, and then fuse the text feature vector and the image feature vector according to this importance to obtain the fused feature;
[0033] A fully connected layer module, communicating with the feature fusion module, is configured to obtain a cardiovascular lesion prediction result of the patient through a fully connected mapping using the fused feature;
[0034] This solution uses an attention mechanism to dynamically adjust the weights of text and images, can calculate the importance of text features to image features, then combines the two features to generate a fused feature, and finally obtains a cardiovascular lesion prediction result through a fully connected layer mapping, realizing accurate and reliable prediction.
[0035] In a possible implementation, the image feature module includes:
[0036] A dense convolution network, its head communicating with the enhancement device, is configured to perform dense convolution and fully connected layer mapping processing on the feature-enhanced image to obtain a dense feature vector;
[0037] A squeeze-and-excitation unit, communicating with the end of the dense convolution network, is configured to perform channel weight adjustment processing on the dense feature vector to obtain the image feature vector;
[0038] This solution ensures the efficiency and accuracy of image feature extraction by setting a dense convolution network and a squeeze-and-excitation unit.
[0039] In a possible implementation, the dense convolution network is a communication network structure formed by connecting in series a plurality of dense convolution models, a plurality of convolution layer models, a plurality of pooling layer models, and a fully connected layer model, where one of the convolution layer models is arranged at the head of the communication network structure, and the fully connected layer model is arranged at the end of the communication network structure;
[0040] In the communication network structure, multiple dense convolutional models are arranged at intervals, and in the order from the head end to the tail end, a convolutional layer model and a pooling layer model are successively arranged between two adjacent dense convolutional models;
[0041] The dense convolutional model is a chain structure formed by connecting multiple mapping models, and the output end of each mapping model communicates with the input ends of all the subsequent mapping models;
[0042] The mapping model is set to operate according to the following functional formula:
[0043] X(T) = C(S(Z(T))),
[0044] In the formula,
[0045] T represents the input image of the mapping model;
[0046] Z(T) represents the first intermediate result, which is the result of batch normalization processing on the input image of the mapping model;
[0047] S(Z(T)) represents the second intermediate result, which is the result of applying the rectified linear unit activation function to the first intermediate result;
[0048] X(T) represents the output image of the mapping model, which is the result of convolution processing on the second intermediate result;
[0049] The dense convolutional network formed by this solution forms a densely connected chain structure by directly communicating and connecting each mapping model of the dense convolutional model with other mapping models. This means that each mapping layer can receive the feature maps of all the previous mapping layers, promoting feature reuse, enhancing gradient flow, and alleviating the problem of gradient disappearance. At the same time, since each mapping layer receives the feature maps from all the previous mapping layers, the parameter efficiency of the network is greatly improved, and the number of parameters of the model is reduced.
[0050] In a possible implementation manner, the squeeze-and-excitation unit is set to perform the following steps:
[0051] B1: Capture the global features of each channel of the dense feature vector through global average pooling operation to obtain a channel squeeze vector;
[0052] B2: Map the channel squeeze vector through a fully connected layer mapping to generate the weights of each channel to obtain excitation weights;
[0053] B3: Perform a channel-wise multiplication operation on the excitation weights and the dense feature vector to obtain the image feature vector;
[0054] The extrusion-incentive unit of this solution first performs channel compression on the input dense feature vector, that is, captures the global features of each channel of the dense feature vector through global average pooling; then performs non-linear mapping processing through a fully connected layer mapping to generate the weights of each channel; finally, these weights are multiplied with the original dense feature vector channel by channel, so as to enhance the attention to important channels and suppress the responses of unimportant channels. This mechanism can effectively improve the feature expression ability and classification performance of the model without significantly increasing the computational complexity.
[0055] In a possible implementation manner, the text feature module is set to perform the following steps:
[0056] C1: Vectorize the clinical data feature set through a text vectorization processing method to obtain a text data vector;
[0057] C2: Adjust the text data vector through a self-attention mechanism adjustment method to obtain the text feature vector.
[0058] Another technical solution of the present invention is to provide a method for predicting cardiovascular lesions by constructing a deep network model, and the method includes the following steps:
[0059] S1: Collect the clinical text data and bilateral earlobe images of the patient through a collection device;
[0060] S2: Extract the earlobe part from the bilateral earlobe images in the way of extracting objects with bounding boxes through a target detection device to obtain an earlobe feature map;
[0061] S3: Perform affine transformation on the earlobe feature map through an enhancement device to obtain transformed images with different perspectives, scales and shapes, and then perform Gaussian blur processing on the transformed images to obtain a feature-enhanced image;
[0062] S4: Perform statistical between-group difference detection, regression analysis and random forest importance screening detection on the clinical text data in sequence through a feature selection device to obtain a clinical data feature set;
[0063] S5: Obtain the fused features by calculating the importance of the clinical data feature set for the feature-enhanced image through an attention mechanism by a prediction device, and then obtain the cardiovascular lesion prediction result of the patient by using the fused features through a fully connected mapping.
[0064] The method disclosed by the present invention abandons the existing technical route for coronary heart disease screening through facial features and earlobe image features of a human face, adopts the clinical text data of patients and bilateral earlobe images, and combines the technical route of statistical analysis and image processing, avoiding the technical defect of large computational complexity caused by dual image processing. Moreover, since the process of obtaining the clinical data feature set from the clinical text data in step S4 is to perform statistical between-group difference detection, regression analysis, and random forest importance screening detection in sequence, not only can important features of the clinical data be obtained, but also it is easy to perform model layout and the model cost is low. In addition, since in step S5, the prediction device calculates the importance of the clinical data feature set for the feature-enhanced image through the attention mechanism, and obtains the cardiovascular lesion prediction result of the patient by using the fused features through full connection mapping, realizing the combination of text features and image features to obtain the cardiovascular lesion prediction result, the accuracy of the prediction result can be guaranteed, achieving the purpose of early screening of cardiovascular lesions. Description of the Drawings
[0065] Figure 1 It is a schematic structural diagram of a cardiovascular lesion prediction system for constructing a deep network model disclosed in an embodiment of the present application;
[0066] Figure 2 It is a schematic overall structure diagram of the target detection device disclosed in an embodiment of the present application;
[0067] Figure 3 It is a schematic structural diagram of the backbone network module disclosed in an embodiment of the present application;
[0068] Figure 4 It is a schematic structural diagram of the neck network module disclosed in an embodiment of the present application;
[0069] Figure 5 It is a schematic structural diagram of the detection head module disclosed in an embodiment of the present application;
[0070] Figure 6 It is a schematic structural diagram of the first cross-connect unit disclosed in an embodiment of the present application;
[0071] Figure 7 It is a schematic structural diagram of any one of the second cross-connect unit, the third cross-connect unit, the fourth cross-connect unit, the fifth cross-connect unit, and the sixth cross-connect unit disclosed in an embodiment of the present application;
[0072] Figure 8 It is a flowchart of the operation of the feature selection device disclosed in an embodiment of the present application;
[0073] Figure 9 It is a schematic structural diagram of the prediction device disclosed in an embodiment of the present application;
[0074] Figure 10Schematic diagram of the dense convolutional network structure disclosed in the embodiments of the present application;
[0075] Figure 11 Schematic diagram of the operation mode of the dense convolutional model disclosed in the embodiments of the present application;
[0076] Figure 12 Flowchart of the operation of the squeeze-excitation unit disclosed in the embodiments of the present application;
[0077] Figure 13 Flowchart of the method disclosed in the embodiments of the present application;
[0078] Figure 14 Confusion matrix of the coronary heart disease prediction model based on clinical text data in the embodiments of the present application;
[0079] Figure 15 Confusion matrix of the coronary heart disease prediction model based on earlobe images in the embodiments of the present application;
[0080] Figure 16 Confusion matrix of a cardiovascular lesion prediction system constructed by a deep network model disclosed in the embodiments of the present application;
[0081] Figure 17 Receiver operating characteristic curve of the coronary heart disease prediction model based on clinical text data in the embodiments of the present application;
[0082] Figure 18 Receiver operating characteristic curve of the coronary heart disease prediction model based on earlobe images in the embodiments of the present application;
[0083] Figure 19 Receiver operating characteristic curve of a cardiovascular lesion prediction system constructed by a deep network model disclosed in the embodiments of the present application. Detailed implementation manners
[0084] First of all, those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can make adjustments according to needs to adapt to specific application scenarios.
[0085] In the embodiments of the present application, unless otherwise clearly specified and limited, the communication or communication connection between the first feature and the second feature means that there is information transmission between the first feature and the second feature. Such information transmission can be either unidirectional or bidirectional, and the ways to achieve the communication connection can be wire electrical connection, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication realized through channels, etc.
[0086] In the embodiments of the present application, unless otherwise clearly specified or limited, the first feature may be in direct contact with the second feature, or the first and second features may be indirectly in contact through an intermediate medium, when the first feature is "on", "under", "in front of", or "behind" the second feature. Moreover, when the first feature is "above", "over", or "on top of" the second feature, it may be directly above or obliquely above the second feature, or merely indicate that the first feature has a higher horizontal height than the second feature. When the first feature is "under", "beneath", or "underneath" the second feature, it may be directly below or obliquely below the second feature, or merely indicate that the first feature has a lower horizontal height than the second feature. When the first feature is "before", "in front of", or "in the front of" the second feature, it may be directly in front of or obliquely in front of the second feature, or merely indicate that the first feature is before the second feature in sequence. When the first feature is "after", "behind", or "at the back of" the second feature, it may be directly behind or obliquely behind the second feature, or merely indicate that the first feature is after the second feature in sequence.
[0087] The technical solution of the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0088] See Figures 1 to 12 , the embodiments of the present application disclose a cardiovascular lesion prediction system for constructing a deep network model. Figure 1 As shown in Figure 1 the structural schematic diagram of the cardiovascular lesion prediction system, the cardiovascular lesion prediction system includes a collection device, a target detection device, an enhancement device, a feature selection device, and a prediction device. Among them, the target detection device communicates with the collection device, the enhancement device communicates with the target detection device, the feature selection device communicates with the collection device, and the prediction device communicates with both the enhancement device and the feature selection device.
[0089] In this cardiovascular lesion prediction system, the collection device is configured to collect the clinical text data and bilateral earlobe images of patients. See Figure 1 , in this embodiment, the collection device includes a text module and a photographing module. The text module communicates with the feature selection device, and the photographing module communicates with the target detection device. The text module is configured to collect the clinical text data of the patient; the clinical text data includes the demographic characteristics, medical history, risk factors, and laboratory examination information of the patient. The demographic characteristics of the patient include age, gender, body mass index (BMI); the medical history of the patient includes a history of hypertension, hyperlipidemia, diabetes, kidney disease, and cerebrovascular disease; the risk factors of the patient include a smoking history, a drinking history, and serological examination results.
[0090] In the acquisition device, the photographing module is set to acquire the images of the earlobes of both ears of the patient. To improve the quality of the acquired images, a high-definition camera can be used for the photographing module. It should be noted that, to ensure the rationality of the cardiovascular disease prediction results, the following factors that are likely to interfere with earlobe wrinkles need to be excluded: ① patients with local ear diseases; ② patients wearing ear ornaments such as earrings; ③ patients with ear wrinkles formed in a short time due to external factors. In addition, patients with congenital heart disease, heart valve disease, or large blood vessel diseases also need to be excluded.
[0091] See Figures 2 to 5 , in this cardiovascular disease prediction system, the target detection device is set to extract the earlobe part from the images of the earlobes of both ears by means of extracting the object through the bounding box to obtain the earlobe feature map. In this embodiment, the mean squared error loss function is used to optimize the parameters of the target detection device, and precision (Pr), recall rate (Rc), and mean average precision (mAP) are used as the parameter evaluation indicators of the target detection device. Figure 2 For the overall structural schematic diagram of the target detection device, as Figure 2 shown, the target detection device in this embodiment is the fifth version of the target detection network model (YOLO v5), which specifically includes a Mosaic module, a backbone network module, a neck network module, and a detection head module arranged in sequence along the running direction (i.e., the direction from the input to the output of the target detection device).
[0092] In the target detection device, the Mosaic module is set to scale the images of the earlobes of both ears to a predetermined standard size by means of adaptive image scaling to obtain the standard size images. In the fifth version of the target detection network model, the Mosaic module can scale the images of the earlobes of both ears by using random scaling, random cropping, and random arrangement methods, and adopt adaptive anchor box calculation. In the training of the Mosaic module, prediction boxes can be output based on the initial anchor boxes, and then compared with the ground truth boxes to calculate the difference between the two, and then updated in the reverse direction to iterate the model parameters.
[0093] In the target detection device, the backbone network module is set to convert the standard size images into multiple feature maps with different scales by means of the fusion of focus operation, convolution operation, and feature pyramid pooling operation. The backbone network module is a chain-like network structure including a Focus unit, a pyramid pooling unit, multiple first convolutional units, multiple first skip units, and multiple output terminals. The number of the first convolutional units is greater than the number of output terminals. The Focus unit is arranged at the head of the chain-like network structure, and the end of the chain-like network structure is a first convolutional unit. One of the output terminals is arranged on the first convolutional unit at the end of the chain-like network structure, and the remaining output terminals are arranged on the other first convolutional units in a one-to-one manner. The structural schematic diagram of the backbone network module constructed in this embodiment is as Figure 3As shown in the figure, it is composed of a focusing unit, a pyramid pooling unit, four convolutional units I, and three cross-connection units I connected in series, and includes three output terminals, namely output terminal P1, output terminal P2, and output terminal P3. Thus, the backbone network module in this embodiment can convert a standard-size image into three feature maps with different scales.
[0094] In the backbone network module, the focusing unit is set to operate as follows: First, the input image is copied, enlarged, and sliced; then these slices are connected in depth to form a complete picture, and then processed through a convolution operation, and finally, the output result is obtained through a two-dimensional batch normalization (BatchNorm2d, also known as two-dimensional batch normalization. In this application, in order to reduce the characters in the drawings, batch normalization is uniformly written as batch normalization, the same below) operation and a SiLU activation function operation.
[0095] In the backbone network module, the convolutional unit (CBL) I is a convolutional block formed by sequentially connecting a convolutional layer model (Convolution), a batch normalization unit (which is called a batch normalization unit (Batch Normalization) in many documents. In this application, in order to reduce the characters in the drawings, it is uniformly written as a batch normalization unit, the same below), and an activation function model that performs a Leaky ReLU activation function operation.
[0096] In the backbone network module, the cross-connection unit I is a first type of cross-connection structure (Cross Stage PartialConnection, abbreviated as CSP1_x structure). The structural schematic diagram of the cross-connection unit I is as Figure 6 shown, including X residual convolutional blocks, a convolutional layer model, an integration unit (Concat) V, a batch normalization model that performs a two-dimensional batch normalization operation, and an activation unit. Among them, the convolutional layer model communicates with the input terminal of the cross-connection unit I. The X residual convolutional blocks form a chain-like structure through series connection. Each residual convolutional block includes a convolutional unit I, a residual unit (Resunit), and a convolutional layer model arranged in sequence according to the running direction. The convolutional unit I in the first residual convolutional block communicates with the input terminal of the cross-connection unit I. The integration unit V communicates with both the convolutional layer model and the convolutional layer model in the last residual convolutional block. The batch normalization model communicates with the integration unit V. The operation performed by the activation unit is a Leaky ReLU activation function operation and communicates with the batch normalization model. This structure of the cross-connection unit can achieve authenticated fusion of image context and obtain an output result through processing by the activation function.
[0097] See Figure 2 and Figure 4, in the target detection device, the neck network module is configured to perform feature fusion on all the feature maps output by the backbone network module through a feature fusion method to obtain multiple feature representations with different scales. Corresponding to the backbone network module in this embodiment, the neck network module in this embodiment is provided with an output end P4, an output end P5, and an output end P6, and is configured to perform feature fusion on three feature maps output by the backbone network module through a feature fusion method to obtain three feature representations with different scales. The structural schematic diagram of the neck network module in this embodiment is as Figure 4 shown. The neck network module includes a bridging unit two, a bridging unit three, a bridging unit four, a bridging unit five, a bridging unit six, a convolutional unit two, a convolutional unit three, a convolutional unit four, a convolutional unit five, a convolutional unit six, a convolutional unit seven, a convolutional unit eight, a convolutional unit nine, a convolutional unit ten, an upsampling unit one, an upsampling unit two, an integration unit one, an integration unit two, an integration unit three, and an integration unit four. The bridging unit two communicates with the output end P3, the convolutional unit two communicates with the bridging unit two, the upsampling unit one communicates with the convolutional unit two, the convolutional unit three communicates with the output end P2, the integration unit one communicates with the upsampling unit one and the convolutional unit three at the same time, the bridging unit three communicates with the integration unit one, the convolutional unit four communicates with the bridging unit three, the upsampling unit two communicates with the convolutional unit four, the convolutional unit five communicates with the output end P1, the integration unit two communicates with the upsampling unit two and the convolutional unit five at the same time, the bridging unit four communicates with the integration unit two, the convolutional unit six communicates with the bridging unit four, the output end P4 communicates with the convolutional unit six, the convolutional unit seven communicates with the convolutional unit six, the integration unit three communicates with the convolutional unit four and the convolutional unit seven at the same time, the bridging unit five communicates with the integration unit three, the convolutional unit eight communicates with the bridging unit five, the output end P5 communicates with the convolutional unit eight, the convolutional unit nine communicates with the convolutional unit eight, the integration unit four communicates with the convolutional unit two and the convolutional unit nine at the same time, the bridging unit six communicates with the integration unit four, and the convolutional unit ten communicates with the bridging unit six, and the output end P6 communicates with the convolutional unit ten.
[0098] In the neck network module of this embodiment, the bridging unit two, the bridging unit three, the bridging unit four, the bridging unit five, and the bridging unit six are all second-type cross-connection structures (abbreviated as CSP2_x), and their structural schematic diagrams are as Figure 7As shown, the spanning connection structure includes X ordinary convolution blocks, a convolution layer model, an integration unit five, a batch normalization model, and an activation unit. Among them, the convolution layer model communicates with the input end of the corresponding bridging unit. The X ordinary convolution blocks are connected in series to form a chain-like structure. Each ordinary convolution block includes two convolution units one and a convolution layer model arranged in sequence according to the running direction. The convolution unit one at the front end of the first ordinary convolution block communicates with the input end of the bridging unit one. The integration unit five communicates with the convolution layer model and the convolution layer model in the last ordinary convolution block at the same time. The batch normalization model communicates with the integration unit five. The operation performed by the activation unit is the Leaky ReLU activation function operation and it communicates with the batch normalization model. The bridging unit with this structure can achieve image context fusion and obtain an output result through processing by the activation function.
[0099] In the neck network module of this embodiment, the structures of convolution unit two, convolution unit three, convolution unit four, convolution unit five, convolution unit six, convolution unit seven, convolution unit eight, convolution unit nine, and convolution unit ten are all the same as that of convolution unit one.
[0100] See Figure 5 , in the object detection device, the detection head module is set to extract the earlobe part from all the feature representations obtained from the neck network module by means of extracting the object through the bounding box to obtain the earlobe feature map. The detection head module in this embodiment includes three detection heads, which communicate with the output ends P4, P5, and P6 respectively, so as to be able to extract the earlobe part from the three feature representations obtained from the neck network module in this embodiment to obtain the earlobe feature map.
[0101] In this cardiovascular disease prediction system, the enhancement device is set to perform an affine transformation on the earlobe feature map to obtain transformed images with different perspectives, scales, and shapes, and then perform Gaussian blur processing on the transformed images to obtain feature-enhanced images. Affine transformation refers to performing translation, rotation, scaling, reflection, and shear operations on the image to generate images with different perspectives, scales, and shapes, so as to achieve the effect of increasing image diversity. Gaussian blur processing can reduce the sensitivity of the enhancement device to subtle changes and noise in the input image, thereby reducing the risk of overfitting, making the detection area of the enhancement device more accurate, ultimately improving the generalization ability of the model, reducing the risk of overfitting, and improving the model adaptation ability of the enhancement device.
[0102] See Figure 1 and Figure 8 , in this cardiovascular disease prediction system, the feature selection device is set to perform statistical inter-group difference detection, regression analysis, and importance ranking of the random forest algorithm on the clinical text data in sequence to obtain the clinical data feature set.
[0103] SeeFigure 8 , in this embodiment, the feature selection device is configured to perform the following steps:
[0104] A1: Perform between-group comparison on the measurement data in the clinical text data by independent sample t-test or Mann-Whitney test, and select features with a statistical difference P-value less than 0.05 to obtain the measurement comparison result; the measurement data in the clinical text data are the data with units in the clinical text data.
[0105] A2: Perform between-group comparison on the count data in the clinical text data by chi-square test, and select features with a statistical difference P-value less than 0.05 to obtain the count comparison result; the count data in the clinical text data are the data without units in the clinical text data.
[0106] A3: Eliminate features with a correlation less than the first threshold from the measurement comparison result and the count comparison result through 10-fold cross-validation and the Least Absolute Shrinkage and Selection Operator (LASSO) method to obtain a strongly correlated feature set; the first threshold can be set according to a self-defined criterion.
[0107] A4: Rank the importance of all features in the strongly correlated feature set through the random forest algorithm to select features with importance stronger than the second threshold to obtain the clinical data feature set; the first threshold can be set according to a self-defined criterion.
[0108] See Figure 1 , Figure 9 , Figure 10 and Figure 11 , in this cardiovascular disease prediction system, the prediction device is configured to calculate the importance of the clinical data feature set for the feature-enhanced image through the attention mechanism, obtain the features after the fusion of the clinical data feature set and the feature-enhanced image based on this importance, and then obtain the cardiovascular disease prediction result of the patient through full connection mapping using the fused features.
[0109] See Figure 9 , in this embodiment, the prediction device includes an image feature module, a text feature module, a feature fusion module, and a full connection layer module. Among them, the image feature module communicates with the enhancement device, the text feature module communicates with the feature selection device, the feature fusion module communicates with both the image feature module and the text feature module, and the full connection layer module communicates with the feature fusion module.
[0110] See Figure 9 and Figure 10, in the prediction device, the image feature module is configured to extract image features from the feature-enhanced image through dense convolution and channel weight adjustment to obtain an image feature vector. The image feature module includes a dense convolution network and a squeeze-and-excitation unit (abbreviated as SE). The head end of the dense convolution network communicates with the enhancement device, the squeeze-and-excitation unit communicates with the end of the dense convolution network, and the squeeze-and-excitation unit also communicates with the feature fusion module. The dense convolution network is configured to perform dense convolution and fully connected layer mapping processing on the feature-enhanced image to obtain a dense feature vector; the squeeze-and-excitation unit is configured to perform channel weight adjustment processing on the dense feature vector to obtain an image feature vector.
[0111] See Figure 10 , in the image feature module, the dense convolution network is a communication network structure formed by connecting a plurality of dense convolution models, a plurality of convolutional layer models, a plurality of pooling layer models, and a fully connected layer model in series. One convolutional layer model is arranged at the head end of the communication network structure, and the fully connected layer model is arranged at the end of the communication network structure. In the communication network structure, the plurality of dense convolution models are arranged at intervals, and in the order from the head end to the end, a convolutional layer model and a pooling layer model are sequentially arranged between two adjacent dense convolution models. Figure 10 shows a dense convolution network with three dense convolution models. The dense convolution model is a chain structure formed by connecting a plurality of mapping models, and the output end of each mapping model communicates with the input ends of all subsequent mapping models; the mapping model is configured to operate according to the following functional formula:
[0112] X(T) = C(S(Z(T))),
[0113] In the formula,
[0114] T represents the input image of the mapping model;
[0115] Z(T) represents the first intermediate result, which is the result of batch normalization processing on the input image of the mapping model;
[0116] S(Z(T)) represents the second intermediate result, which is the result of linear rectified activation function processing on the first intermediate result;
[0117] X(T) represents the output image of the mapping model, which is the result of convolution processing on the second intermediate result.
[0118] Figure 11 shows the operation rules of multiple mapping models in the dense convolution model, and the red arrow represents the skip connection.
[0119] See Figure 12, in the image feature module, the squeeze-and-excitation unit is configured to perform the following steps:
[0120] B1: Capture the global features of each channel of the dense feature vector through global average pooling operation to obtain the channel squeeze vector.
[0121] B2: Map the channel squeeze vector through a fully connected layer mapping to generate the weights of each channel, and obtain the excitation weights.
[0122] B3: Perform a channel-wise multiplication operation on the excitation weights and the dense feature vector to obtain the image feature vector.
[0123] In the prediction device, the text feature module is configured to adjust the clinical data feature set through a self-attention mechanism adjustment method to obtain the text feature vector. In this embodiment, the text feature module is configured to perform the following steps:
[0124] C1: Vectorize the clinical data feature set through a text vectorization processing method to obtain the text data vector.
[0125] C2: Adjust the text data vector through a self-attention mechanism adjustment method to obtain the text feature vector. The process of adjusting the text data vector through the self-attention mechanism adjustment method includes initializing weights, calculating attention scores, scaling the attention scores, activating with the softmax function, and weighted merging value vectors. This adjustment method is a prior art and will not be elaborated here. Those skilled in the art can refer to relevant books on machine learning to obtain the corresponding processing process.
[0126] In the prediction device, the feature fusion module is configured to calculate the importance of the text feature vector to the image feature vector through an attention mechanism, take this importance as the importance of the clinical data feature set to the feature-enhanced image, and then fuse the text feature vector and the image feature vector according to this importance to obtain the fused feature. In this embodiment, the method of fusing the text feature vector and the image feature vector according to the importance is: perform weighted averaging on the text feature vector and the image feature vector according to the value of the importance to obtain the fused feature. The fully connected layer module is configured to obtain the cardiovascular lesion prediction result of the patient through a fully connected mapping using the fused feature. The prediction result is presented in the form of the probability of suffering from coronary heart disease in this embodiment.
[0127] For the optimization of the model parameters of the prediction device, in this embodiment, the mean squared error loss function is adopted, and the accuracy, sensitivity, specificity, receiver operating characteristic curve (ROC), and the area under the receiver operating characteristic curve (AUC) are used as the model evaluation indicators of the prediction device to evaluate its generalization ability, and the results are visualized through a confusion matrix.
[0128] Next, the cardiovascular disease prediction method corresponding to the cardiovascular disease prediction system will be further disclosed. Figure 13 As shown in the flowchart of this method, the method includes the following steps:
[0129] S1: Collect the clinical text data and bilateral earlobe images of the patient through the collection device.
[0130] S2: Extract the earlobe part from the bilateral earlobe images in the way of extracting objects with bounding boxes through the target detection device to obtain an earlobe feature map.
[0131] S3: Perform affine transformation on the earlobe feature map through the enhancement device to obtain transformed images with different perspectives, scales, and shapes, and then perform Gaussian blur processing on the transformed images to obtain a feature-enhanced image.
[0132] S4: Sequentially perform statistical between-group difference detection, regression analysis, and random forest screening importance detection on the clinical text data through the feature selection device to obtain a set of clinical data features.
[0133] S5: Obtain the fused features by calculating the importance of the set of clinical data features for the feature-enhanced image through the attention mechanism of the prediction device, and then obtain the cardiovascular disease prediction result of the patient by using the fused features through full connection mapping. This prediction result is presented as the probability of having coronary heart disease.
[0134] Next, the technical effects of the cardiovascular disease prediction system will be further described. Figure 14 This is the confusion matrix of the coronary heart disease prediction model based on clinical text data (abbreviated as Model A). Figure 15 This is the confusion matrix of the coronary heart disease prediction model based on earlobe images (abbreviated as Model B). Figure 16 This is the confusion matrix of this cardiovascular disease prediction system. Figure 17 This is the receiver operating characteristic curve of Model A. Figure 18 This is the receiver operating characteristic curve of Model B. Figure 19It is the working characteristic curve of the cardiovascular lesion prediction system. Through the detection of the test set, for Model A, the AUC = 0.708, accuracy = 68.15%, sensitivity = 65.82%, and specificity = 70.51. For Model B, the AUC = 0.730, accuracy = 69.43%, sensitivity = 72.61%, and specificity = 66.24%. For the cardiovascular lesion prediction system (specifically calculating the prediction device), the AUC = 0.740, accuracy = 72.93%, sensitivity = 74.52%, and specificity = 71.34%. The results show that the comprehensive performance of the cardiovascular lesion prediction system is better than that of Model A and Model B.
[0135] The cardiovascular lesion prediction system constructed by the deep network model disclosed in this embodiment sets up an acquisition device to collect the clinical text data and bilateral earlobe images of patients. For the bilateral earlobe images, the target detection device extracts the earlobe parts in the bilateral earlobe images in the way of extracting objects with bounding boxes to obtain earlobe feature maps. Subsequently, the enhancement device performs affine transformation and Gaussian blur processing on the earlobe feature maps to obtain feature-enhanced images, and the feature-enhanced images are received by the set prediction device. For the clinical text data, the feature selection device performs statistical inter-group difference detection, regression analysis, and random forest screening importance detection on it in sequence to obtain a clinical data feature set, and the clinical data feature set is received by the set prediction device. Finally, the prediction device calculates the importance of the clinical data feature set for the feature-enhanced images through the attention mechanism, and obtains the fused features. Then, the full connection mapping is used to obtain the cardiovascular lesion prediction results of the patients by using the fused features. It can be seen that the cardiovascular lesion prediction system disclosed in the present invention abandons the existing technical route of screening for coronary heart disease through facial features and earlobe image features, adopts the clinical text data and bilateral earlobe images of patients, and combines the technical routes of statistical analysis and image processing, avoiding the technical defect of large computational complexity caused by dual image processing. Moreover, since the process of obtaining the clinical data feature set from the clinical text data is to perform statistical inter-group difference detection, regression analysis, and random forest screening importance detection in sequence, not only can the important features of the clinical data be obtained, but also the model layout is easy and the model cost is low. In addition, since the prediction device calculates the importance of the clinical data feature set for the feature-enhanced images through the attention mechanism, and obtains the cardiovascular lesion prediction results of the patients by using the fused features through full connection mapping, it realizes the combination of text features and image features to obtain the cardiovascular lesion prediction results, which can ensure the accuracy of the prediction results and achieve the purpose of early screening for cardiovascular lesions.
[0136] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms indicating directions or positional relationships such as "inside" and "outside" are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application.
[0137] In the description of the present application, the description with reference to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0138] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A cardiovascular lesion prediction system for constructing a deep network model, characterized in that, Comprising: A collection device, configured to collect clinical text data and bilateral earlobe images of a patient; A target detection device, communicating with the collection device, configured to extract the earlobe part from the bilateral earlobe images by means of bounding box extraction of an object to obtain an earlobe feature map; An enhancement device, communicating with the target detection device, configured to perform an affine transformation on the earlobe feature map to obtain transformed images with different perspectives, scales, and shapes, and then perform Gaussian blur processing on the transformed images to obtain a feature-enhanced image; A feature selection device, communicating with the collection device, configured to perform statistical inter-group difference detection, regression analysis, and importance ranking of a random forest algorithm on the clinical text data in sequence to obtain a clinical data feature set; A prediction device, communicating with the enhancement device and the feature selection device simultaneously, configured to calculate the importance of the clinical data feature set for the feature-enhanced image through an attention mechanism, obtain the features after fusion of the clinical data feature set and the feature-enhanced image based on this importance, and then obtain the cardiovascular lesion prediction result of the patient by means of full connection mapping using the fused features; The prediction device includes: An image feature module, communicating with the enhancement device, configured to perform image feature extraction on the feature-enhanced image by means of dense convolution and channel weight adjustment to obtain an image feature vector; A text feature module, communicating with the feature selection device, configured to adjust the clinical data feature set by means of a self-attention mechanism adjustment to obtain a text feature vector; A feature fusion module, communicating with the image feature module and the text feature module simultaneously, configured to calculate the importance of the text feature vector for the image feature vector through an attention mechanism, use this importance as the importance of the clinical data feature set for the feature-enhanced image, and then perform feature fusion of the text feature vector and the image feature vector according to this importance to obtain the fused features; A fully connected layer module, communicating with the feature fusion module, configured to obtain the cardiovascular lesion prediction result of the patient by means of full connection mapping using the fused features; The image feature module includes: A dense convolution network, whose head communicates with the enhancement device, configured to perform dense convolution and full connection layer mapping processing on the feature-enhanced image to obtain a dense feature vector; A squeeze-and-excitation unit, communicating with the end of the dense convolution network, configured to perform channel weight adjustment processing on the dense feature vector to obtain the image feature vector; The dense convolution network is a communication network structure formed by connecting in series a plurality of dense convolution models, a plurality of convolution layer models, a plurality of pooling layer models, and a full connection layer model, wherein one of the convolution layer models is arranged at the head of the communication network structure, and the full connection layer model is arranged at the end of the communication network structure; In the communication network structure, the multiple dense convolutional models are arranged at intervals, and in the order from the head end to the tail end, a convolutional layer model and a pooling layer model are successively arranged between two adjacent dense convolutional models; The dense convolutional model is a chain-like structure formed by connecting multiple mapping models, and the output end of each mapping model communicates with the input ends of all the subsequent mapping models.
2. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 1, wherein, The acquisition device includes: A text module, communicating with the feature selection device, and configured to acquire clinical text data of a patient; the clinical text data includes demographic characteristics, medical history, risk factors, and laboratory examination information of the patient; A photographing module, communicating with the target detection device, and configured to acquire images of the earlobes of both ears of the patient.
3. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 2, characterized in that, The target detection device includes, successively arranged along the running direction: A mosaic module, configured to scale the earlobe images of both ears to a predetermined standard size by means of adaptive image scaling to obtain a standard size image; A backbone network module, configured to convert the standard size image into multiple feature maps with different scales by a fusion method of a focusing operation, a convolutional operation, and a feature pyramid pooling operation; A neck network module, configured to perform feature fusion on all the feature maps output by the backbone network module by a feature fusion method to obtain multiple feature representations with different scales; A detection head module, configured to extract the earlobe part from all the feature representations obtained by the neck network module by a method of extracting objects with bounding boxes to obtain an earlobe feature map.
4. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 2 or 3, characterized in that The feature selection device is configured to perform the following steps: A1: Perform an inter-group comparison on the measurement data in the clinical text data by an independent samples t-test method or a Mann-Whitney test method, and select features with a statistical difference P-value less than 0.05 to obtain a measurement comparison result; A2: Perform an inter-group comparison on the count data in the clinical text data by a chi-square test method, and select features with a statistical difference P-value less than 0.05 to obtain a count comparison result; A3: Eliminate features with a correlation less than a first threshold from the measurement comparison result and the count comparison result by a 10-fold cross-validation and least absolute shrinkage and selection operator method to obtain a strongly correlated feature set; A4: Rank the importance of all the features in the strongly correlated feature set by a random forest algorithm to select features with an importance stronger than a second threshold to obtain the clinical data feature set.
5. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 4, wherein, The mapping model is configured to operate according to the following functional formula: X(T)=C(S(Z(T))), wherein, T represents the input image of the mapping model; Z(T) represents a first intermediate result, which is the result of batch normalization processing on the input image of the mapping model; S(Z(T)) represents a second intermediate result, which is the result of linear rectifier activation function processing on the first intermediate result; X(T) represents the output image of the mapping model, which is the result of convolutional processing on the second intermediate result.
6. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 5, characterized in that, The squeeze-and-excitation unit is configured to perform the following steps: B1: Capture the global features of each channel of the dense feature vector through global average pooling operation to obtain a channel squeeze vector; B2: Perform mapping processing on the channel squeeze vector through a fully connected layer mapping to generate the weights of each channel, obtaining excitation weights; B3: Perform per-channel multiplication operation on the excitation weights and the dense feature vector to obtain the image feature vector.
7. The cardiovascular lesion prediction system constructed based on the deep network model according to claim 6, wherein, The text feature module is set to perform the following steps: C1: Vectorize the clinical data feature set through a text vectorization processing method to obtain a text data vector; C2: Adjust the text data vector through a self-attention mechanism adjustment method to obtain the text feature vector.
8. A method for predicting cardiovascular lesions by constructing a deep network model, characterized in that, A cardiovascular lesion prediction system applicable to the construction of the deep network model according to any one of claims 1-7, comprising the following steps: S1: Collect the clinical text data and bilateral earlobe images of the patient through a collection device; S2: Extract the earlobe part from the bilateral earlobe images in the manner of extracting objects with bounding boxes through a target detection device to obtain an earlobe feature map; S3: Perform affine transformation on the earlobe feature map through an enhancement device to obtain transformed images with different perspectives, scales and shapes, and then perform Gaussian blur processing on the transformed images to obtain feature enhanced images; S4: Perform statistical between-group difference detection, regression analysis and random forest screening importance detection on the clinical text data in sequence through a feature selection device to obtain a clinical data feature set; S5: Calculate the importance of the clinical data feature set for the feature enhanced image through an attention mechanism by a prediction device, and obtain the fused features, and then obtain the cardiovascular lesion prediction result of the patient by using the fused features through a fully connected mapping.
Citation Information
Patent Citations
Multi-mode-based intelligent auxiliary prediction and diagnosis platform for diabetes and complications thereof
CN116386860A
Semantic navigation and lesion mapping from digital breast tomosynthesis
US20140348404A1