Cardiovascular lesion prediction system and method constructed by deep network model

The cardiovascular lesion prediction system constructed through deep network models combines clinical text data and earlobe images, and uses attention mechanism and full connection mapping to solve the problems of high model cost and high computational complexity in the existing technology, achieving fast and convenient cardiovascular lesion evaluation and accurate prediction results.

CN119943414AActive Publication Date: 2025-05-06NINGBO UNIV
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510428591.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When the prior art deeply explores the wrinkle characteristics of earlobes to predict coronary heart disease, there are technical defects of high model cost and high computational complexity, resulting in cardiovascular lesions not being fast and convenient enough.

Method used

A cardiovascular lesion prediction system built using a deep network model. This system obtains the patient's clinical text data and binaural earlobe images through the acquisition device. Combined with object detection, enhancement, feature selection and prediction devices, it uses attention mechanism and full connection mapping to achieve the fusion of text features and image features and obtain the cardiovascular lesion prediction results.

Benefits of technology

It reduces the cost of the model, avoids the high computational complexity caused by dual image processing, realizes fast and convenient cardiovascular lesions evaluation, and ensures the accuracy of prediction results, achieving the purpose of early screening of cardiovascular lesions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943414A_ABST
    Figure CN119943414A_ABST
Patent Text Reader

Abstract

The invention relates to a cardiovascular lesion prediction system and method constructed by a deep network model, and the system is provided with an acquisition device which is used for collecting clinical text data and binaural earlobe images of a patient. The method comprises the following steps: extracting an earlobe part in a binaural earlobe image by a set target detection device in a bounding box object extraction manner to obtain an earlobe feature map, and then carrying out affine transformation and Gaussian blur processing on the earlobe feature map by a set enhancement device to obtain a feature enhanced image. As for clinical text data, a feature selection device is arranged to sequentially carry out statistical inter-group difference detection, regression analysis and random forest screening importance detection on the clinical text data so as to obtain a clinical data feature set. And finally, obtaining a cardiovascular lesion prediction result of the patient through a prediction device according to fusion of the feature enhanced image and the clinical data feature set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and artificial intelligence technology, and in particular to a cardiovascular disease prediction system and method based on a deep network model. Background Art

[0002] The impact of cardiovascular disease on human health is self-evident. Coronary atherosclerotic heart disease (CAD), also known as coronary heart disease, is the most representative of its impact. Since the grading of earlobe wrinkles is significantly positively correlated with the severity of coronary heart disease, and earlobe images are a common body surface feature that is easy to obtain and non-invasive, screening for coronary heart disease based on earlobe wrinkles has potential application prospects. Traditional technical solutions usually judge earlobe wrinkles based on experience, which is subjective and lacks standard classification definitions and objective quantitative basis.

[0003] However, the existing technology discloses a technical solution of "artificial intelligence + imaging omics" to screen for coronary heart disease through facial features and earlobe image features, and to evaluate the cardiovascular lesions triggered by coronary heart disease. This technical solution applies artificial intelligence technology to the processing, analysis and interpretation of medical imaging data, analyzes the imaging data of a large sample of cases, deeply mines the disease-specific imaging features, and predicts the results on this basis.

[0004] However, this "artificial intelligence + imaging omics" technical solution has the technical defect of high model cost and high computational complexity in the process of deeply mining earlobe wrinkle characteristics for the prediction of coronary heart disease due to the need for double image processing, which makes it impossible to quickly and conveniently evaluate cardiovascular lesions. Summary of the invention

[0005] The technical problem to be solved by the present invention is how to overcome the technical defects in the prior art, that is, the high model cost and large computational complexity in the process of deeply mining earlobe wrinkle features for predicting coronary heart disease. In order to overcome the above defects of the prior art, the present invention provides a cardiovascular lesion prediction system and method constructed by a deep network model, comprising a cardiovascular lesion prediction system constructed by a deep network model and a cardiovascular lesion prediction method constructed by a deep network model.

[0006] The present invention provides a cardiovascular disease prediction system constructed by a deep network model, comprising: A collection device configured to collect clinical text data and binaural earlobe images of the patient; an object detection device, in communication with the acquisition device, configured to extract the earlobe parts from the earlobe images of both ears by extracting objects through a bounding box to obtain an earlobe feature map; an enhancement device, in communication with the target detection device, configured to perform an affine transformation on the earlobe feature map to obtain a transformed image with different perspectives, scales and shapes, and then perform a Gaussian blur process on the transformed image to obtain a feature-enhanced image; A feature selection device, in communication with the acquisition device, is configured to sequentially perform statistical inter-group difference detection, regression analysis, and importance ranking of a random forest algorithm on the clinical text data to obtain a clinical data feature set; The prediction device communicates with the enhancement device and the feature selection device at the same time, and is configured to calculate the importance of the clinical data feature set to the feature-enhanced image through an attention mechanism, obtain the fused features of the clinical data feature set and the feature-enhanced image based on the importance, and then obtain the patient's cardiovascular disease prediction result using the fused features through a fully connected mapping.

[0007] The cardiovascular lesion prediction system constructed by the deep network model disclosed in the present invention aims at the technical defects listed above. By setting a collection device to collect the clinical text data and binaural earlobe images of patients, the target detection device set extracts the earlobe parts in the binaural earlobe images in the way of extracting objects by a bounding box, and obtains the earlobe feature map. Then, the earlobe feature map is subjected to affine transformation and Gaussian blur processing by the enhancement device set to obtain a feature enhanced image, which is received by the prediction device set. For clinical text data, a feature selection device is set to perform statistical inter-group difference detection, regression analysis and random forest screening importance detection on it in turn to obtain a clinical data feature set, which is received by the prediction device set. Finally, the prediction device calculates the importance of the clinical data feature set to the feature enhanced image by an attention mechanism, and obtains the fused features, and then uses the fused features by a fully connected mapping to obtain the patient's cardiovascular lesion prediction results. It can be seen that the cardiovascular lesion prediction system disclosed in the present invention abandons the existing technical route of screening for coronary heart disease through facial features and earlobe image features, adopts the patient's clinical text data and binaural earlobe images, and combines the technical route of statistical analysis and image processing, avoiding the technical defect of large computational complexity caused by double image processing. In addition, since the process of obtaining the clinical data feature set through clinical text data is to perform statistical inter-group difference detection, regression analysis and random forest screening importance detection in sequence, not only can the important features of the clinical data be obtained, but also the model layout is easy, and the model cost is low. In addition, since the prediction device performs the calculation of the importance of the clinical data feature set to the feature enhanced image through the attention mechanism, and obtains the patient's cardiovascular lesion prediction result by using the fused features through the fully connected mapping, it realizes the combination of text features and image features to obtain the cardiovascular lesion prediction result, which can ensure the accuracy of the prediction result and achieve the purpose of early screening of cardiovascular lesions.

[0008] In a possible implementation, the collection device includes: a text module, in communication with the feature selection device, configured to collect clinical text data of the patient; the clinical text data including the patient's demographic characteristics, medical history, risk factors and laboratory test information; a photographing module, communicating with the target detection device and configured to capture images of both earlobes of the patient; The above scheme can realize the parallel acquisition of the patient's clinical text data and bilateral earlobe images, ensuring the orderliness of the separate processing of text data and images.

[0009] In a possible implementation, the target detection device includes: A mosaic module is configured to scale the binaural earlobe image to a predetermined standard size by adaptive image scaling to obtain a standard size image; A backbone network module is configured to convert the standard size image into a plurality of feature maps with different scales by integrating a focusing operation, a convolution operation and a feature pyramid pooling operation; The neck network module is configured to perform feature fusion on all feature maps output by the backbone network module to obtain multiple feature representations with different scales; A detection head module is configured to extract the earlobe part from all feature representations obtained by the neck network module by extracting objects through a bounding box to obtain an earlobe feature map; The target detection device corresponding to the structure in this scheme can realize accurate detection of the earlobe part in the image, and can also use the bounding box to identify the required object and its position in the inspection image.

[0010] In a possible implementation, the feature selection device is configured to perform the following steps: A1: Perform inter-group comparison on the measurement data in the clinical text data by using the independent sample t test or the Mann-Whitney test, and select features with a statistically significant difference P value less than 0.05 to obtain measurement comparison results; A2: Perform inter-group comparison on the count data in the clinical text data by using the chi-square test, and select features with a statistically significant difference P value less than 0.05 to obtain the count comparison results; A3: Eliminate features whose correlation is less than a first threshold from the measurement comparison result and the counting comparison result through 10-fold cross validation and minimum absolute shrinkage and selection operator method to obtain a set of strongly correlated features; A4: sorting the importance of all features in the strongly correlated feature set by a random forest algorithm to select features whose importance is greater than a second threshold, thereby obtaining the clinical data feature set; This solution extracts measurement data and count data from clinical text data, selects corresponding statistical analysis methods based on their characteristics, obtains a set of clinical data features, and selects features with high correlation.

[0011] In a possible implementation, the prediction device includes: An image feature module, communicating with the enhancement device, is configured to extract image features from the feature-enhanced image by dense convolution and channel weight adjustment to obtain an image feature vector; a text feature module, communicating with the feature selection device, and configured to adjust the clinical data feature set by a self-attention mechanism to obtain a text feature vector; a feature fusion module, communicating with the image feature module and the text feature module at the same time, configured to calculate the importance of the text feature vector to the image feature vector through an attention mechanism, taking the importance as the importance of the clinical data feature set to the feature-enhanced image, and then performing feature fusion of the text feature vector and the image feature vector according to the importance to obtain the fused feature; a fully connected layer module, communicating with the feature fusion module, and configured to obtain a cardiovascular disease prediction result of the patient using the fused features through fully connected mapping; This solution uses the attention mechanism to dynamically adjust the weights of text and images. It can calculate the importance of text features to image features, and then combine the two features to generate fused features. Finally, it obtains the cardiovascular disease prediction results through a fully connected layer mapping, achieving accurate and reliable prediction.

[0012] In a possible implementation, the image feature module includes: A dense convolutional network, whose head end is in communication with the enhancement device, is configured to perform dense convolution and fully connected layer mapping processing on the feature enhanced image to obtain a dense feature vector; a squeeze-excitation unit, in communication with an end of the dense convolutional network, configured to perform a channel weight adjustment process on the dense feature vector to obtain the image feature vector; This scheme ensures the efficiency and accuracy of image feature extraction by setting up a dense convolutional network and a squeeze-excitation unit.

[0013] In a possible implementation, the dense convolutional network is a communication network structure formed by connecting a plurality of dense convolutional models, a plurality of convolutional layer models, a plurality of pooling layer models and a fully connected layer model in series, wherein one of the convolutional layer models is arranged at the head end of the communication network structure, and the fully connected layer model is arranged at the end of the communication network structure; In the communication network structure, the plurality of dense convolution models are arranged at intervals, and in order from the head end to the end end, one convolution layer model and one pooling layer model are arranged in sequence between two adjacent dense convolution models; The dense convolution model is a chain structure formed by connecting multiple mapping models, and the output end of each mapping model communicates with the input ends of all subsequent mapping models; The mapping model is set up to run as follows: X(T)=C(S(Z(T))), In the formula, T represents the input image of the mapping model; Z(T) represents a first intermediate result, which is a result of performing batch normalization processing on the input image of the mapping model; S(Z(T)) represents the second intermediate result, which is the result of performing linear rectification activation function processing on the first intermediate result; X(T) represents the output image of the mapping model, which is the result of convolution processing on the second intermediate result; The dense convolutional network formed by this scheme forms a densely connected chain structure by allowing each mapping model of the dense convolutional model to communicate directly with other mapping models. This means that each mapping layer can receive the feature maps of all the previous mapping layers, which promotes feature reuse, enhances gradient flow, and alleviates the gradient disappearance problem. At the same time, since each mapping layer receives feature maps from all previous mapping layers, the parameter efficiency of the network is greatly improved and the number of model parameters is reduced.

[0014] In a possible implementation, the squeeze-excitation unit is configured to perform the following steps: B1: Capture the global features of each channel of the dense feature vector through a global average pooling operation to obtain a channel squeeze vector; B2: Mapping the channel squeezing vector through a fully connected layer mapping to generate a weight for each channel to obtain an excitation weight; B3: performing channel-by-channel multiplication operation on the excitation weight and the dense feature vector to obtain the image feature vector; The squeeze-excitation unit of this scheme first performs channel compression on the input dense feature vector, that is, the global features of each channel of the dense feature vector are captured through global average pooling; then a fully connected layer is used for nonlinear mapping to generate the weight of each channel; finally, these weights are multiplied channel by channel with the original dense feature vector, thereby enhancing the focus on important channels and suppressing the response of unimportant channels. This mechanism can effectively improve the feature expression ability and classification performance of the model without significantly increasing the amount of calculation.

[0015] In a possible implementation, the text feature module is configured to perform the following steps: C1: vectorizing the clinical data feature set by text vectorization processing to obtain a text data vector; C2: The text data vector is adjusted by a self-attention mechanism to obtain the text feature vector.

[0016] Another technical solution of the present invention is to provide a method for predicting cardiovascular lesions using a deep network model, the method comprising the following steps: S1: Collecting clinical text data and bilateral earlobe images of patients through a collection device; S2: extracting the earlobe parts from the binaural earlobe images by using a target detection device in a manner of extracting objects by a bounding box to obtain an earlobe feature map; S3: performing affine transformation on the earlobe feature map by an enhancement device to obtain a transformed image with different viewing angles, scales and shapes, and then performing Gaussian blur processing on the transformed image to obtain a feature enhanced image; S4: performing statistical inter-group difference detection, regression analysis and random forest screening importance detection on the clinical text data in sequence by a feature selection device to obtain a clinical data feature set; S5: The prediction device calculates the importance of the clinical data feature set to the feature enhanced image by an attention mechanism to obtain fused features, and then uses the fused features through a fully connected mapping to obtain the patient's cardiovascular disease prediction results.

[0017] The method disclosed by the present invention abandons the existing technical route of screening for coronary heart disease through facial features and earlobe image features of human faces, adopts the clinical text data and binaural earlobe images of patients, and combines the technical route of statistical analysis and image processing, thereby avoiding the technical defect of large computational complexity caused by double image processing. In addition, since the process of obtaining the clinical data feature set through clinical text data in step S4 is to sequentially perform statistical inter-group difference detection, regression analysis and random forest screening importance detection, not only can the important features of clinical data be obtained, but also the model layout is easy, and the model cost is low. In addition, since the prediction device in step S5 calculates the importance of the clinical data feature set to the feature enhanced image through the attention mechanism, and the fused features are used to obtain the patient's cardiovascular lesion prediction results through full connection mapping, the text features are combined with the image features to obtain the cardiovascular lesion prediction results, which can ensure the accuracy of the prediction results and achieve the purpose of early screening of cardiovascular lesions. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of the structure of a cardiovascular disease prediction system constructed by a deep network model disclosed in an embodiment of the present application; Figure 2 Schematic diagram of the overall structure of the target detection device disclosed in the embodiments of the present application; Figure 3 This is a schematic diagram of the backbone network module structure disclosed in the embodiments of the present application; Figure 4 This is a schematic diagram of the structure of the neck network module disclosed in the embodiments of the present application; Figure 5 This is a schematic diagram of the structure of the detection head module disclosed in the embodiments of the present application; Figure 6 It is a structural schematic diagram of a jumper unit disclosed in an embodiment of the present application; Figure 7 It is a structural schematic diagram of any one of the jumper unit 2, jumper unit 3, jumper unit 4, jumper unit 5 and jumper unit 6 disclosed in the embodiments of the present application; Figure 8 It is an operation flow chart of the feature selection device disclosed in the embodiment of the present application; Fig. 9 It is a schematic diagram of the structure of the prediction device disclosed in the embodiments of the present application; Fig.10 Schematic diagram of the dense convolutional network structure disclosed in the embodiments of the present application; Fig.11 A schematic diagram of the operation mode of the dense convolution model disclosed in the embodiments of the present application; Fig.12 It is an operation flow chart of the squeezing-excitation unit disclosed in the embodiment of the present application; Fig.13 It is a flow chart of the method disclosed in the embodiments of this application; Fig.14 The confusion matrix of the coronary heart disease prediction model based on clinical text data in the embodiment of the present application; Fig.15 is a confusion matrix of a coronary heart disease prediction model based on earlobe images in an embodiment of the present application; Fig.16 A confusion matrix of a cardiovascular disease prediction system constructed by a deep network model disclosed in an embodiment of the present application; Fig.17 is the operating characteristic curve of the coronary heart disease prediction model based on clinical text data in the embodiment of the present application; Fig.18 is the operating characteristic curve of the coronary heart disease prediction model based on earlobe images in the embodiment of the present application; Fig.19 This is a working characteristic curve of a cardiovascular disease prediction system constructed using a deep network model disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0019] First, those skilled in the art should understand that these implementations are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can make adjustments to them as needed to adapt to specific application scenarios.

[0020] In the embodiments of the present application, unless otherwise clearly specified and limited, the communication or communication connection between the first feature and the second feature refers to the transmission of information between the first feature and the second feature. This information transmission can be either unidirectional or bidirectional, and the communication connection can be achieved by electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc.

[0021] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "above", "below", "in front of" or "behind" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "above" and "above" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. The first feature being "below", "below" and "below" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature. The first feature being "before", "in front of" and "in front of" the second feature may mean that the first feature is directly in front of or obliquely in front of the second feature, or simply means that the first feature is before the second feature in order. The first feature being "after", "behind" and "behind" the second feature may mean that the first feature is directly behind or obliquely behind the second feature, or simply means that the first feature is after the second feature in order.

[0022] The technical solution of the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0023] See also Figures 1 to 12 The present application embodiment discloses a cardiovascular disease prediction system constructed by a deep network model. Figure 1 This is a schematic diagram of the cardiovascular disease prediction system structure. Figure 1 As shown, the cardiovascular disease prediction system includes an acquisition device, a target detection device, an enhancement device, a feature selection device and a prediction device, wherein the target detection device communicates with the acquisition device, the enhancement device communicates with the target detection device, the feature selection device communicates with the acquisition device, and the prediction device communicates with the enhancement device and the feature selection device at the same time.

[0024] In the cardiovascular disease prediction system, the collecting device is configured to collect clinical text data and bilateral earlobe images of the patient. Figure 1In this embodiment, the acquisition device includes a text module and a camera module, the text module communicates with the feature selection device, and the camera module communicates with the target detection device. The text module is configured to collect clinical text data of the patient; the clinical text data includes the patient's demographic characteristics, medical history, risk factors, and laboratory test information. The patient's demographic characteristics include age, gender, and body mass index (Body Mass Index, BMI); the patient's medical history includes a history of hypertension, a history of hyperlipidemia, a history of diabetes, a history of kidney disease, and a history of cerebrovascular disease; the patient's risk factors include a history of smoking, a history of drinking, and serological test results.

[0025] In the acquisition device, the camera module is configured to acquire images of both earlobes of the patient. In order to improve the quality of the obtained images, the camera module may use a high-definition camera. It should be noted that in order to ensure the rationality of the prediction results of cardiovascular lesions, the following factors that may interfere with earlobe wrinkles need to be excluded: ① patients with local ear lesions; ② patients wearing ear ornaments such as earrings; ③ patients whose ear wrinkles are formed in a short period of time due to external factors. In addition, patients with congenital heart disease, valvular heart disease or large vessel disease need to be excluded.

[0026] See also Figure 2 to Figure 5 In the cardiovascular disease prediction system, the target detection device is configured to extract the earlobe part from the earlobe image of both ears by extracting the object through a bounding box to obtain the earlobe feature map. This embodiment uses the mean square error loss function to optimize the parameters of the target detection device, and uses precision (Pr), recall (Rc) and mean average precision (mAP) as parameter evaluation indicators of the target detection device. Figure 2 is a schematic diagram of the overall structure of the target detection device, such as Figure 2 As shown, the target detection device in this embodiment is the fifth version of the target detection network model (YOLO v5), which specifically includes a mosaic module, a backbone network module, a neck network module and a detection head module which are sequentially arranged along the running direction (i.e., the direction from the input to the output of the target detection device).

[0027] In the target detection device, the mosaic module is set to scale the binaural earlobe image to a predetermined standard size by adaptive image scaling to obtain a standard size image. In the fifth version of the target detection network model, the mosaic module can scale the binaural earlobe image by random scaling, random cropping, and random arrangement, and use adaptive anchor frame calculation. In the training of the mosaic module, the predicted frame can be output based on the initial anchor frame, and then compared with the real frame, the difference between the two is calculated, and then the model parameters are updated in reverse to iterate.

[0028] In the target detection device, the backbone network module is configured to convert a standard size image into multiple feature maps of different scales by integrating focusing operation, convolution operation and feature pyramid pooling operation. The backbone network module is a chain network structure including a focusing unit (Focus), a pyramid pooling unit, multiple convolution units one, multiple cross-connection units one and multiple output terminals. The number of convolution units one is greater than the number of output terminals. The focusing unit is arranged at the head end of the chain network structure, and the end of the chain network structure is a convolution unit one, one of which is arranged on the convolution unit one at the end of the chain network structure, and the remaining output terminals are arranged on other convolution units one in a one-to-one manner. The schematic diagram of the backbone network module structure constructed in this embodiment is shown in FIG. Figure 3 As shown, it is composed of a focusing unit, a pyramid pooling unit, four convolution units 1 and three cross-connection units 1 connected in series, and includes three output terminals, namely output terminal P1, output terminal P2 and output terminal P3. Thus, the backbone network module in this embodiment can convert a standard size image into three feature maps with different scales.

[0029] In the backbone network module, the focusing unit is set to operate as follows: first, the input image is copied, enlarged and sliced; then, these slices are connected in depth to form a complete picture, followed by processing through convolution operations, and finally, the output result is obtained through two-dimensional batch normalization (BacthNorm2d, also known as two-dimensional batch normalization, in order to reduce the size of characters in the drawings, this application uniformly writes batch normalization as batch normalization, the same below) operations and SiLU activation function operations.

[0030] In the backbone network module, the convolution unit (CBL) is a convolution block composed of a convolution layer model (Convolution), a batch normalization unit (called a batch normalization unit in many documents, and in order to reduce the size of characters in the drawings, this application uniformly writes a batch normalization unit, the same below)) and an activation function model that performs the Leaky ReLU activation function operation, which are connected in sequence.

[0031] In the backbone network module, the cross-connection unit 1 is a first-class cross-connection structure (Cross Stage PartialConnection, referred to as CSP1_x structure). The structural diagram of the cross-connection unit 1 is as follows: Figure 6As shown, it includes X residual convolution blocks, a convolution layer model, an integration unit (Concat) 5, a batch normalization model that performs a two-dimensional batch normalization operation, and an activation unit, wherein the convolution layer model communicates with the input end of the jump unit 1, and the X residual convolution blocks are connected in series to form a chain structure, and each residual convolution block includes a convolution unit 1, a residual unit (Resunit) and a convolution layer model arranged in sequence according to the running direction, and the convolution unit 1 in the first residual convolution block communicates with the input end of the jump unit 1, the integration unit 5 communicates with the convolution layer model and the convolution layer model in the last residual convolution block at the same time, the batch normalization model communicates with the integration unit 5, and the activation unit performs the operation of the Leaky ReLU activation function and communicates with the batch normalization model. The jump unit of this structure can realize the authenticated fusion of the image context, and is processed by the activation function to obtain the output result.

[0032] See also Figure 2 and Figure 4 In the target detection device, the neck network module is configured to perform feature fusion on all feature maps output by the backbone network module by means of feature fusion to obtain multiple feature representations with different scales. Corresponding to the backbone network module in this embodiment, the neck network module in this embodiment is provided with output terminals P4, P5 and P6, which are configured to perform feature fusion on the three feature maps output by the backbone network module by means of feature fusion to obtain three feature representations with different scales. The structural schematic diagram of the neck network module in this embodiment is shown in FIG. Figure 4As shown, the neck network module includes jumper unit 2, jumper unit 3, jumper unit 4, jumper unit 5, jumper unit 6, convolution unit 2, convolution unit 3, convolution unit 4, convolution unit 5, convolution unit 6, convolution unit 7, convolution unit 8, convolution unit 9, convolution unit 10, upsampling unit 1, upsampling unit 2, integration unit 1, integration unit 2, integration unit 3 and integration unit 4. Jumper unit 2 communicates with output terminal P3, convolution unit 2 communicates with jumper unit 2, upsampling unit 1 communicates with convolution unit 2, convolution unit 3 communicates with output terminal P2, integration unit 1 communicates with upsampling unit 1 and convolution unit 3 at the same time, jumper unit 3 communicates with integration unit 1, convolution unit 4 communicates with jumper unit 3, upsampling unit Element 2 communicates with convolution unit 4, convolution unit 5 communicates with output terminal P1, integration unit 2 communicates with upsampling unit 2 and convolution unit 5 at the same time, jumper unit 4 communicates with integration unit 2, convolution unit 6 communicates with jumper unit 4, output terminal P4 communicates with convolution unit 6, convolution unit 7 communicates with convolution unit 6, integration unit 3 communicates with convolution unit 4 and convolution unit 7 at the same time, jumper unit 5 communicates with integration unit 3, convolution unit 8 communicates with jumper unit 5, output terminal P5 communicates with convolution unit 8, convolution unit 9 communicates with convolution unit 8, integration unit 4 communicates with convolution unit 2 and convolution unit 9 at the same time, jumper unit 6 communicates with integration unit 4, convolution unit 10 communicates with jumper unit 6, and output terminal P6 communicates with convolution unit 10.

[0033] In the neck network module of this embodiment, the jumper unit 2, the jumper unit 3, the jumper unit 4, the jumper unit 5, and the jumper unit 6 are all second-type crossover connection structures (abbreviated as CSP2_x), and the structural diagram thereof is shown in FIG. Figure 7 As shown, the cross-connection structure includes X common convolution blocks, a convolution layer model, an integration unit five, a batch normalization model and an activation unit, wherein the convolution layer model communicates with the input end of the cross-connection unit, and the X common convolution blocks are connected in series to form a chain structure, and each common convolution block includes two convolution units one and a convolution layer model arranged in sequence according to the running direction, and the convolution unit one at the front end of the common convolution block at the first position communicates with the input end of the cross-connection unit one, the integration unit five communicates with the convolution layer model and the convolution layer model in the common convolution block at the end at the same time, the batch normalization model communicates with the integration unit five, and the activation unit performs the operation of the Leaky ReLU activation function and communicates with the batch normalization model. The cross-connection unit of this structure can realize image context fusion, and is processed by the activation function to obtain the output result.

[0034] In the neck network module of this embodiment, the structures of convolution unit two, convolution unit three, convolution unit four, convolution unit five, convolution unit six, convolution unit seven, convolution unit eight, convolution unit nine and convolution unit ten are the same as convolution unit one.

[0035] See also Figure 5 In the target detection device, the detection head module is configured to extract the earlobe part from all feature representations obtained by the neck network module by extracting the object through a bounding box to obtain an earlobe feature map. The detection head module in this embodiment includes three detection heads, which communicate with output terminals P4, P5, and P6 respectively, so that the earlobe part can be extracted from the three feature representations obtained by the neck network module in this embodiment to obtain the earlobe feature map.

[0036] In the cardiovascular lesion prediction system, the enhancement device is configured to perform an affine transformation on the earlobe feature map to obtain a transformed image with different perspectives, scales and shapes, and then perform Gaussian blur processing on the transformed image to obtain a feature-enhanced image. Affine transformation refers to the translation, rotation, scaling, reflection and shearing operations on an image to generate images with different perspectives, scales and shapes, thereby increasing the diversity of the image. Gaussian blur processing can reduce the sensitivity of the enhancement device to subtle changes and noise in the input image, thereby reducing the risk of overfitting, making the detection area of ​​the enhancement device more accurate, and ultimately improving the generalization ability of the model, reducing the risk of overfitting, and improving the model adaptability of the enhancement device.

[0037] See also Figure 1 and Figure 8 ,In the cardiovascular lesion prediction system, the feature selection device is configured to sequentially perform statistical ,intergroup difference detection, regression analysis and importance ranking of the ,random forest algorithm on the clinical text data to obtain the ,clinical data feature set.

[0038] See also Figure 8 In this embodiment, the feature selection device is configured to perform the following steps: A1: Perform inter-group comparison on the quantitative data in clinical text data by independent sample t-test or Mann-Whitney test, and select features with statistical difference P value less than 0.05 to obtain quantitative comparison results; quantitative data in clinical text data refers to data with units in clinical text data.

[0039] A2: The count data in the clinical text data are compared between groups using the chi-square test, and features with statistical differences P values ​​less than 0.05 are selected to obtain the count comparison results; the count data in the clinical text data are the data without units in the clinical text data.

[0040] A3: Through 10-fold cross validation and the least absolute shrinkage and selection operator (LASSO), the features with correlation less than the first threshold are eliminated from the measurement comparison results and the counting comparison results to obtain a set of strongly correlated features; the first threshold can be set according to a custom standard.

[0041] A4: All features in the strongly correlated feature set are ranked by importance using a random forest algorithm to select features whose importance is greater than a second threshold, thereby obtaining the clinical data feature set; the first threshold can be set according to a custom standard.

[0042] See also Figure 1 , Fig. 9 , Fig.10 and Fig.11 In this cardiovascular lesion prediction system, the prediction device is configured to calculate the importance of the clinical data feature set to the feature enhanced image through an attention mechanism, obtain the fused features of the clinical data feature set and the feature enhanced image based on this importance, and then obtain the patient's cardiovascular lesion prediction result using the fused features through a fully connected mapping.

[0043] See also Fig. 9 In this embodiment, the prediction device includes an image feature module, a text feature module, a feature fusion module and a fully connected layer module, wherein the image feature module communicates with the enhancement device, the text feature module communicates with the feature selection device, the feature fusion module communicates with the image feature module and the text feature module at the same time, and the fully connected layer module communicates with the feature fusion module.

[0044] See also Fig. 9 and Fig.10 In the prediction device, the image feature module is configured to extract image features from the feature enhanced image by dense convolution and channel weight adjustment to obtain an image feature vector. The image feature module includes a dense convolution network and a squeeze-and-excitation unit (SE), the head end of the dense convolution network communicates with the enhancement device, the squeeze-and-excitation unit communicates with the end of the dense convolution network, and the squeeze-and-excitation unit also communicates with the feature fusion module. The dense convolution network is configured to perform dense convolution and full connection layer mapping processing on the feature enhanced image to obtain a dense feature vector; the squeeze-and-excitation unit is configured to perform channel weight adjustment processing on the dense feature vector to obtain an image feature vector.

[0045] See also Fig.10In the image feature module, the dense convolution network is a communication network structure formed by connecting multiple dense convolution models, multiple convolution layer models, multiple pooling layer models and a fully connected layer model in series, wherein one convolution layer model is set at the head end of the communication network structure, and the fully connected layer model is set at the end of the communication network structure. In the communication network structure, multiple dense convolution models are arranged at intervals, and a convolution layer model and a pooling layer model are arranged between two adjacent dense convolution models in order from the head end to the end end. Fig.10 A dense convolutional network with three dense convolutional models is shown in . The dense convolutional model is a chain structure formed by connecting multiple mapping models, and the output of each mapping model communicates with the input of all subsequent mapping models; the mapping model is set to run as follows: X(T)=C(S(Z(T))), In the formula, T represents the input image of the mapping model; Z(T) represents the first intermediate result, which is the result of batch normalization processing on the input image of the mapping model; S(Z(T)) represents the second intermediate result, which is the result of performing linear rectification activation function processing on the first intermediate result; X(T) represents the output image of the mapping model, which is the result of convolution processing on the second intermediate result.

[0046] Fig.11 The operating rules of multiple mapping models in a dense convolutional model are shown, and the red arrows represent the spanning output.

[0047] See also Fig.12 , in the image feature module, the squeeze-excitation unit is configured to perform the following steps: B1: The global features of each channel of the dense feature vector are captured through the global average pooling operation to obtain the channel squeeze vector.

[0048] B2: The channel squeeze vector is mapped through the fully connected layer mapping to generate the weight of each channel and obtain the excitation weight.

[0049] B3: Multiply the excitation weights by the dense feature vector channel by channel to obtain the image feature vector.

[0050] In the prediction device, the text feature module is configured to adjust the clinical data feature set by a self-attention mechanism to obtain a text feature vector. In this embodiment, the text feature module is configured to perform the following steps: C1: Vectorize the clinical data feature set through text vectorization processing to obtain text data vectors.

[0051] C2: The text data vector is adjusted by the self-attention mechanism to obtain the text feature vector. The process of adjusting the text data vector by the self-attention mechanism includes initializing weights, calculating attention scores, scaling attention beams, activating the softmax function, and weighted merging value vectors. This adjustment method is a prior art and will not be expanded here. Those skilled in the art can refer to relevant books on machine learning to obtain the corresponding processing process.

[0052] In the prediction device, the feature fusion module is configured to calculate the importance of the text feature vector to the image feature vector through the attention mechanism, and use this importance as the importance of the clinical data feature set to the feature enhanced image, and then perform feature fusion of the text feature vector and the image feature vector according to this importance to obtain the fused feature. In this embodiment, the method of performing feature fusion of the text feature vector and the image feature vector according to the importance is: weighted averaging the text feature vector and the image feature vector according to the value of the importance to obtain the fused feature. The fully connected layer module is configured to obtain the patient's cardiovascular lesion prediction result using the fused feature through a fully connected mapping. The prediction result is presented in the form of the probability of suffering from coronary heart disease in this embodiment.

[0053] For the model parameter optimization of the prediction device, the mean square error loss function is used in this embodiment, and accuracy, sensitivity, specificity, operating characteristic curve (ROC) and area under the operating characteristic curve (AUC) are used as model evaluation indicators of the prediction device to evaluate its generalization ability, and the results are visualized through the confusion matrix.

[0054] The following will further disclose the cardiovascular lesion prediction method corresponding to the cardiovascular lesion prediction system. Fig.13 is a flow chart of the method, the method comprising the following steps: S1: Collect clinical text data and binaural earlobe images of the patient through a collection device.

[0055] S2: Extracting the earlobe parts from the earlobe images of both ears by using a target detection device in a manner of extracting objects by a bounding box to obtain an earlobe feature map.

[0056] S3: performing affine transformation on the earlobe feature map through an enhancement device to obtain a transformed image with different perspectives, scales and shapes, and then performing Gaussian blur processing on the transformed image to obtain a feature enhanced image.

[0057] S4: The clinical text data is subjected to statistical inter-group difference detection, regression analysis and random forest screening importance detection in sequence through the feature selection device to obtain the clinical data feature set.

[0058] S5: The prediction device calculates the importance of the clinical data feature set to the feature-enhanced image using an attention mechanism to obtain fused features, and then uses the fused features through a fully connected mapping to obtain the patient's cardiovascular disease prediction results. The prediction results are presented as the probability of suffering from coronary heart disease.

[0059] The technical effects of the cardiovascular disease prediction system are further explained below. Fig.14 is the confusion matrix of the coronary heart disease prediction model based on clinical text data (referred to as model A), Fig.15 is the confusion matrix of the coronary heart disease prediction model based on earlobe images (referred to as model B), Fig.16 This is the confusion matrix of the cardiovascular disease prediction system; Fig.17 is the operating characteristic curve of model A, Fig.18 is the operating characteristic curve of model B, Fig.19 The working characteristic curve of the cardiovascular lesion prediction system. Through the test set detection, the AUC of model A is 0.708, the accuracy is 68.15%, the sensitivity is 65.82%, and the specificity is 70.51. The AUC of model B is 0.730, the accuracy is 69.43%, the sensitivity is 72.61%, and the specificity is 66.24%. The AUC of the cardiovascular lesion prediction system (specifically calculated is the prediction device) is 0.740, the accuracy is 72.93%, the sensitivity is 74.52%, and the specificity is 71.34%. The results show that the comprehensive performance of the cardiovascular lesion prediction system is better than that of model A and model B.

[0060] The cardiovascular lesion prediction system constructed by the deep network model disclosed in this embodiment is configured to collect the patient's clinical text data and binaural earlobe images by setting a collection device. For binaural earlobe images, the target detection device is configured to extract the earlobe part in the binaural earlobe image in the manner of extracting objects by a bounding box, and obtain an earlobe feature map. Then, the earlobe feature map is subjected to affine transformation and Gaussian blur processing by the enhancement device, and a feature-enhanced image is obtained. The feature-enhanced image is received by the prediction device configured. For clinical text data, a feature selection device is configured to perform statistical inter-group difference detection, regression analysis, and random forest screening importance detection in sequence to obtain a clinical data feature set, and the clinical data feature set is received by the prediction device configured. Finally, the prediction device calculates the importance of the clinical data feature set to the feature-enhanced image by an attention mechanism, and obtains the fused features, and then uses the fused features by a fully connected mapping to obtain the patient's cardiovascular lesion prediction result. It can be seen that the cardiovascular lesion prediction system disclosed in the present invention abandons the existing technical route of screening for coronary heart disease through facial features and earlobe image features, adopts the patient's clinical text data and binaural earlobe images, and combines the technical route of statistical analysis and image processing, avoiding the technical defect of large computational complexity caused by double image processing. In addition, since the process of obtaining the clinical data feature set through clinical text data is to perform statistical inter-group difference detection, regression analysis and random forest screening importance detection in sequence, not only can the important features of the clinical data be obtained, but also the model layout is easy, and the model cost is low. In addition, since the prediction device performs the calculation of the importance of the clinical data feature set to the feature enhanced image through the attention mechanism, and obtains the patient's cardiovascular lesion prediction result by using the fused features through the fully connected mapping, it realizes the combination of text features and image features to obtain the cardiovascular lesion prediction result, which can ensure the accuracy of the prediction result and achieve the purpose of early screening of cardiovascular lesions.

[0061] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description, and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0062] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" etc. means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0063] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A cardiovascular disease prediction system constructed by a deep network model, characterized in that: include: A collection device configured to collect clinical text data and binaural earlobe images of the patient; an object detection device, in communication with the acquisition device, configured to extract the earlobe parts from the earlobe images of both ears by extracting objects through a bounding box to obtain an earlobe feature map; an enhancement device, in communication with the target detection device, configured to perform an affine transformation on the earlobe feature map to obtain a transformed image with different perspectives, scales and shapes, and then perform a Gaussian blur process on the transformed image to obtain a feature-enhanced image; A feature selection device, in communication with the acquisition device, is configured to sequentially perform statistical inter-group difference detection, regression analysis, and importance ranking of a random forest algorithm on the clinical text data to obtain a clinical data feature set; The prediction device communicates with the enhancement device and the feature selection device at the same time, and is configured to calculate the importance of the clinical data feature set to the feature-enhanced image through an attention mechanism, obtain the fused features of the clinical data feature set and the feature-enhanced image based on the importance, and then obtain the patient's cardiovascular disease prediction result using the fused features through a fully connected mapping.

2. The cardiovascular disease prediction system constructed by the deep network model according to claim 1, characterized in that: The collection device comprises: a text module, in communication with the feature selection device, configured to collect clinical text data of the patient; the clinical text data including the patient's demographic characteristics, medical history, risk factors and laboratory test information; The photographing module communicates with the target detection device and is configured to collect images of both earlobes of the patient.

3. The cardiovascular disease prediction system constructed by the deep network model according to claim 2, characterized in that: The target detection device includes: A mosaic module is configured to scale the binaural earlobe image to a predetermined standard size by adaptive image scaling to obtain a standard size image; A backbone network module is configured to convert the standard size image into a plurality of feature maps with different scales by integrating a focusing operation, a convolution operation and a feature pyramid pooling operation; The neck network module is configured to perform feature fusion on all feature maps output by the backbone network module to obtain multiple feature representations with different scales; The detection head module is configured to extract the earlobe part from all feature representations obtained by the neck network module by extracting objects through a bounding box to obtain an earlobe feature map.

4. The cardiovascular disease prediction system constructed by the deep network model according to claim 2 or 3, characterized in that: The feature selection device is configured to perform the following steps: A1: Perform inter-group comparison on the measurement data in the clinical text data by using the independent sample t test or the Mann-Whitney test, and select features with a statistically significant difference P value less than 0.05 to obtain measurement comparison results; A2: Perform inter-group comparison on the count data in the clinical text data by using the chi-square test, and select features with a statistically significant difference P value less than 0.05 to obtain the count comparison results; A3: Eliminate features whose correlation is less than a first threshold from the measurement comparison result and the counting comparison result through 10-fold cross validation and minimum absolute shrinkage and selection operator method to obtain a set of strongly correlated features; A4: All features in the strongly correlated feature set are ranked by importance using a random forest algorithm to select features whose importance is greater than a second threshold, thereby obtaining the clinical data feature set.

5. The cardiovascular disease prediction system constructed by the deep network model according to claim 4, characterized in that: The prediction device comprises: An image feature module, communicating with the enhancement device, is configured to extract image features from the feature-enhanced image by dense convolution and channel weight adjustment to obtain an image feature vector; a text feature module, communicating with the feature selection device, and configured to adjust the clinical data feature set by a self-attention mechanism to obtain a text feature vector; a feature fusion module, communicating with the image feature module and the text feature module at the same time, configured to calculate the importance of the text feature vector to the image feature vector through an attention mechanism, taking the importance as the importance of the clinical data feature set to the feature-enhanced image, and then performing feature fusion of the text feature vector and the image feature vector according to the importance to obtain the fused feature; The fully connected layer module communicates with the feature fusion module and is configured to obtain the patient's cardiovascular disease prediction result by using the fused features through fully connected mapping.

6. The cardiovascular disease prediction system constructed by the deep network model according to claim 5, characterized in that: The image feature module comprises: A dense convolutional network, whose head end is in communication with the enhancement device, is configured to perform dense convolution and fully connected layer mapping processing on the feature enhanced image to obtain a dense feature vector; A squeeze-excitation unit, in communication with an end of the dense convolutional network, is configured to perform a channel weight adjustment process on the dense feature vector to obtain the image feature vector.

7. The cardiovascular disease prediction system constructed by the deep network model according to claim 6, characterized in that: The dense convolution network is a communication network structure formed by connecting a plurality of dense convolution models, a plurality of convolution layer models, a plurality of pooling layer models and a fully connected layer model in series, wherein one of the convolution layer models is arranged at the head end of the communication network structure, and the fully connected layer model is arranged at the end of the communication network structure; In the communication network structure, the plurality of dense convolution models are arranged at intervals, and in order from the head end to the end end, one convolution layer model and one pooling layer model are sequentially arranged between two adjacent dense convolution models; The dense convolution model is a chain structure formed by connecting multiple mapping models, and the output end of each mapping model communicates with the input ends of all subsequent mapping models; The mapping model is set up to run as follows: X(T)=C(S(Z(T))), In the formula, T represents the input image of the mapping model; Z(T) represents a first intermediate result, which is a result of performing batch normalization processing on the input image of the mapping model; S(Z(T)) represents the second intermediate result, which is the result of performing linear rectification activation function processing on the first intermediate result; X(T) represents the output image of the mapping model, which is the result of convolution processing on the second intermediate result.

8. The cardiovascular disease prediction system constructed by the deep network model according to claim 6 or 7, characterized in that: The squeeze-excitation unit is configured to perform the following steps: B1: Capture the global features of each channel of the dense feature vector through a global average pooling operation to obtain a channel squeeze vector; B2: Mapping the channel squeezing vector through a fully connected layer mapping to generate a weight for each channel to obtain an excitation weight; B3: performing channel-by-channel multiplication operation on the excitation weight and the dense feature vector to obtain the image feature vector.

9. The cardiovascular disease prediction system constructed by the deep network model according to claim 8, characterized in that: The text feature module is configured to perform the following steps: C1: vectorizing the clinical data feature set by text vectorization processing to obtain a text data vector; C2: The text data vector is adjusted by a self-attention mechanism to obtain the text feature vector.

10. A method for predicting cardiovascular disease using a deep network model, characterized in that: A cardiovascular disease prediction system suitable for constructing a deep network model according to any one of claims 1 to 9, comprising the following steps: S1: Collecting clinical text data and bilateral earlobe images of patients through a collection device; S2: extracting the earlobe parts from the earlobe images of both ears by using a target detection device in a manner of extracting objects by a bounding box to obtain an earlobe feature map; S3: performing affine transformation on the earlobe feature map by an enhancement device to obtain a transformed image with different viewing angles, scales and shapes, and then performing Gaussian blur processing on the transformed image to obtain a feature enhanced image; S4: performing statistical inter-group difference detection, regression analysis and random forest screening importance detection on the clinical text data in sequence by a feature selection device to obtain a clinical data feature set; S5: The prediction device calculates the importance of the clinical data feature set to the feature enhanced image by an attention mechanism to obtain fused features, and then uses the fused features to obtain the patient's cardiovascular disease prediction results through fully connected mapping.

Citation Information

Patent Citations

  • Heterogeneous Feature Fusion Based Risk Prediction Method, Model and System for Coronary Heart Disease

    CN109117864A

  • Pathological feature and clinical information fusion method and system

    CN116344070A

  • Multi-mode-based intelligent auxiliary prediction and diagnosis platform for diabetes and complications thereof

    CN116386860A

  • Coronary heart disease diagnosis system based on digital twinborn technology

    CN116631603A

  • Risk prediction method based on multi-modal data fusion

    CN117708746A