Multimodal AI-Driven Cardiovascular Image Detection Method, Device and Equipment

Through the multimodal AI-driven cardiovascular image detection method, combined with multiple cardiovascular images and multimodal AI models, the problem of incomplete information detection is solved, and multi-dimensional evaluation and more comprehensive detection results are achieved.

CN119672016BActive Publication Date: 2025-06-13ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510182820.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

When detecting a single-modal cardiovascular image in the prior art, the information is not comprehensive and the evaluation results cannot be obtained from multi-dimensional and more complete image data structures.

Method used

The cardiovascular image detection method based on multimodal AI is used to obtain an image set composed of multiple cardiovascular images (such as ultrasound, X-ray, CT, MRI), and determine the target multimodal AI model set, and perform multimodal image detection and result fusion.

Benefits of technology

It achieves more comprehensive cardiovascular image detection results from multiple dimensions, providing more complete evaluation information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672016B_ABST
    Figure CN119672016B_ABST
Patent Text Reader

Abstract

The present invention discloses a cardiovascular image detection method, device and equipment based on multi-modal AI drive. First, an image set of cardiovascular images to be detected is obtained and decrypted and updated. Based on the composition ratio of image types in the image set of cardiovascular images to be detected and the multi-modal AI model determination strategy, a set of target multi-modal AI models is determined. For each cardiovascular image to be detected, a target multi-modal AI model is obtained from the set of target multi-modal AI models, and the image detection result corresponding to the cardiovascular image to be detected is determined and the results are fused to obtain a fused detection result. Through the above method, after determining the set of target multi-modal AI models based on the composition ratio of image types for the image set of cardiovascular images to be detected including multiple types of cardiovascular images to be detected, multi-modal image detection and result fusion are correspondingly performed, and a fused detection result that can reflect multi-dimensional evaluation results and has more comprehensive corresponding information is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent decision-making in artificial intelligence, and particularly to a cardiovascular image detection method, device, and equipment driven by multi-modal AI. Background Art

[0002] Currently, after acquiring a user's cardiovascular image, an AI model of a single model (such as a convolutional neural network, etc.) is generally used to perform image detection on the cardiovascular image. The human cardiovascular system is a complex and interconnected whole. Single-modal images such as cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, cardiovascular MRI images, etc. often only provide information about a certain aspect of the cardiovascular system. More specifically, for example, cardiovascular X-ray images mainly present the general outline of the heart and the morphology of some large blood vessels. It can be seen that when performing image detection on a single-modal image using an AI model of a single model to obtain a detection result corresponding to the cardiovascular system as the object to be evaluated, the corresponding information is incomplete, and it is impossible to obtain a more comprehensive detection result with more dimensions of evaluation results and more complete image data structure from more dimensions. Summary of the Invention

[0003] Embodiments of the present invention provide a cardiovascular image detection method, device, and equipment driven by multi-modal AI, aiming to solve the problem that when performing image detection on a single-modal image using an AI model of a single model to obtain a detection result corresponding to the cardiovascular system as the object to be evaluated in the existing technical methods, the corresponding information is incomplete, and it is impossible to obtain a more comprehensive detection result with more dimensions of evaluation results and more complete image data structure from more dimensions.

[0004] In a first aspect, embodiments of the present invention provide a cardiovascular image detection method driven by multi-modal AI, which includes:

[0005] In response to an image detection instruction, obtain a set of cardiovascular images to be detected corresponding to the image detection instruction; wherein, the set of cardiovascular images to be detected includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images;

[0006] If it is determined that the set of cardiovascular images to be detected is an encrypted image set, generate target decryption information based on the target user information corresponding to the image detection instruction, and decrypt the set of cardiovascular images to be detected based on the target decryption information to update the set of cardiovascular images to be detected;

[0007] Determine a set of target multi-modal AI models corresponding to the set of cardiovascular images to be detected based on the composition ratio of the image types in the set of cardiovascular images to be detected and a preset multi-modal AI model determination strategy;

[0008] For each cardiovascular image to be detected in the cardiovascular image set to be detected, obtain a target multi-modal AI model corresponding to the cardiovascular image to be detected from the target multi-modal AI model set, and input the cardiovascular image to be detected into the corresponding target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected;

[0009] Based on a preset detection result fusion model and the image detection results corresponding to each cardiovascular image to be detected in the cardiovascular image set to be detected, obtain a fusion detection result.

[0010] In a second aspect, an embodiment of the present invention provides a cardiovascular image detection device driven by multi-modal AI, which includes:

[0011] A cardiovascular image set to be detected acquisition unit, configured to obtain a cardiovascular image set to be detected corresponding to the image detection instruction in response to the image detection instruction; wherein, the cardiovascular image set to be detected includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images;

[0012] An image set decryption unit, configured to, if it is determined that the cardiovascular image set to be detected is an encrypted image set, generate target decryption information based on the target user information corresponding to the image detection instruction, and decrypt the cardiovascular image set to be detected based on the target decryption information to update the cardiovascular image set to be detected;

[0013] A target multi-modal AI model set acquisition unit, configured to determine a target multi-modal AI model set corresponding to the cardiovascular image set to be detected based on the image type composition ratio in the cardiovascular image set to be detected and a preset multi-modal AI model determination strategy;

[0014] An image detection result acquisition unit, configured to, for each cardiovascular image to be detected in the cardiovascular image set to be detected, obtain a target multi-modal AI model corresponding to the cardiovascular image to be detected from the target multi-modal AI model set, and input the cardiovascular image to be detected into the corresponding target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected;

[0015] A detection result fusion unit, configured to obtain a fusion detection result based on a preset detection result fusion model and the image detection results corresponding to each cardiovascular image to be detected in the cardiovascular image set to be detected.

[0016] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer device executes the computer program, it implements the multi-modal AI-driven cardiovascular image detection method as described in the first aspect above.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the multi-modal AI-driven cardiovascular image detection method as described in the first aspect above.

[0018] The embodiment of the present invention provides a multi-modal AI-driven cardiovascular image detection method, device, and equipment. The method includes: in response to an image detection instruction, obtaining a set of cardiovascular images to be detected corresponding to the image detection instruction; where the set of cardiovascular images to be detected includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images; if it is determined that the set of cardiovascular images to be detected is an encrypted image set, generating target decryption information based on the target user information corresponding to the image detection instruction, and decrypting the set of cardiovascular images to be detected based on the target decryption information to update the set of cardiovascular images to be detected; determining a target multi-modal AI model set corresponding to the set of cardiovascular images to be detected based on the composition ratio of the image types in the set of cardiovascular images to be detected and a preset multi-modal AI model determination strategy; for each cardiovascular image to be detected in the set of cardiovascular images to be detected, obtaining a target multi-modal AI model corresponding to the cardiovascular image to be detected in the target multi-modal AI model set, and inputting the cardiovascular image to be detected into the corresponding target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected; obtaining a fused detection result based on a preset detection result fusion model and the image detection results corresponding to each cardiovascular image to be detected in the set of cardiovascular images to be detected. Through the above method, after determining the target multi-modal AI model set based on the composition ratio of the image types for the set of cardiovascular images to be detected including multiple types of cardiovascular images to be detected, multi-modal image detection and detection result fusion can be performed accordingly, and a fused detection result that can reflect multi-dimensional evaluation results and has more comprehensive corresponding information can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1Schematic diagram of the application scenario of the cardiovascular image detection method driven by multimodal AI provided by the embodiments of the present invention;

[0021] Figure 2 Schematic flowchart of the cardiovascular image detection method driven by multimodal AI provided by the embodiments of the present invention;

[0022] Figure 3 Schematic sub - flowchart of the cardiovascular image detection method driven by multimodal AI provided by the embodiments of the present invention;

[0023] Figure 4 Schematic sub - flowchart of the cardiovascular image detection method driven by multimodal AI provided by the embodiments of the present invention;

[0024] Figure 5 Schematic sub - flowchart of the cardiovascular image detection method driven by multimodal AI provided by the embodiments of the present invention;

[0025] Figure 6 Schematic block diagram of the cardiovascular image detection device driven by multimodal AI provided by the embodiments of the present invention;

[0026] Figure 7 Schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0029] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] Please refer to Figure 1 and Figure 2 , Figure 1 which is provided for an embodiment of the present invention Figure 1 is a schematic diagram of an application scenario of a multi-modal AI-driven cardiovascular image detection method provided for an embodiment of the present invention; Figure 2 is a schematic flowchart of a multi-modal AI-driven cardiovascular image detection method provided for an embodiment of the present invention; this multi-modal AI-driven cardiovascular image detection method is applied to the server 10, and the server 10 is communicatively connected to a user terminal 20 used by a doctor user. As Figure 2 shown, the method includes steps S110 to S170.

[0032] S110. In response to an image detection instruction, obtain a set of cardiovascular images to be detected corresponding to the image detection instruction.

[0033] Among them, the set of cardiovascular images to be detected includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images.

[0034] In this embodiment, in order to obtain cardiovascular images of the user to be detected from more dimensions, at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images (the full name of CT is Computed Tomography, indicating computed tomography), and cardiovascular MRI images (the full name of MRI is Nuclear Magnetic Resonance Imaging, indicating nuclear magnetic resonance imaging) can be obtained for the same user to be detected and combined into a set of cardiovascular images to be detected. Among them, in order to ensure the validity of the images in the set of cardiovascular images to be detected, the image acquisition time of each cardiovascular image to be detected included in the set of cardiovascular images to be detected can be further judged, and after determining the maximum image acquisition interval duration, it is compared with a preset interval duration, and when the maximum image acquisition interval duration is less than or equal to the preset interval duration, it is determined that the set of cardiovascular images to be detected corresponds to a valid image set detection result.

[0035] Among them, if each cardiovascular image to be detected in the set of cardiovascular images to be detected is received by the server in the form of an encrypted image, it can be pre-agreed that the image acquisition time of each cardiovascular image to be detected uses plaintext information, and other information of each cardiovascular image to be detected can be encrypted.

[0036] S120. If it is determined that the cardiovascular image set to be detected is an encrypted image set, generate target decryption information based on the target user information corresponding to the image detection instruction, and decrypt the cardiovascular image set to be detected based on the target decryption information to update the cardiovascular image set to be detected.

[0037] In this embodiment, if it is determined in the server that the cardiovascular image set to be detected is an encrypted image set, it means that the server cannot directly process each cardiovascular image to be detected in the cardiovascular image set to be detected. At this time, it is necessary to first generate target decryption information, that is, generate a decryption key, based on the target user information corresponding to the image detection instruction, and decrypt the cardiovascular image set to be detected based on the target decryption information, so as to update the cardiovascular image set to be detected. Moreover, at this time, the updated cardiovascular image set to be detected is in plain text form and can be further processed by the server for image processing.

[0038] In one embodiment, as Figure 3 shown, step S120 includes:

[0039] S121. Obtain the last digit identification bit of the user information corresponding to the target user information;

[0040] S122. Obtain the target decryption model corresponding to the last digit identification bit of the user information from multiple decryption models locally based on the last digit identification bit of the user information, and use it as the target decryption information;

[0041] S123. Decrypt the cardiovascular image set to be detected based on the target decryption information to obtain the decrypted cardiovascular image set to be detected, and update the cardiovascular image set to be detected.

[0042] In this embodiment, to improve data security, multiple encryption models can be pre-deployed in the server. Then, after each user uses their user terminal to establish a communication connection with the server, when the user terminal uploads a piece of user information each time, in the server, a corresponding encryption model can be obtained based on the last identification bit of the user information and sent to the user terminal for data encryption. Moreover, the server also saves the decryption model corresponding to the encryption model for the corresponding user information. After that, when a user terminal uploads the corresponding target user information (such as including user ID, user name, user gender, etc.), the last identification bit of the user information corresponding to the target user information is obtained (for example, the last character in the user ID can be used as the last identification bit of the user information); then, the target decryption model with the last identification bit of the user information is obtained from the multiple decryption models locally with the last identification bit of the user information and used as the target decryption information; finally, the to-be-detected cardiovascular image set is decrypted according to the target decryption information to obtain the decrypted to-be-detected cardiovascular image set in plaintext form, and the to-be-detected cardiovascular image set is updated accordingly.

[0043] S130. Determine a target multi-modal AI model set corresponding to the to-be-detected cardiovascular image set based on the composition ratio of image types in the to-be-detected cardiovascular image set and a preset multi-modal AI model determination strategy.

[0044] In this embodiment, after obtaining the decrypted to-be-detected cardiovascular image set in plaintext form, the composition ratio of image types therein can be determined first. For example, the total number of images included in the to-be-detected cardiovascular image set is recorded as N (where N is a positive integer), and the first total number of vascular ultrasound images is recorded as N1 and the second total number of cardiovascular CT images is recorded as N2 (where both N1 and N2 are positive integers). Then, the composition ratio of image types can be determined according to the ratios of the first total number of vascular ultrasound images N1 and the second total number of cardiovascular CT images N2 to the total number of images N included in the to-be-detected cardiovascular image set. More specifically, the proportion of the image type of vascular ultrasound images is N1 / N, and the proportion of the image type of cardiovascular CT images is N2 / N. After knowing the composition ratio of image types, the preset multi-modal AI model determination strategy locally in the server can be combined to determine the target multi-modal AI model set corresponding to the to-be-detected cardiovascular image set, realizing the determination of the target multi-modal AI model set including the target multi-modal AI model in a more flexible manner.

[0045] In one embodiment, as the first embodiment of step S130, as Figure 4 shown, step S130 includes:

[0046] S131. Obtain the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively, and determine the image type composition ratio based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio;

[0047] S132. If it is determined based on the multi-modal AI model determination strategy that the first image type ratio in the image type composition ratio exceeds the first preset ratio and the second image type ratio exceeds the second preset ratio, obtain the first pre-trained AI model and the second pre-trained AI model to form the target multi-modal AI model set;

[0048] S133. If it is determined based on the multi-modal AI model determination strategy that the third image type ratio in the image type composition ratio exceeds the third preset ratio and the fourth image type ratio exceeds the fourth preset ratio, obtain the third pre-trained AI model and the fourth pre-trained AI model to form the target multi-modal AI model set.

[0049] In this embodiment, as the first embodiment of step S130, after determining the image type composition ratio in the to-be-detected cardiovascular image set in the server, the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set can be specifically obtained. When the image type ratio is 0, it means that the number of corresponding image types in the to-be-detected cardiovascular image set is 0. After determining the above four image type ratios, the image type composition ratio can be specifically determined.

[0050] After determining the image type composition ratio, if it is specifically determined that the first image type ratio in the image type composition ratio exceeds the first preset ratio (for example, the first preset ratio is set to 20%) and the second image type ratio exceeds the second preset ratio (for example, the second preset ratio is set to 30%), obtain the first pre-trained AI model (such as a residual network model, i.e., ResNet model) and the second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) to form the target multi-modal AI model set.

[0051] After determining the composition ratio of image types, if it is specifically determined that the ratio of the third image type in the composition ratio of the image types exceeds a third preset ratio (for example, setting the third preset ratio to 30%) and the ratio of the fourth image type exceeds a fourth preset ratio (for example, setting the fourth preset ratio to 40%), then obtain a third pre-trained AI model (such as a 3D U-Net model) and a fourth pre-trained AI model (such as a 3D CNN model, that is, a three-dimensional convolutional neural network model) to form the target multi-modal AI model set. It can be seen that through the above method, based on the composition ratio of image types and the preset multi-modal AI model determination strategy, the target multi-modal AI model set suitable for the current cardiovascular image set to be detected can be quickly determined.

[0052] In one embodiment, after step S131, it further includes:

[0053] If it is determined based on the multi-modal AI model determination strategy that the ratio of the first image type in the composition ratio of the image types exceeds a first preset ratio and the ratio of the third image type exceeds a third preset ratio, then obtain a first pre-trained AI model and a third pre-trained AI model to form the target multi-modal AI model set;

[0054] If it is determined based on the multi-modal AI model determination strategy that the ratio of the second image type in the composition ratio of the image types exceeds a second preset ratio and the ratio of the fourth image type exceeds a fourth preset ratio, then obtain a second pre-trained AI model and a fourth pre-trained AI model to form the target multi-modal AI model set.

[0055] In this embodiment, similarly, after determining the composition ratio of image types, if it is specifically determined that the ratio of the first image type in the composition ratio of the image types exceeds a first preset ratio and the ratio of the third image type exceeds a third preset ratio, then obtain a first pre-trained AI model (such as a residual network model, that is, a ResNet model) and a third pre-trained AI model (such as a 3D U-Net model) to form the target multi-modal AI model set.

[0056] After determining the composition ratio of image types, if it is specifically determined that the ratio of the second image type in the composition ratio of the image types exceeds a second preset ratio and the ratio of the fourth image type exceeds a fourth preset ratio, then obtain a second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) and a fourth pre-trained AI model (such as a 3D CNN model, that is, a three-dimensional convolutional neural network model) to form the target multi-modal AI model set. It can be seen that through the above method, based on the composition ratio of image types and the preset multi-modal AI model determination strategy, the target multi-modal AI model set suitable for the current cardiovascular image set to be detected can also be quickly determined.

[0057] In one embodiment, as the second embodiment of step S130, step S130 includes:

[0058] Obtain the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively, and determine the image type composition ratio based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio.

[0059] Based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio in the image type composition ratio, determine the output result weight values corresponding to each AI model in the initial preset multi-modal AI model set, and determine the target multi-modal AI model set based on the output result weight values corresponding to each AI model and the initial preset multi-modal AI model set.

[0060] In this embodiment, as the second embodiment of step S130, after determining the image type composition ratio in the to-be-detected cardiovascular image set in the server, it is also possible to specifically obtain the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively. At this time, the first image type ratio can be directly used as the first output result weight value corresponding to the first pre-trained AI model (such as a residual network model, i.e., ResNet model) in the initial preset multi-modal AI model set, the second image type ratio can be used as the second output result weight value corresponding to the second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) in the initial preset multi-modal AI model set, the third image type ratio can be used as the third output result weight value corresponding to the third pre-trained AI model (such as a 3D U-Net model) in the initial preset multi-modal AI model set, and the fourth image type ratio can be used as the fourth output result weight value corresponding to the fourth pre-trained AI model (such as a 3D CNN model, i.e., a three-dimensional convolutional neural network model) in the initial preset multi-modal AI model set. After determining the output result weight values corresponding to each AI model in the initial preset multi-modal AI model set in the above manner, the target multi-modal AI model set for subsequent image detection can be determined.

[0061] S140. For each cardiovascular image to be detected in the set of cardiovascular images to be detected, obtain the corresponding target multi-modal AI model from the set of target multi-modal AI models, and input the cardiovascular image to be detected into the corresponding target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected.

[0062] In this embodiment, after determining the set of target multi-modal AI models in the server, image detection can be performed on each cardiovascular image to be detected in the set of cardiovascular images to be detected to obtain corresponding image detection results. For example, taking the image detection process of a cardiovascular image to be detected in the set of cardiovascular images to be detected as an example, after obtaining the corresponding target multi-modal AI model of the cardiovascular image to be detected in the set of target multi-modal AI models (specifically, the corresponding target multi-modal AI model can be selected from the first pre-trained AI model to the fourth pre-trained AI model according to the image type of the cardiovascular image to be detected), then input the cardiovascular image to be detected into the target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected. Among them, the image detection result corresponding to the cardiovascular image to be detected is a probability value with a value range of [0, 1], and represents the probability value of the cardiovascular image to be detected corresponding to the corresponding image classification result. The image detection processes of other cardiovascular images to be detected in the set of cardiovascular images to be detected also refer to the above process. After completing the image detection of each cardiovascular image to be detected in the set of cardiovascular images to be detected, subsequent operations can be performed.

[0063] In one embodiment, before step S140, it further includes:

[0064] For each cardiovascular image to be detected in the set of cardiovascular images to be detected, perform image preprocessing on the cardiovascular image to be detected based on a preset set of image preprocessing strategies to update the cardiovascular image to be detected.

[0065] In this embodiment, still referring to the above example, taking the image preprocessing process of a cardiovascular image to be detected in the cardiovascular image set to be detected as an example, after obtaining the image type of the cardiovascular image to be detected, the target image preprocessing strategy corresponding to the cardiovascular image to be detected can also be obtained from the image preprocessing strategy set. For example, if the image type of the cardiovascular image to be detected is a cardiovascular ultrasound image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a first preset cropped image size), image denoising, and normalization processing; if the image type of the cardiovascular image to be detected is a cardiovascular X-ray image, the corresponding target image preprocessing strategy includes image denoising and normalization processing; if the image type of the cardiovascular image to be detected is a cardiovascular CT image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a second preset cropped image size), image denoising, and normalization processing; if the image type of the cardiovascular image to be detected is a cardiovascular MRI image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a third preset cropped image size), image denoising, and normalization processing. After preprocessing the cardiovascular image to be detected through the target image preprocessing strategy in the image preprocessing strategy set, the image preprocessing of the cardiovascular image to be detected is realized in a timely manner, which is convenient for more accurate image detection in the subsequent process.

[0066] S150. Based on a preset detection result fusion model and the image detection results corresponding to each cardiovascular image to be detected in the cardiovascular image set to be detected, obtain a fusion detection result.

[0067] In this embodiment, after obtaining the image detection results corresponding to each cardiovascular image to be detected in the cardiovascular image set to be detected in the server, the detection result fusion model can also be used to fuse the image detection results to obtain a fusion detection result. It can be seen that through the above method, the fusion detection result corresponding to the cardiovascular image set to be detected can be quickly obtained.

[0068] In one embodiment, as Figure 5 shown, step S150 includes:

[0069] S151. Based on the detection result fusion model, arrange the image detection results corresponding to each cardiovascular image to be detected in the cardiovascular image set to be detected in the chronological order of image acquisition time, and form a multi-modal feature vector.

[0070] S152. Input the multi-modal feature vector into the long short-term memory network corresponding to the detection result fusion model to obtain the fusion detection result corresponding to the multi-modal feature vector.

[0071] In this embodiment, after obtaining the image detection results corresponding to each to-be-detected vascular image in the to-be-detected cardiovascular image set (i.e., after corresponding to a probability value ranging from 0 to 1), the above image detection results can also be arranged in the chronological order of image acquisition time and composed into a multi-modal feature vector. Then, the multi-modal feature vector is input into the long short-term memory network corresponding to the detection result fusion model (more specifically, the self-attention mechanism can also be fused in this long short-term memory network) to obtain the fusion detection result corresponding to the multi-modal feature vector (and the fusion detection result is also a probability value ranging from 0 to 1). Finally, the fusion detection result can also be compared with multiple preset value ranges (the union of the above multiple preset value ranges is also [0, 1]) to determine the target preset value range corresponding to the fusion detection result in the multiple preset value ranges, and use the classification prompt information corresponding to the target preset value range (such as corresponding to the low-risk degree type) as the current classification prompt information sent by the server to the user terminal.

[0072] It can be seen that the embodiment implementing this method can, after determining the target multi-modal AI model set based on the composition ratio of image types for the to-be-detected cardiovascular image set including multiple types of to-be-detected cardiovascular images, perform multi-modal image detection and detection result fusion accordingly, so as to obtain a fusion detection result that can reflect multi-dimensional evaluation results and has more comprehensive corresponding information.

[0073] An embodiment of the present invention also provides a cardiovascular image detection device driven by multi-modal AI. The cardiovascular image detection device driven by multi-modal AI can be configured in a server, and the cardiovascular image detection device driven by multi-modal AI is used to execute any embodiment of the foregoing cardiovascular image detection method driven by multi-modal AI. Specifically, please refer to Figure 6 , Figure 6 which is a schematic block diagram of the cardiovascular image detection device driven by multi-modal AI provided by an embodiment of the present invention. As Figure 6 shown, the cardiovascular image detection device 100 driven by multi-modal AI includes a to-be-detected cardiovascular image set acquisition unit 110, an image set decryption unit 120, a target multi-modal AI model set acquisition unit 130, an image detection result acquisition unit 140, and a detection result fusion unit 150.

[0074] The to-be-detected cardiovascular image set acquisition unit 110 is configured to obtain a to-be-detected cardiovascular image set corresponding to the image detection instruction in response to the image detection instruction.

[0075] Among them, the to-be-detected cardiovascular image set includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images.

[0076] In this embodiment, in order to obtain cardiovascular images of the user to be detected from more dimensions, at least two of the cardiovascular ultrasound image, cardiovascular X-ray image, cardiovascular CT image (the full name of CT is Computed Tomography, indicating computed tomography), and cardiovascular MRI image (the full name of MRI is Nuclear Magnetic Resonance Imaging, indicating nuclear magnetic resonance imaging) can be obtained for the same user to be detected and form a set of cardiovascular images to be detected. Among them, in order to ensure the validity of the images in the set of cardiovascular images to be detected, the image acquisition time of each cardiovascular image to be detected included in the set of cardiovascular images to be detected can be further determined, and after determining the maximum image acquisition interval duration, it is compared with the preset interval duration, and when the maximum image acquisition interval duration is less than or equal to the preset interval duration, it is determined that the set of cardiovascular images to be detected corresponds to a valid image set detection result.

[0077] Among them, when each cardiovascular image to be detected in the set of cardiovascular images to be detected is received by the server in the form of an encrypted image, it can be pre-agreed that the image acquisition time of each cardiovascular image to be detected uses plaintext information, and other information of each cardiovascular image to be detected can be encrypted.

[0078] The image set decryption unit 120 is configured to, if it is determined that the set of cardiovascular images to be detected is an encrypted image set, generate target decryption information based on the target user information corresponding to the image detection instruction, and decrypt the set of cardiovascular images to be detected based on the target decryption information to update the set of cardiovascular images to be detected.

[0079] In this embodiment, if it is determined in the server that the set of cardiovascular images to be detected is an encrypted image set, it means that the server cannot directly process each cardiovascular image to be detected in the set of cardiovascular images to be detected. At this time, it is necessary to first generate target decryption information, that is, generate a decryption key, based on the target user information corresponding to the image detection instruction, and decrypt the set of cardiovascular images to be detected based on the target decryption information, so as to update the set of cardiovascular images to be detected. Moreover, at this time, the updated set of cardiovascular images to be detected is in plaintext form and can be further processed by the server for image processing.

[0080] In one embodiment, the image set decryption unit 120 is specifically configured to:

[0081] Obtain the last digit identification bit of the user information corresponding to the target user information;

[0082] Obtain the target decryption model corresponding to the last digit identification bit of the user information from multiple decryption models locally based on the last digit identification bit of the user information, and use it as the target decryption information;

[0083] Decrypt the cardiovascular image set to be detected based on the target decryption information to obtain the decrypted cardiovascular image set to be detected, and update the cardiovascular image set to be detected.

[0084] In this embodiment, to improve data security, multiple encryption models can be pre-deployed in the server. Then, after each user uses their user terminal to establish a communication connection with the server, when the user terminal uploads a user information each time, in the server, an encryption model corresponding to the last identification bit of the user information can be obtained based on the last identification bit of the user information and sent to the user terminal for data encryption. Moreover, the server also saves the decryption model corresponding to the encryption model for the corresponding user information. After that, when a user terminal uploads the corresponding target user information (such as including user ID, user name, user gender, etc.), the last identification bit of the user information corresponding to the target user information is obtained (for example, the last character in the user ID can be used as the last identification bit of the user information); then, the target decryption model with the last identification bit of the user information is obtained from the multiple decryption models locally with the last identification bit of the user information, and used as the target decryption information; finally, the cardiovascular image set to be detected is decrypted according to the target decryption information to obtain the decrypted cardiovascular image set to be detected in plaintext form, and the cardiovascular image set to be detected is updated accordingly.

[0085] The target multi-modal AI model set acquisition unit 130 is configured to determine a target multi-modal AI model set corresponding to the cardiovascular image set to be detected based on the image type composition ratio in the cardiovascular image set to be detected and a preset multi-modal AI model determination strategy.

[0086] In this embodiment, after obtaining the decrypted cardiovascular image set to be detected in plaintext form, the image type composition ratio can be determined first. For example, the total number of images included in the cardiovascular image set to be detected is recorded as N (where N is a positive integer), and the first total number of vascular ultrasound images is recorded as N1 and the second total number of cardiovascular CT images is recorded as N2 (where both N1 and N2 are positive integers). Then, the image type composition ratio can be determined according to the ratio of the first total number of vascular ultrasound images N1 and the second total number of cardiovascular CT images N2 to the total number of images N included in the cardiovascular image set to be detected. More specifically, the image type proportion of vascular ultrasound images is N1 / N and the image type proportion of cardiovascular CT images is N2 / N. After knowing the image type composition ratio, the target multi-modal AI model set corresponding to the cardiovascular image set to be detected can be determined in combination with the multi-modal AI model determination strategy preset locally in the server, realizing the determination of the target multi-modal AI model set including the target multi-modal AI model in a more flexible manner.

[0087] In one embodiment, as the first embodiment of the target multi-modal AI model set acquisition unit 130, the target multi-modal AI model set acquisition unit 130 is specifically configured to:

[0088] Obtain the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively, and determine the image type composition ratio based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio;

[0089] If it is determined based on the multi-modal AI model determination strategy that the first image type ratio in the image type composition ratio exceeds the first preset ratio and the second image type ratio exceeds the second preset ratio, then obtain the first pre-trained AI model and the second pre-trained AI model to form the target multi-modal AI model set;

[0090] If it is determined based on the multi-modal AI model determination strategy that the third image type ratio in the image type composition ratio exceeds the third preset ratio and the fourth image type ratio exceeds the fourth preset ratio, then obtain the third pre-trained AI model and the fourth pre-trained AI model to form the target multi-modal AI model set.

[0091] In this embodiment, as the first embodiment of the target multi-modal AI model set acquisition unit 130, after determining the image type composition ratio in the to-be-detected cardiovascular image set in the server, the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set can be specifically obtained, where when an image type ratio is 0, it means that the number of corresponding images in the to-be-detected cardiovascular image set is 0. After determining the above four image type ratios, the image type composition ratio can be specifically determined.

[0092] After determining the image type composition ratio, if it is specifically determined that the first image type ratio in the image type composition ratio exceeds the first preset ratio (for example, setting the first preset ratio to 20%) and the second image type ratio exceeds the second preset ratio (for example, setting the second preset ratio to 30%), then obtain the first pre-trained AI model (such as a residual network model, i.e., a ResNet model) and the second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) to form the target multi-modal AI model set.

[0093] After determining the composition ratio of the image types, if it is specifically determined that the ratio of the third image type in the composition ratio of the image types exceeds a third preset ratio (for example, setting the third preset ratio to 30%) and the ratio of the fourth image type exceeds a fourth preset ratio (for example, setting the fourth preset ratio to 40%), then obtain a third pre-trained AI model (such as a 3D U-Net model) and a fourth pre-trained AI model (such as a 3D CNN model, that is, a three-dimensional convolutional neural network model) to form the target multi-modal AI model set. It can be seen that through the above method, based on the composition ratio of the image types and the preset multi-modal AI model determination strategy, the target multi-modal AI model set applicable to the current cardiovascular image set to be detected can be quickly determined.

[0094] In one embodiment, the target multi-modal AI model set acquisition unit 130 is further configured to:

[0095] If it is determined based on the multi-modal AI model determination strategy that the ratio of the first image type in the composition ratio of the image types exceeds a first preset ratio and the ratio of the third image type exceeds a third preset ratio, then obtain a first pre-trained AI model and a third pre-trained AI model to form the target multi-modal AI model set;

[0096] If it is determined based on the multi-modal AI model determination strategy that the ratio of the second image type in the composition ratio of the image types exceeds a second preset ratio and the ratio of the fourth image type exceeds a fourth preset ratio, then obtain a second pre-trained AI model and a fourth pre-trained AI model to form the target multi-modal AI model set.

[0097] In this embodiment, similarly, after determining the composition ratio of the image types, if it is specifically determined that the ratio of the first image type in the composition ratio of the image types exceeds a first preset ratio and the ratio of the third image type exceeds a third preset ratio, then obtain a first pre-trained AI model (such as a residual network model, that is, a ResNet model) and a third pre-trained AI model (such as a 3D U-Net model) to form the target multi-modal AI model set.

[0098] After determining the composition ratio of the image types, if it is specifically determined that the ratio of the second image type in the composition ratio of the image types exceeds a second preset ratio and the ratio of the fourth image type exceeds a fourth preset ratio, then obtain a second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) and a fourth pre-trained AI model (such as a 3D CNN model, that is, a three-dimensional convolutional neural network model) to form the target multi-modal AI model set. It can be seen that through the above method, based on the composition ratio of the image types and the preset multi-modal AI model determination strategy, the target multi-modal AI model set applicable to the current cardiovascular image set to be detected can also be quickly determined.

[0099] In one embodiment, as the second embodiment of the target multimodal AI model set obtaining unit 130, the target multimodal AI model set obtaining unit 130 is specifically configured to:

[0100] Obtain a first image type ratio, a second image type ratio, a third image type ratio, and a fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively, and determine an image type composition ratio based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio;

[0101] Determine the output result weight values corresponding to each AI model in the initial preset multimodal AI model set based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio in the image type composition ratio, and determine the target multimodal AI model set based on the output result weight values corresponding to each AI model and the initial preset multimodal AI model set.

[0102] In this embodiment, as the second embodiment of the target multimodal AI model set obtaining unit 130, after determining the image type composition ratio in the to-be-detected cardiovascular image set in the server, it is also possible to specifically obtain the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image, and the cardiovascular MRI image in the to-be-detected cardiovascular image set respectively. At this time, the first image type ratio can be directly used as the first output result weight value corresponding to the first pre-trained AI model (such as a residual network model, i.e., a ResNet model) in the initial preset multimodal AI model set, the second image type ratio can be used as the second output result weight value corresponding to the second pre-trained AI model (such as a convolutional neural network model combined with an attention mechanism) in the initial preset multimodal AI model set, the third image type ratio can be used as the third output result weight value corresponding to the third pre-trained AI model (such as a 3D U-Net model) in the initial preset multimodal AI model set, and the fourth image type ratio can be used as the fourth output result weight value corresponding to the fourth pre-trained AI model (such as a 3DCNN model, i.e., a three-dimensional convolutional neural network model) in the initial preset multimodal AI model set. After determining the output result weight values corresponding to each AI model in the initial preset multimodal AI model set in the above manner, the target multimodal AI model set for image detection can be determined.

[0103] An image detection result acquisition unit 140 is configured to, for each cardiovascular image to be detected in the cardiovascular images to be detected, obtain a target multi-modal AI model corresponding to the cardiovascular image to be detected from the target multi-modal AI model set, and input the cardiovascular image to be detected into the corresponding target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected.

[0104] In this embodiment, after the target multi-modal AI model set is determined in the server, image detection can be performed on each cardiovascular image to be detected in the cardiovascular images to be detected to obtain corresponding image detection results. For example, taking the image detection process of a cardiovascular image to be detected in the cardiovascular images to be detected as an example, after obtaining the target multi-modal AI model corresponding to the cardiovascular image to be detected in the target multi-modal AI model set (specifically, the corresponding target multi-modal AI model can be selected from the first pre-trained AI model to the fourth pre-trained AI model according to the image type of the cardiovascular image to be detected), then input the cardiovascular image to be detected into the target multi-modal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected. Among them, the image detection result corresponding to the cardiovascular image to be detected is a probability value with a value range of [0, 1], and represents the probability value of the cardiovascular image to be detected corresponding to the corresponding image classification result. The image detection processes of other cardiovascular images to be detected in the cardiovascular images to be detected also refer to the above process. After the image detection of each cardiovascular image to be detected in the cardiovascular images to be detected is completed, subsequent operations can be performed.

[0105] In one embodiment, the cardiovascular image detection device 100 driven by multi-modal AI further includes:

[0106] An image preprocessing unit is configured to, for each cardiovascular image to be detected in the cardiovascular images to be detected, perform image preprocessing on the cardiovascular image to be detected based on a preset image preprocessing strategy set to update the cardiovascular image to be detected.

[0107] In this embodiment, still referring to the above example, taking the image preprocessing process of a to-be-detected cardiovascular image in the to-be-detected cardiovascular image set as an example, after obtaining the image type of the to-be-detected cardiovascular image, a target image preprocessing strategy corresponding to the to-be-detected cardiovascular image can also be obtained from the image preprocessing strategy set. For example, if the image type of the to-be-detected cardiovascular image is a cardiovascular ultrasound image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a first preset cropped image size), image denoising, and normalization processing; if the image type of the to-be-detected cardiovascular image is a cardiovascular X-ray image, the corresponding target image preprocessing strategy includes image denoising and normalization processing; if the image type of the to-be-detected cardiovascular image is a cardiovascular CT image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a second preset cropped image size), image denoising, and normalization processing; if the image type of the to-be-detected cardiovascular image is a cardiovascular MRI image, the corresponding target image preprocessing strategy includes image cropping (corresponding to a third preset cropped image size), image denoising, and normalization processing. After preprocessing the to-be-detected cardiovascular image through the target image preprocessing strategy in the image preprocessing strategy set, the image preprocessing of the to-be-detected cardiovascular image is realized in a timely manner, which is convenient for more accurate image detection in the subsequent process.

[0108] The detection result fusion unit 150 is configured to obtain a fused detection result based on a preset detection result fusion model and the image detection results corresponding to each to-be-detected vascular image in the to-be-detected cardiovascular image set.

[0109] In this embodiment, after obtaining the image detection results corresponding to each to-be-detected vascular image in the to-be-detected cardiovascular image set in the server, the detection result fusion model can also be used to fuse the image detection results to obtain a fused detection result. It can be seen that through the above method, the fused detection result corresponding to the to-be-detected cardiovascular image set can be obtained quickly.

[0110] In one embodiment, the detection result fusion unit 150 is specifically configured to:

[0111] Arrange the image detection results corresponding to each to-be-detected vascular image in the to-be-detected cardiovascular image set in the chronological order of image acquisition time based on the detection result fusion model, and form a multi-modal feature vector;

[0112] Input the multi-modal feature vector into the long short-term memory network corresponding to the detection result fusion model to obtain the fused detection result corresponding to the multi-modal feature vector.

[0113] In this embodiment, after obtaining the image detection results corresponding to each to-be-detected vascular image in the to-be-detected cardiovascular image set (i.e., after corresponding to a probability value ranging from 0 to 1), the above-mentioned various image detection results can also be arranged in the chronological order of image acquisition time and composed into a multi-modal feature vector. Then, the multi-modal feature vector is input into the long short-term memory network corresponding to the detection result fusion model (more specifically, a self-attention mechanism can also be fused in this long short-term memory network) to obtain the fusion detection result corresponding to the multi-modal feature vector (and the fusion detection result is also a probability value ranging from 0 to 1). Finally, the fusion detection result can also be compared with multiple preset value ranges (the union of the above-mentioned multiple preset value ranges is also [0, 1]) to determine the target preset value range corresponding to the fusion detection result in the multiple preset value ranges, and use the classification prompt information corresponding to the target preset value range (such as corresponding to the low-risk degree type) as the current classification prompt information sent by the server to the user terminal.

[0114] It can be seen that implementing the embodiment of this device can, after determining the target multi-modal AI model set based on the composition ratio of image types for the to-be-detected cardiovascular image set including multiple types of to-be-detected cardiovascular images, correspondingly perform multi-modal image detection and detection result fusion to obtain a fusion detection result that can reflect multi-dimensional evaluation results and has more comprehensive corresponding information.

[0115] The above-mentioned cardiovascular image detection device driven by multi-modal AI can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 7 shown.

[0116] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of a computer device provided by an embodiment of the present invention. This computer device integrates any one of the cardiovascular image detection devices driven by multi-modal AI provided by the embodiments of the present invention.

[0117] Refer to Figure 7 , this computer device 400 includes a processor 402, a memory, and a network interface 405 connected through a system bus 401. Among them, the memory can include a storage medium 403 and an internal memory 404.

[0118] The storage medium 403 can store an operating system 4031 and a computer program 4032. This computer program 4032 includes program instructions, and when these program instructions are executed, the processor 402 can be made to execute the above-mentioned cardiovascular image detection method driven by multi-modal AI.

[0119] The processor 402 is used to provide computing and control capabilities to support the operation of the entire computer device.

[0120] The internal memory 404 provides an environment for the operation of the computer program 4032 in the storage medium 403. When the computer program 4032 is executed by the processor 402, the processor 402 can be caused to execute the above-mentioned cardiovascular image detection method driven by multimodal AI.

[0121] The network interface 405 is used for network communication with other devices. Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0122] Among them, the processor 402 is used to run the computer program 4032 stored in the memory to implement the above-mentioned cardiovascular image detection method driven by multimodal AI.

[0123] It should be understood that in the embodiment of the present invention, the processor 402 may be a central processing unit (CPU), and the processor 402 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0124] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer device to implement the process steps of the embodiments of the above methods.

[0125] Therefore, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by the processor, the processor is caused to execute the above-mentioned cardiovascular image detection method driven by multimodal AI.

[0126] The computer-readable storage medium may be a USB flash drive, a portable hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which are all computer-readable storage media that can store program codes.

[0127] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0128] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0129] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0131] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A cardiovascular image detection method based on multimodal AI driving, characterized in that: include: In response to an image detection instruction, acquiring a cardiovascular image set to be detected corresponding to the image detection instruction; wherein the cardiovascular image set to be detected includes at least two of cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images; If it is determined that the cardiovascular image set to be detected is an encrypted image set, generating target decryption information based on target user information corresponding to the image detection instruction, and decrypting the cardiovascular image set to be detected based on the target decryption information to update the cardiovascular image set to be detected; Based on the image type composition ratio in the cardiovascular image set to be detected and the preset multimodal AI model determination strategy, determine the target multimodal AI model set corresponding to the cardiovascular image set to be detected; the multimodal AI model determination strategy is used to determine the target multimodal AI model set from multiple initial preset multimodal AI models or the initial preset multimodal AI model set based on the image type composition ratio corresponding to cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images and cardiovascular MRI images in the cardiovascular image set to be detected and the corresponding ratio of each image type, and based on the size relationship of each image type ratio; For each cardiovascular image to be detected in the cardiovascular image set to be detected, a target multimodal AI model corresponding to the cardiovascular image to be detected is obtained in the target multimodal AI model set, and the cardiovascular image to be detected is input into the corresponding target multimodal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected; Based on a preset detection result fusion model and the image detection result corresponding to each blood vessel image to be detected in the cardiovascular image set to be detected, a fusion detection result is obtained.

2. The cardiovascular image detection method based on multimodal AI drive according to claim 1, characterized in that: The generating target decryption information based on the target user information corresponding to the image detection instruction, and decrypting the cardiovascular image set to be detected based on the target decryption information to update the cardiovascular image set to be detected, comprises: Obtain the last identification digit of the user information corresponding to the target user information; Based on the last identification bit of the user information, a target decryption model corresponding to the last identification bit of the user information is obtained from multiple local decryption models, and used as the target decryption information; The cardiovascular image set to be detected is decrypted based on the target decryption information to obtain the decrypted cardiovascular image set to be detected, and the cardiovascular image set to be detected is updated.

3. The cardiovascular image detection method based on multimodal AI drive according to claim 1, characterized in that: The step of determining a target multimodal AI model set corresponding to the cardiovascular image set to be detected based on the image type composition ratio in the cardiovascular image set to be detected and a preset multimodal AI model determination strategy includes: Acquire a first image type ratio, a second image type ratio, a third image type ratio and a fourth image type ratio respectively corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image and the cardiovascular MRI image in the cardiovascular image set to be detected, and determine the image type composition ratio by using the first image type ratio, the second image type ratio, the third image type ratio and the fourth image type ratio; If it is determined based on the multimodal AI model determination strategy that the first image type ratio in the image type composition ratio exceeds a first preset ratio, and the second image type ratio exceeds a second preset ratio, then obtaining a first pre-trained AI model and a second pre-trained AI model to form the target multimodal AI model set; If it is determined based on the multimodal AI model determination strategy that the proportion of the third image type in the image type composition ratio exceeds the third preset ratio, and the proportion of the fourth image type exceeds the fourth preset ratio, then the third pre-trained AI model and the fourth pre-trained AI model are obtained to form the target multimodal AI model set.

4. The cardiovascular image detection method based on multimodal AI drive according to claim 3, characterized in that: After the step of obtaining the first image type ratio, the second image type ratio, the third image type ratio and the fourth image type ratio respectively corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image and the cardiovascular MRI image in the cardiovascular image set to be detected, and determining the image type composition ratio by using the first image type ratio, the second image type ratio, the third image type ratio and the fourth image type ratio, the method further includes: If it is determined based on the multimodal AI model determination strategy that the first image type ratio in the image type composition ratio exceeds a first preset ratio, and the third image type ratio exceeds a third preset ratio, then obtaining a first pre-trained AI model and a third pre-trained AI model to form the target multimodal AI model set; If it is determined based on the multimodal AI model determination strategy that the proportion of the second image type in the image type composition ratio exceeds the second preset ratio, and the proportion of the fourth image type exceeds the fourth preset ratio, then the second pre-trained AI model and the fourth pre-trained AI model are obtained to form the target multimodal AI model set.

5. The cardiovascular image detection method based on multimodal AI drive according to claim 1, characterized in that: The step of determining a target multimodal AI model set corresponding to the cardiovascular image set to be detected based on the image type composition ratio in the cardiovascular image set to be detected and a preset multimodal AI model determination strategy includes: Acquire a first image type ratio, a second image type ratio, a third image type ratio and a fourth image type ratio respectively corresponding to the cardiovascular ultrasound image, the cardiovascular X-ray image, the cardiovascular CT image and the cardiovascular MRI image in the cardiovascular image set to be detected, and determine the image type composition ratio by using the first image type ratio, the second image type ratio, the third image type ratio and the fourth image type ratio; Based on the first image type ratio, the second image type ratio, the third image type ratio, and the fourth image type ratio in the image type composition ratio, determine the output result weight value corresponding to each AI model in the initial preset multimodal AI model set, and determine the target multimodal AI model set based on the output result weight value corresponding to each AI model and the initial preset multimodal AI model set.

6. The cardiovascular image detection method based on multimodal AI drive according to claim 1, characterized in that: Before the step of acquiring, for each cardiovascular image to be detected in the cardiovascular image set, a target multimodal AI model corresponding to the cardiovascular image to be detected in the target multimodal AI model set, and inputting the cardiovascular image to be detected into the corresponding target multimodal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected, the method further includes: For each cardiovascular image to be detected in the set of cardiovascular images to be detected, image preprocessing is performed on the cardiovascular image to be detected based on a preset set of image preprocessing strategies to update the cardiovascular image to be detected.

7. The cardiovascular image detection method based on multimodal AI drive according to claim 1, characterized in that: The obtaining of the fusion detection result based on the preset detection result fusion model and the image detection result corresponding to each blood vessel image to be detected in the cardiovascular image set to be detected includes: Based on the detection result fusion model, the image detection results corresponding to each blood vessel image to be detected in the cardiovascular image set to be detected are arranged in the temporal order of image acquisition time to form a multimodal feature vector; The multimodal feature vector is input into the long short-term memory network corresponding to the detection result fusion model to obtain the fusion detection result corresponding to the multimodal feature vector.

8. A cardiovascular image detection device based on multimodal AI drive, characterized in that: include: a cardiovascular image set acquisition unit to be detected, used for acquiring a cardiovascular image set to be detected corresponding to the image detection instruction in response to the image detection instruction; wherein the cardiovascular image set to be detected includes at least two of the following: cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images; an image set decryption unit, configured to generate target decryption information based on target user information corresponding to the image detection instruction if it is determined that the cardiovascular image set to be detected is an encrypted image set, and decrypt the cardiovascular image set to be detected based on the target decryption information to update the cardiovascular image set to be detected; a target multimodal AI model set acquisition unit, configured to determine a target multimodal AI model set corresponding to the cardiovascular image set to be detected based on the image type composition ratio in the cardiovascular image set to be detected and a preset multimodal AI model determination strategy; the multimodal AI model determination strategy is configured to determine the target multimodal AI model set from a plurality of initial preset multimodal AI models or an initial preset multimodal AI model set based on the image type composition ratio corresponding to cardiovascular ultrasound images, cardiovascular X-ray images, cardiovascular CT images, and cardiovascular MRI images in the cardiovascular image set to be detected and the corresponding image type ratios, and based on the size relationship of the image type ratios; an image detection result acquisition unit, for acquiring, for each cardiovascular image to be detected in the cardiovascular image set, a target multimodal AI model corresponding to the cardiovascular image to be detected in the target multimodal AI model set, and inputting the cardiovascular image to be detected into the corresponding target multimodal AI model to obtain an image detection result corresponding to the cardiovascular image to be detected; The detection result fusion unit is used to obtain a fusion detection result based on a preset detection result fusion model and the image detection result corresponding to each blood vessel image to be detected in the cardiovascular image set to be detected.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer device executes the computer program, the multimodal AI-driven cardiovascular image detection method described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multimodal AI-driven cardiovascular image detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Training method and apparatus for image fusion processing model, device, and storage medium

    US20210166088A1

  • Multi-model fusion-based object detection method and apparatus, device and medium

    WO2023071121A1