A carotid artery auxiliary diagnosis method and system based on ultrasonic visual language
Patent Information
- Application Number
- CN202610988640.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]颈动脉超声凭借无创、实时、低成本、可重复的特点,是临床筛查颈动脉内膜增厚、粥样硬化、斑块、血管狭窄及评估脑血管事件风险的重要手段,临床检查需覆盖双侧多段颈动脉并结合多切面、多模态成像,诊断过程依赖多血管结构、斑块特征及血流参数的综合交叉研判,具备强关联性与结构化特点;当前颈动脉超声智能诊断技术多基于单帧、单切面图像,依托卷积神经网络、视觉Transformer、视觉语言模型实现斑块识别、参数测量、狭窄分级等任务,但该方式割裂了同一患者多血管、多切面的关联信息,与临床综合诊断逻辑不符,易出现漏诊、误诊,诊断准确性与鲁棒性有限,同时主流的视觉语言预训练对比学习方法将非配对样本统一作为负样本训练,忽略了颈动脉超声诊断标签固有的层级关联与语义连续关系,易使模型学习到违背临床先验的特征,降低诊断适配精度,且现有技术未能充分挖掘超声报告中的结构化医学先验知识、未构建标准化可计算的诊断知识图谱,导致模型预训练缺乏医学专业约束、特征可解释性差,同时模型仅学习单图像特征,忽视了颈动脉天然的解剖拓扑结构,无法利用多血管段的结构约束与证据互补关系,难以有效表征患者整体动脉粥样硬化负荷,整体存在跨模态学习不合理、医学知识融合不足、血管拓扑特征缺失等缺陷,难以满足临床高精度、高鲁棒性、高可解释性的诊断需求
第一,本发明将颈动脉超声诊断知识从自然语言描述升级为可计算知识图谱,使解剖位置、病灶属性、血流参数和诊断结论之间的医学关系能够显式参与模型训练。
Smart Images

Figure CN122800191A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent diagnosis and treatment technology, and more specifically, to a carotid artery-assisted diagnosis method and system based on ultrasound visual language. Background Technology
[0002] Carotid ultrasound, with its non-invasive, real-time, low-cost, and repeatable characteristics, is an important tool for clinical screening of carotid intimal thickening, atherosclerosis, plaques, vascular stenosis, and assessment of cerebrovascular event risk. Clinical examination requires coverage of multiple segments of both carotid arteries and the use of multi-sectional and multi-modal imaging. The diagnostic process relies on the comprehensive cross-analysis of multiple vascular structures, plaque characteristics, and blood flow parameters, exhibiting strong correlation and structured characteristics. Current intelligent diagnostic technologies for carotid ultrasound are mostly based on single-frame, single-section images, relying on convolutional neural networks, visual Transformers, and visual language models to achieve tasks such as plaque recognition, parameter measurement, and stenosis classification. However, this approach fragments the interconnected information of multiple vessels and multiple sections within the same patient, which is inconsistent with the logic of comprehensive clinical diagnosis, easily leading to missed diagnoses and misdiagnoses, and has limited diagnostic accuracy and robustness. Furthermore, mainstream visual language... Pre-training contrastive learning methods uniformly treat unpaired samples as negative samples for training, ignoring the inherent hierarchical associations and semantic continuity of carotid ultrasound diagnostic labels. This can easily lead the model to learn features that violate clinical priors, reducing diagnostic accuracy. Furthermore, current technologies fail to fully exploit the structured medical prior knowledge in ultrasound reports and do not construct standardized, computable diagnostic knowledge graphs. This results in a lack of medical professional constraints and poor feature interpretability during model pre-training. Additionally, the model only learns single-image features, ignoring the natural anatomical topology of the carotid artery. It cannot utilize the structural constraints and complementary evidence relationships of multiple vessel segments, making it difficult to effectively characterize the overall atherosclerotic burden of patients. Overall, it suffers from defects such as unreasonable cross-modal learning, insufficient integration of medical knowledge, and missing vascular topological features, making it difficult to meet the clinical diagnostic needs for high accuracy, high robustness, and high interpretability. Summary of the Invention
[0003] To address the aforementioned problems, the present invention aims to provide a carotid artery-assisted diagnostic method and system based on ultrasound visual language, thereby resolving the existing issues in current carotid artery ultrasound intelligent diagnostic technology.
[0004] To achieve the above technical objectives, this application provides a carotid artery-assisted diagnostic method based on ultrasound visual language, comprising the following steps: Using the constructed carotid artery ultrasound multi-vessel segment image and text dataset, a carotid artery diagnostic knowledge graph was built, and after embedding the knowledge graph using TransE, a patient-level vascular heterogeneous instance graph was constructed. Based on patient-level vascular heterogeneous instance maps, the HGT Heterogeneous Graph Transformer is used for relational modeling. After performing image-text-atlas trimodal pre-training, the transfer of downstream tasks of carotid ultrasound is performed to generate interpretable diagnostic paths for carotid artery assisted diagnosis.
[0005] Preferably, when constructing the carotid artery ultrasound multi-vessel segment image and text dataset, the original DICOM video stream, B-mode keyframes, color Doppler keyframes, spectral Doppler images, doctor's diagnostic reports, structured measurements, and expert annotation information are collected. Using the patient's unique identifier as the primary key, all images and reports from the same patient and the same examination are bound as patient-level samples, and the left and right sides, vessel segments, section types, acquisition timestamps, image modes, and device sources are recorded to construct the carotid artery ultrasound multi-vessel segment image and text dataset.
[0006] Preferably, when constructing a multi-segment image dataset of carotid artery ultrasound, for the video stream, a joint screening of inter-frame differences and image sharpness is performed to remove redundant similar frames and obviously blurred frames.
[0007] Preferably, when constructing the carotid artery diagnostic knowledge graph, anatomical structure nodes, left and right side nodes, section nodes, image pattern nodes, lesion nodes, ultrasound sign nodes, measurement parameter nodes, risk level nodes, and diagnostic conclusion nodes are used as entity nodes of the graph; belonging to, located at, corresponding section, manifested as, supported by parameters, indicating risk, supporting diagnosis, and follow-up changes are used as relationship types of the graph to construct the carotid artery diagnostic knowledge graph.
[0008] Preferably, when using TransE for knowledge graph embedding, the TransE model is used to vectorize the entity nodes and relation edges in the knowledge graph. For any triple, TransE is used to constrain the head entity vector, relation vector, and tail entity vector.
[0009] Preferably, when constructing a patient-level heterogeneous vascular instance graph, the patient-level heterogeneous vascular instance graph is constructed based on the patient's node set, internal relation edge set, node type and edge type set, and node feature set. The node types include patient nodes, left and right side nodes, vascular segment nodes, cross-section nodes, keyframe nodes, lesion nodes, attribute nodes, measurement parameter nodes, and diagnostic conclusion candidate nodes.
[0010] Preferably, when performing relationship modeling, a four-layer HGT heterogeneous graph Transformer is used to perform relationship modeling on the patient-level vascular instance graph. HGT is used to set independent projection matrices for different node types and different relationship types to handle the heterogeneous relationships between patient nodes, vascular segment nodes, keyframe nodes, lesion nodes, attribute nodes and measurement parameter nodes.
[0011] Preferably, during trimodal pre-training, the DINOv2 ViT-B / 14 visual encoder, Chinese-MedBERT-base text encoder, TransE knowledge graph embedding, and HGT patient-level vascular graph neural network are used in combination for trimodal pre-training to map image representation, text representation, knowledge graph entity representation, and patient-level graph representation to a unified 512-dimensional common space.
[0012] Preferably, when performing the downstream task transfer of carotid ultrasound, the visual encoder, TransE knowledge embedding and HGT graph neural network are transferred to the downstream task of carotid ultrasound, wherein, for image-level tasks, joint prediction is performed using keyframe image representation, its corresponding vessel segment representation and patient-level graph representation. For the segmentation task, the multi-layer visual features output by DINOv2 ViT-B / 14 are used to construct the FPN decoder, which outputs patch segmentation masks and intima-media complex segmentation masks. For patient-level multi-segment lesion burden assessment, attention-weighted convergence of the node representations of each segment on the left and right sides is performed to obtain the patient-level lesion burden vector. In the plaque vulnerability prediction task, lesion node representation, corresponding keyframe visual representation, plaque segmentation morphological features and knowledge graph attributes are embedded into the common input vulnerability task head. The task head is used to output the probability of stable plaques, suspected unstable plaques and high-risk vulnerable plaques, and at the same time output the main basis nodes. In the stenosis grading task, the morphological stenosis rate, lumen diameter change and spectral Doppler velocity parameters are input into the stenosis grading task header. If the image morphology indicates mild stenosis but the spectral velocity indicates moderate stenosis, it is marked by the conflict edge in the patient-level instance image, and the doctor is prompted to review it in conjunction with the original spectral image during the report generation stage.
[0013] Based on the same inventive concept, this invention also discloses a carotid artery-assisted diagnostic system based on ultrasound visual language, comprising: The front-end preprocessing module is used to construct a carotid artery diagnostic knowledge graph using the constructed carotid ultrasound multi-vessel segment image and text dataset, and then construct a patient-level vascular heterogeneous instance graph after embedding the knowledge graph using TransE. The downstream auxiliary diagnostic module is used to perform relationship modeling based on patient-level vascular heterogeneous instance maps using HGT Heterogeneous Graph Transformer. After performing image-text-atlas trimodal pre-training, it performs transfer of carotid ultrasound downstream tasks to generate interpretable diagnostic paths for carotid artery auxiliary diagnosis.
[0014] The present invention discloses the following technical effects: First, this invention upgrades carotid ultrasound diagnostic knowledge from natural language description to a computable knowledge graph, enabling the medical relationships between anatomical location, lesion attributes, blood flow parameters, and diagnostic conclusions to explicitly participate in model training.
[0015] Second, the present invention uses a single carotid ultrasound examination of the same patient as a patient-level heterogeneous instance image, enabling the model to learn the correlation between multiple vascular segments, multiple cross-sections and multiple lesions.
[0016] Third, this invention improves the model's ability to understand left-right differences, vascular segment distribution, plaque multiplication, and the relationship between local lesions and overall lesion burden by using the HGT Heterogeneous Graph Transformer to perform relationship type-aware message transmission in patient-level vascular maps.
[0017] Fourth, this invention introduces semantic soft label loss based on knowledge graph distance, enabling the model to distinguish between completely irrelevant negative samples and medically semantically similar samples, thus alleviating the problem of semantically similar labels being mistakenly pushed away in traditional contrastive learning.
[0018] Fifth, this invention can output a diagnostic evidence path consisting of vascular segments, lesions, signs, parameters, and conclusions, which facilitates review by clinicians. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram illustrating the execution logic of the method described in this invention.
[0021] Figure 2 This is a schematic diagram of the carotid artery diagnostic knowledge graph construction process described in this invention.
[0022] Figure 3 This is a schematic diagram of an example of patient-level vascular heterogeneity described in this invention.
[0023] Figure 4 This is a schematic diagram of the image-text-graph three-modal pre-training process described in this invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0025] like Figure 1 As shown, this invention provides a carotid artery-assisted diagnosis method based on ultrasound visual language. First, a multi-segment image-text dataset containing carotid artery ultrasound images, video keyframes, report text, structured measurements, and expert annotations is constructed. Second, a carotid artery diagnostic knowledge graph is built, representing anatomical structures, scanning sections, image patterns, lesion types, ultrasound signs, quantitative parameters, risk levels, and diagnostic conclusions as entity nodes and relational edges. Then, a patient-level heterogeneous instance graph is constructed for each patient, unifying the left and right sides, vessel segments, sections, keyframes, lesions, attributes, and measurement parameters within the patient into a graph structure. TransE is then used to obtain entity and relational embeddings from the knowledge graph, and HGT heterogeneous graph Transformer is used to perform message passing on the patient-level instance graph, obtaining lesion-level, vessel segment-level, and patient-level representations. Finally, a trimodal pre-training is performed using a DINOv2ViT-B / 14 visual encoder, a Chinese-MedBERT-base text encoder, TransE knowledge embedding, and an HGT graph neural network, and the pre-trained model is transferred to downstream carotid artery ultrasound diagnostic tasks.
[0026] Step S100: Construct a multi-segment image and text dataset of carotid artery ultrasound.
[0027] In step S100 of one embodiment, this step collects carotid ultrasound examination data from multiple medical institutions.
[0028] In one embodiment, the carotid ultrasound examination data includes: raw DICOM video stream, B-mode keyframes, color Doppler keyframes, spectral Doppler images, physician diagnostic reports, structured measurements, and expert annotation information.
[0029] For example, the present invention uses a patient's unique identifier as the primary key to bind all images and reports of the same patient in the same examination as a patient-level sample, and records the left and right sides, blood vessel segments, section types, acquisition timestamps, image modes and device sources.
[0030] For example, for a raw DICOM video stream, the image matrix, pixel pitch, frame rate, acquisition time, probe frequency, and device manufacturer information are read first.
[0031] For example, in order for subsequent models to be able to handle different devices and different imaging parameters simultaneously, the pixel grayscale is normalized to a fixed range and the original physical scale of the image is preserved so that the thickness of the inner membrane, plaque thickness, plaque area and lumen stenosis rate can be automatically calculated later.
[0032] In one implementation, for a video stream, a joint screening of inter-frame differences and image sharpness is performed to remove redundant similar frames and obviously blurry frames.
[0033] For example, let the image of the t-th frame be I. t The image height and width are H and W respectively, and the inter-frame difference score is D. t Defined as: ; Image sharpness score Q t The variance of the response is expressed using Laplace's method: Where Lap represents the Laplace operator and Var represents the variance.
[0034] For example, to comprehensively consider inter-frame variations, sharpness, and the visibility of blood vessel regions, the system further calculates a keyframe candidate score R. t : Among them, D t and Q t The cap at the top indicates the normalization result, C t The visibility score of the vascular region is represented by α, β and δ, which are weighting coefficients. In this embodiment, α is set to 0.35, β is set to 0.35 and δ is set to 0.30.
[0035] For example, if R t Greater than the keyframe threshold If the time interval between adjacent selected keyframes is greater than the minimum interval threshold, then frame t is saved as a keyframe. .
[0036] In one embodiment, the candidate keyframes are further processed to locate blood vessel regions and desensitize images, removing privacy information such as names, examination numbers, and equipment numbers.
[0037] For example, the vascular region localization uses the YOLOv8n detection network. After detecting the effective imaging window of the carotid artery, it is cropped and scaled to an input image of 224 pixels by 224 pixels to adapt to the DINOv2 ViT-B / 14 visual encoder.
[0038] For example, for spectral Doppler images, velocity scale, sampling gate position, and angle correction information are additionally retained for subsequent construction of blood flow parameter nodes.
[0039] For example, the expert annotations include left and right sides, common carotid artery, carotid sinus, internal carotid artery, external carotid artery, long axis section, short axis section, color Doppler, spectral Doppler, presence or absence of plaque, plaque location, plaque echo, surface regularity, calcification or acoustic shadowing, ulceration signs, intimal-media thickening, degree of luminal stenosis, and plaque vulnerability level.
[0040] For example, structured measurements include intima-media thickness, plaque thickness, plaque area, lumen diameter, peak systolic velocity, end-diastolic velocity, and stenosis rate.
[0041] In one embodiment, to ensure the consistency of data from multiple institutions, an examination integrity record is established for each case. This record includes whether both the left and right sides have been scanned, whether the common carotid artery, carotid sinus, internal carotid artery, and external carotid artery all have available keyframes, whether both the long axis and short axis can be used to locate the lesion, and whether there are resolvable parameters for color Doppler and spectral Doppler.
[0042] For example, if a blood vessel segment lacks the necessary cross-section, the segment is marked as having incomplete information instead of being simply classified as a negative sample, thereby avoiding misclassification of uncollected areas as normal blood vessel segments during the training phase.
[0043] In one embodiment, for repeated occurrences of the same lesion in multiple sections, lesion matching is performed using the location of the vascular segment, left and right sides, relative anatomical location, plaque thickness, and report description.
[0044] For example, if multiple keyframes are identified as different viewing sections of the same lesion, only one lesion node is generated in the patient-level instance map, and the multiple keyframe nodes are connected to this lesion node. This operation reduces the problem of the same plaque being counted repeatedly and enables the subsequent graph neural network to learn complementary evidence of the same lesion under different sections.
[0045] In one implementation, the original report is split into segments such as examination findings, measurement descriptions, diagnostic suggestions, and follow-up recommendations, and the left and right sides, vessel segments, lesion attributes, and numerical parameters are identified through a rule dictionary.
[0046] For example, if the report contains a description such as "mixed echo plaque visible on the posterior wall of the left carotid sinus", it is parsed into structured fields such as left side, carotid sinus, posterior wall, and mixed echo plaque, and a weak pairing relationship is established with the corresponding image keyframe.
[0047] For example, if a report contains only patient-level conclusions and lacks single-frame descriptions, the text is retained for patient-report comparison pre-training.
[0048] For example, to prevent data leakage, the present invention divides the training set, validation set and test set according to the patient's unique identifier, ensuring that all images, reports and graph structures of the same patient exist in only the same data subset.
[0049] For example, for cross-institutional validation, the system also establishes independent external test sets according to hospital origin and equipment manufacturer to evaluate the robustness of the model under different devices, different operators and different image styles.
[0050] Step S200: Construct a knowledge graph for carotid artery diagnosis.
[0051] In step S200 of one embodiment, this step constructs a static carotid artery diagnostic knowledge graph to represent medical priors in carotid artery ultrasound diagnosis.
[0052] In one implementation, the knowledge graph is represented as: Among them, V K E represents the set of entity nodes. K Let R represent the set of relation edges. K Represents a set of relation types.
[0053] For example, entity nodes include anatomical structure nodes, left and right side nodes, cross-sectional nodes, image pattern nodes, lesion nodes, ultrasound sign nodes, measurement parameter nodes, risk level nodes, and diagnostic conclusion nodes.
[0054] For example, relationship types include belonging, located, corresponding aspect, manifest as, supported by parameters, indicating risk, supporting diagnosis and follow-up changes, etc.
[0055] In one implementation, each fact in the knowledge graph is represented as a triple: For example, a ternary set can be constructed such as "plaque-located-carotid sinus", "hypoechoic-indicating-risk-increased vulnerability", "irregular surface-supporting diagnosis-unstable plaque", and "increased peak systolic velocity-supporting diagnosis-moderate to severe stenosis".
[0056] For example, for anatomical structures, the knowledge graph expresses the subordinate relationships between the common carotid artery, carotid sinus, internal carotid artery, and external carotid artery.
[0057] For example, for lesion attributes, the knowledge graph expresses the relationship between echo type, calcification, acoustic shadowing, surface morphology, and ulceration signs and plaque properties.
[0058] For example, for quantitative parameters, the knowledge graph expresses the relationship between intima-media thickness, plaque thickness, stenosis rate, and blood flow velocity and the diagnostic conclusion.
[0059] In one embodiment, the present invention employs a rule dictionary, a Chinese-MedBERT-base named entity recognition model, and Qwen3-8B-Instruct to assist in extracting entities and relationships from historical reports.
[0060] In one embodiment, after entity extraction, the present invention maps descriptions such as "carotid bifurcation", "carotid sinus", and "bulbar" to a unified entity according to the synonym normalization rule, and normalizes descriptions such as "mixed echo", "mixed echo", and "heterogeneous echo" into standard attribute nodes in the knowledge graph.
[0061] In one implementation, after relation extraction, two senior ultrasound physicians perform manual correction to form a stable carotid artery diagnostic knowledge graph.
[0062] In one embodiment, to enable the knowledge graph to support the construction of semantic soft tags, the present invention assigns relation weights to different relation types.
[0063] For example, the lower weight of the anatomical subordinate relationship indicates that the nodes within the same anatomical level are semantically close; the higher weight of the risk warning relationship and the diagnostic conclusion support relationship indicates that they represent a semantic jump from the attribute to the diagnostic conclusion.
[0064] For example, the relation weights are initialized by expert experience and can be fine-tuned based on the frequency of co-occurrence of entities in the training set.
[0065] In one embodiment, in the hierarchical design of the knowledge graph, the present invention divides entities into a basic anatomy layer, a data acquisition section layer, an image sign layer, a measurement parameter layer, a diagnostic conclusion layer, and a risk interpretation layer.
[0066] For example, the basic anatomical layer is used to express the spatial relationships between the left and right sides and each vascular segment.
[0067] For example, the acquisition section layer is used to express the different roles of long axis, short axis, color Doppler, and spectral Doppler in diagnosis.
[0068] For example, the image feature layer is used to express plaque echo, surface morphology, calcification acoustic shadowing, and ulceration features.
[0069] For example, the measurement parameter layer is used to express intima-media thickness, plaque thickness, stenosis rate, and blood flow velocity.
[0070] For example, the diagnostic conclusion layer is used to express the presence of plaque, the degree of stenosis, and the level of vulnerability.
[0071] For example, the risk interpretation layer is used to express risk warnings and follow-up recommendations for cerebrovascular events.
[0072] In one embodiment, to avoid the knowledge graph from relying too heavily on a single text source, the present invention cross-validates expert rules, structured annotations, historical reports, and model extraction results.
[0073] For example, if an entity relationship is only extracted by a large language model and is not supported by rules or expert annotations, this invention marks it as a relationship to be reviewed; if a relationship is supported by multiple reports and expert rules at the same time, it is given higher credibility.
[0074] For example, during pre-training, relations with higher credibility participate in TransE training and semantic soft label construction, while relations with lower credibility are only used as candidate paths and are not directly used as strong supervision signals.
[0075] Step S300 uses TransE to embed the knowledge graph.
[0076] In step S300 of one embodiment, the TransE model is used to vectorize entity nodes and relation edges in the knowledge graph.
[0077] For example, for any triple (h,r,t), the TransE constraint head entity vector, relation vector, and tail entity vector satisfy the following relationship: Among them, e h The embedding vector of the head entity h, e r Let e be the embedding vector of relation r. t This represents the embedding vector of the tail entity t.
[0078] For example, the triplet scoring function is defined as: ; For example, during training, an interval sorting loss is applied to the set of positive triples T and the set of negative triples T': .
[0079] Wherein, γ is a preset positive interval hyperparameter, used to specify the minimum interval between the negative triplet score and the positive triplet score. In this embodiment, γ is taken as 1.0.
[0080] For example, negative samples are generated by randomly replacing the head or tail entity, but if the replaced triple already exists in the knowledge graph, it is not considered a negative sample.
[0081] In one embodiment, to prevent the model from excessively widening medically similar nodes, the present invention introduces type constraints when generating negative samples, that is, anatomical structure nodes are only replaced with nodes of the same type of anatomical structure, and lesion attribute nodes are only replaced with nodes of the same type of attribute, thereby improving training stability.
[0082] In one implementation, the entity embedding dimension is set to 256, the relation embedding dimension is set to 256, the number of TransE training rounds is set to 500, the optimizer is Adam, and the initial learning rate is set to 1e-3.
[0083] In one implementation, after training is completed, knowledge graph entity embedding is used to initialize vascular segment nodes, lesion nodes, attribute nodes, and diagnostic conclusion candidate nodes in the patient-level instance graph; relation embedding is used to initialize relation type parameters in the HGT heterogeneous graph Transformer.
[0084] Step S400: Construct a patient-level vascular heterogeneity instance diagram.
[0085] In step S400 of one embodiment, this step dynamically constructs a patient-level vascular heterogeneity instance map for each patient.
[0086] In one implementation, for the i-th patient, its patient-level instance graph is represented as follows: Among them, V i E represents the set of nodes for this patient. i T represents the set of relation edges within the patient. i X represents the set of node types and edge types. i Represents the set of node features.
[0087] For example, node types include patient nodes, left and right side nodes, blood vessel segment nodes, section nodes, keyframe nodes, lesion nodes, attribute nodes, measurement parameter nodes, and diagnostic conclusion candidate nodes.
[0088] Specifically, the patient node is connected to the left and right nodes; the left and right nodes are connected to the vascular segment nodes of the common carotid artery, carotid sinus, internal carotid artery and external carotid artery, respectively.
[0089] Specifically, the vascular segment nodes are connected to the long axis section, short axis section, color Doppler section, and spectral Doppler section.
[0090] Specifically, the aspect node is connected to the corresponding keyframe node.
[0091] Specifically, the vascular segment nodes are connected to the lesion nodes.
[0092] Specifically, lesion nodes are connected to attribute nodes such as echogenicity, calcification, acoustic shadowing, surface regularity, ulceration signs, and degree of stenosis.
[0093] Specifically, lesion nodes or vascular segment nodes are connected to measurement parameter nodes.
[0094] In one embodiment, for vascular segments where no lesions are found, the present invention still retains the vascular segment node, section node, and keyframe node, and connects them to normal attribute nodes such as "no obvious plaques found" and "no obvious stenosis found in the lumen." Through the above operations, the model can avoid learning the graph structure only on positive lesions and ignoring the contribution of normal vascular segments to the overall patient-level judgment.
[0095] In one embodiment, during the graph construction process, a skeleton graph is first established based on the left and right sides and vessel segment labels. Then, image nodes are attached to the corresponding vessel segment nodes according to the keyframe section type and timestamp. Subsequently, the system reads the segmentation mask and classification results. If a plaque or intimal thickening exists in a certain image, a lesion node is generated for the corresponding vessel segment. If lesions from multiple sections meet the conditions of being on the same side, in the same vessel segment, in similar location, and in similar thickness, they are merged into the same lesion node. The merged lesion node stores visual evidence from multiple sections, and can simultaneously absorb long axis morphological information and short axis area information during subsequent HGT message transmission.
[0096] In one embodiment, for spectral Doppler images, the present invention does not simply treat them as ordinary image nodes, but instead generates additional blood flow parameter nodes.
[0097] For example, the blood flow parameter node includes information such as peak systolic velocity, end-diastolic velocity, resistance index, peak systolic ratio, and angle correction status. This node is connected to the corresponding vessel segment node and the stenosis degree candidate node, enabling the model to simultaneously refer to morphological and hemodynamic evidence when determining the degree of stenosis.
[0098] In one implementation, the visual representation of the keyframe node is extracted by DINOv2 ViT-B / 14.
[0099] For example, let the keyframe image be I. j Its visual representation is: Among them, f img This refers to the DINOv2 ViT-B / 14 vision encoder. This represents the parameters of the visual encoder.
[0100] For example, the semantic representation of the report text fragment is extracted from the Chinese-MedBERT-base: .
[0101] In one embodiment, for the structured measurement value p m First, numerical normalization is performed, and then the node features of the measurement parameters are obtained through linear mapping: .
[0102] For any node v, its initial node characteristics are defined as: Among them, z v For visual features, u v For text semantic features, q v For the characteristics of the measured value, k v For TransE knowledge graph entity embedding, e type(v) For node type embedding, e pos(v) The location of the blood vessel segment is embedded, and the symbol || indicates feature splicing.
[0103] For example, if a node does not have a certain type of feature, the corresponding vector is set to zero, and its origin is distinguished by node type embedding and position embedding.
[0104] Step S500 uses HGT Heterogeneous Graph Transformer to model relationships.
[0105] In step S500 of one embodiment, this step uses a four-layer HGT heterogeneous graph Transformer to model the relationships of patient-level vascular instance graphs.
[0106] For example, HGT sets up independent projection matrices for different node types and different relationship types, which can handle heterogeneous relationships between patient nodes, blood vessel segment nodes, keyframe nodes, lesion nodes, attribute nodes and measurement parameter nodes.
[0107] In one implementation, for the l-th layer HGT, node v receives information from its neighbor node u under relation r, and its attention weight is defined as: Where τ is the node type mapping function, τ(v) represents the node type of the target node v, and τ(u) represents the node type of the source node u. Node types include patient nodes, left and right side nodes, blood vessel segment nodes, section nodes, keyframe nodes, lesion nodes, attribute nodes, measurement parameter nodes, and diagnostic conclusion candidate nodes. Qτ(v) represents the query projection matrix corresponding to the target node type τ(v), Kτ(u),r represents the key projection matrix jointly determined by the source node type τ(u) and the relation type r, and d represents the feature dimension of a single attention head. In this embodiment, the hidden dimension is 768 and the number of attention heads is 8, therefore d is taken as 96.
[0108] For example, the message aggregation result of node v is represented as: .
[0109] For example, the node feature update process is represented as follows: in, This represents a value matrix corresponding to the source node type and relation type. This represents the set of neighbors connected to node v through relation r, LN represents layer normalization, FFN represents feedforward network, where HGT hidden dimension is set to 768, attention head number is set to 8, number of layers is set to 4, and dropout is set to 0.1.
[0110] In one embodiment, after HGT encoding, the present invention extracts lesion node representation, vascular segment node representation and patient node representation respectively.
[0111] For example, lesion node representation is used for plaque attribute identification and vulnerability prediction; vessel segment node representation is used for segmental stenosis grading and lesion burden assessment; and patient node representation is used for patient-level report alignment and multi-segment lesion burden assessment.
[0112] For example, through this hierarchical representation, the model is able to simultaneously preserve information about local lesions and the patient's overall vascular status.
[0113] Step S600 Image-Text-Graph Three-Modal Pre-training.
[0114] In step S600 of one embodiment, this step combines the DINOv2 ViT-B / 14 visual encoder, the Chinese-MedBERT-base text encoder, the TransE knowledge graph embedding, and the HGT patient-level vascular graph neural network for trimodal pre-training.
[0115] In one embodiment, the present invention maps image representations, text representations, knowledge graph entity representations, and patient-level graph representations to a unified 512-dimensional public space.
[0116] For example, for the i-th image sample, its image projection vector is: ; Its corresponding text projection vector is: ; Image-text contrast loss is expressed in the form of InfoNCE: Where sim represents cosine similarity. The value represents the temperature coefficient, and N represents the number of samples in the training batch.
[0117] In one embodiment, for the knowledge graph node set C corresponding to the image sample i The system embeds its entities into average pooling to obtain the graph semantic vector c. i : .
[0118] For example, the image-map contrast loss is defined as: .
[0119] For example, for the i-th patient, HGT outputs a patient-level vascular representation g. i The textual representation of a complete diagnostic report is r i The patient image-report contrast loss is defined as: Where M represents the number of patients in the patient-level training batch.
[0120] For example, this loss aligns multi-vessel instance images of the same patient with the complete diagnostic report in the semantic space, avoiding semantic mismatches caused by inconsistencies in granularity between individual images and the complete report.
[0121] In one implementation, to enable the model to learn medical semantic proximity relationships in the knowledge graph, the system constructs soft labels based on the path distance in the knowledge graph.
[0122] For example, the weighted path distance between any two diagnostic labels or attribute nodes a and b is defined as: Where P represents the path from node a to node b, w r This represents the weight of relation r. Semantic similarity is obtained based on this distance: ; Normalize all candidate labels within a batch to obtain the soft label distribution: ; The model's predicted distribution is denoted as Q. ab The semantic soft label loss is defined as: .
[0123] In one embodiment, the present invention introduces an edge relationship prediction task in a patient instance graph.
[0124] For example, for nodes u and v and relation r, the probability of edge existence is expressed as: ; The relationship prediction loss is expressed as a binary cross-entropy form: ; In summary, the total loss function for trimodal pre-training is: in, to For the loss weights, where, Set it to 0.5. Set to 1.0. Set it to 0.5. Set it to 0.2. Set it to 0.3.
[0125] For example, the optimizer is trained using AdamW with an initial learning rate of 1e. -4 The weight decay is 0.05, the batch size is 128, and the number of pre-training rounds is 100.
[0126] In one embodiment, regarding the training strategy, the present invention first freezes the first 8 Transformer blocks of DINOv2 ViT-B / 14, and only fine-tunes the subsequent blocks, projection layers, HGT, and text-side projection layers; in the second half of the pre-training process, all visual encoder parameters are gradually unfrozen to avoid excessive damage to basic visual features caused by small-scale medical data.
[0127] For example, the Chinese-MedBERT-base text encoder uses a small learning rate for updates to maintain the semantic stability of medical text.
[0128] In one embodiment, in order to improve applicability in low-annotation scenarios, the present invention uses both expert-annotated samples and weakly annotated samples consisting only of report text during pre-training.
[0129] For example, for samples without pixel-level lesion masks, the present invention can still construct knowledge graph nodes and patient-level instance graphs based on the reported entity extraction results; for samples with pixel-level annotations, additional visual features of the lesion region and segmentation auxiliary loss are added.
[0130] In one embodiment, the present invention employs a patient-level sampling strategy instead of single-frame-level random sampling in terms of batch construction.
[0131] For example, each training batch contains several patients, and a fixed number of keyframes are randomly selected from each patient, while retaining their complete patient-level instance graph.
[0132] For example, if the number of keyframes for a patient exceeds the upper limit, the present invention samples according to the principles of prioritizing vascular segment coverage, prioritizing lesion positivity, and prioritizing image quality; if the number of keyframes is less than the upper limit, mask nodes are used to fill in the gaps.
[0133] For example, in this way, the model can see the multi-segment vascular structure inside the patient in each training session, instead of just seeing independent images.
[0134] In one embodiment, regarding cross-modal alignment, the present invention sets up three granularities of positive sample relationships, wherein the first is a local positive sample relationship between a keyframe image and its corresponding report fragment; the second is a lesion-level positive sample relationship between a lesion node and its corresponding attribute text; and the third is a patient-level positive sample relationship between a patient-level instance image and a complete diagnostic report.
[0135] For example, training positive sample relationships at different granularities together can reduce the problem that a single frame image cannot fully cover the semantics of the entire report.
[0136] In one embodiment, in order to improve the model's ability to process semantically similar labels, the present invention no longer equates all unpaired samples with strong negative samples.
[0137] For example, for labels that are close in the knowledge graph, such as mild stenosis and moderate stenosis, hypoechoic plaques and vulnerable plaques, the present invention provides soft supervision of intermediate intensity; for labels that are far in the knowledge graph, such as normal intima and severe stenosis, no plaques and ulcerative plaques, stronger negative sample constraints are provided, so that the common feature space can be more in line with the logic of clinical diagnosis.
[0138] Step S700 map enhancement for downstream task prediction. In step S700 of one embodiment, after pre-training is completed, the present invention transfers the visual encoder, TransE knowledge embedding, and HGT graph neural network to the downstream task of carotid ultrasound.
[0139] For example, for image-level tasks, the present invention uses keyframe images to represent z. i Characteristic b of its corresponding vascular segment i and patient-level graph representation g i Conduct joint forecasting: Where k represents the specific task. The cap number above represents the predicted probability of the k-th task.
[0140] For example, the classification tasks include long and short axis section recognition, vessel segment recognition, image pattern recognition, plaque presence or absence recognition, plaque location determination, intimal-media thickening determination, stenosis degree grading, plaque vulnerability prediction, and patient-level multi-vessel segment lesion burden assessment.
[0141] In one embodiment, for the segmentation task, the present invention uses the multilayer visual features output by DINOv2 ViT-B / 14 to construct an FPN decoder, which outputs a patch segmentation mask and an inner membrane mid-layer complex segmentation mask.
[0142] For example, the segmentation loss consists of the Dice loss and the cross-entropy loss: Where Y represents the manually labeled mask, and the cap above Y represents the model prediction mask. Based on the segmentation mask and the physical scale of the image, this invention further calculates quantitative indicators such as plaque thickness, plaque area, intima-media thickness, and stenosis rate.
[0143] In one embodiment, for patient-level multi-segmental lesion burden assessment, the present invention performs attention-weighted convergence on the node representations of each segment on both sides to obtain a patient-level lesion burden vector. This vector is used to predict whether the patient has multi-segmental lesions, the number of segments involved, whether the lesions are symmetrical on both sides, and the overall atherosclerotic burden level. Among them, B i This represents the set of vascular segment nodes for the i-th patient. The attention weight of the vessel segment node v is represented by the design that enables the model to automatically focus on vessel segments with more obvious lesions and higher diagnostic value, while retaining the constraint effect of normal vessel segments on the overall lesion load assessment.
[0144] In one embodiment, in the plaque vulnerability prediction task, the present invention embeds lesion node representation, corresponding keyframe visual representation, plaque segmentation morphological features, and knowledge graph attributes into a common input vulnerability task header.
[0145] For example, the task header outputs the probabilities of stable plaques, suspected unstable plaques, and high-risk vulnerable plaques, along with the main evidence nodes. If the model primarily relies on low echogenicity, surface irregularities, and ulceration signs to make judgments, these nodes will receive higher weight in subsequent evidence path generation.
[0146] In one embodiment, in a stenosis severity grading task, the present invention inputs morphological stenosis rate, lumen diameter variation, and spectral Doppler velocity parameters into the stenosis severity task head.
[0147] For example, if the image morphology suggests mild stenosis but the spectral velocity suggests moderate stenosis, the present invention marks the conflicting edges in the patient-level instance image and prompts the doctor to review it in conjunction with the original spectrogram during the report generation stage. This design avoids the model from directly outputting overly certain conclusions when multimodal evidence is inconsistent.
[0148] Step S800 can explain the diagnostic path generation.
[0149] In step S800 of one embodiment, this step retrieves diagnostic evidence paths in the carotid artery diagnostic knowledge graph based on the model prediction results.
[0150] In one embodiment, for a predicted conclusion y, the present invention starts from the activated vascular segment nodes, lesion nodes, attribute nodes, and measurement parameter nodes in the patient-level instance graph and retrieves candidate paths leading to the conclusion node y.
[0151] For example, the scoring function for candidate path P is defined as: Where conf(e) represents the model confidence corresponding to the edge or node in the path. This represents the weight of the evidence type, and len(P) represents the path length. The path length penalty coefficient is represented, and in this invention, the highest-scoring paths are selected as the diagnostic criteria.
[0152] In one embodiment, when the model determines that the plaque in the left carotid sinus has a vulnerability risk, the present invention can output the following evidence path: left side - carotid sinus - plaque - hypoechoic - irregular surface - local stenosis - increased vulnerability risk. This path can be input into Qwen3-8B-Instruct along with the model's predicted probability and quantitative measurement value, which will generate an explanatory text that conforms to clinical reporting conventions.
[0153] In one implementation, the large language model is used only for evidence path transcription, duplicate information merging, and report language generation, and does not directly replace the task model for original diagnostic prediction.
[0154] For example, this invention inputs a structured context into a large language model, including patient-level prediction results, prediction results for each vascular segment, plaque segmentation measurements, knowledge graph paths, and model confidence. The large language model is then required to output structured text containing "ultrasound findings," "ultrasound suggestions," and "AI evidence path." If the task model confidence is below a threshold, the system generates a prompt in the report suggesting that "the doctor review or supplement the relevant cross-sections."
[0155] In one implementation, to prevent inconsistencies between the explanatory text and model evidence, this invention performs structured validation on the output of the large language model. Validation includes verifying whether the vascular segment mentioned in the report exists in the patient-level instance image, whether plaque attributes originate from model predictions or expert annotations, whether quantitative values are consistent with measurement parameter nodes, and whether diagnostic prompts can be found in the knowledge graph through at least one evidence path. If validation fails, the system does not directly output the text but instead reorganizes the prompt words, requiring the large language model to regenerate it based on the given evidence.
[0156] In one embodiment, the present invention also retains the node confidence and edge confidence of each diagnostic path.
[0157] For example, when a doctor views a report on the front end, he can click on a conclusion, and the system will automatically highlight the corresponding original keyframe, patch segmentation region, and knowledge graph path.
[0158] For example, in this way, the output of the present invention is no longer just a black box classification label, but a chain of diagnostic evidence that can be traced back to a specific blood vessel segment, specific section, specific lesion, and specific sign.
[0159] Step S900 algorithm deployment, training parameters and testing process.
[0160] In step S900 of one embodiment, the present invention can be deployed on a hospital intranet server or a local workstation.
[0161] For example, the backend uses the Python Flask framework to encapsulate interfaces for DICOM parsing, keyframe extraction, image encoding, text encoding, knowledge graph querying, HGT graph neural network inference, downstream task prediction, and report generation.
[0162] For example, the front end supports importing carotid ultrasound DICOM data, browsing keyframes, displaying structured vascular segments, displaying plaque and intimal media segmentation masks, displaying diagnostic paths, and secondary editing of reports.
[0163] For example, the training phase is implemented using the PyTorch framework.
[0164] For example, the DINOv2 ViT-B / 14 input image resolution is 224 pixels by 224 pixels, and the patch size is 14; the maximum text length of Chinese-MedBERT-base is set to 256 tokens; the TransE entity and relation embedding dimension is set to 256; the HGT hidden layer dimension is set to 768, the number of attention heads is set to 8, and the number of layers is set to 4.
[0165] For example, training was performed using eight NVIDIA A800 graphics cards, with mixed-precision training enabled and the gradient clipping threshold set to 1.0.
[0166] For example, during the testing phase, various downstream tasks are evaluated on an independent test set. Classification tasks are evaluated using accuracy, AUC, weighted F1, sensitivity, and specificity; segmentation tasks are evaluated using Dice coefficient, IoU, and HD95; patient-level report generation is evaluated using structured field accuracy, diagnostic concordance rate, and clinical expert ratings; and interpretable pathways are evaluated using pathway hit rate, pathway integrity score, and expert readability score.
[0167] For example, to verify the effectiveness of the knowledge graph and patient-level graph neural network, an ablation experiment was set up. The first group removed the knowledge graph embeddings, retaining only the image-text comparison pre-training; the second group retained the knowledge graph but removed the patient-level instance graphs, using only single-frame images for prediction; the third group retained the patient-level instance graphs but removed the semantic soft-label loss; and the fourth group used the complete system. The contributions of the knowledge graph, HGT, and semantic soft-label loss in this invention were verified by comparing the performance of each group in plaque vulnerability prediction, multi-segment lesion burden assessment, and reporting consistency.
[0168] For example, the present invention also conducts cross-device robustness tests, using device manufacturers or probe models not appearing in the training set as external test sources to evaluate the performance differences between the basic visual language model and the complete system of the present invention. In this invention, because the present invention explicitly preserves vascular segments, lesion attributes and quantitative measurement parameters in patient-level instance images, the model's dependence on the image style of a single device is reduced, and it can maintain more stable diagnostic performance in cross-device scenarios.
[0169] For example, this invention further incorporates physician-interactive verification, where the same batch of cases is independently interpreted by the system, junior physicians, and senior physicians, before an expert panel reaches a consensus. Evaluation metrics include diagnostic accuracy, report completeness, lesion localization consistency, and average image reading time. If the evidence path output by the system aligns with the expert panel's consensus path, it is considered a path hit; if the diagnostic conclusion is correct but the path omits crucial evidence, it is considered a partial hit. This evaluation method can simultaneously measure the model's diagnostic accuracy and interpretive reliability.
[0170] In one embodiment, in terms of security control, the present invention sets up a confidence threshold, a missing section prompt, and an abnormal input detection mechanism.
[0171] For example, when the input image quality is too low, key vascular segments are missing, spectral Doppler angle correction is abnormal, or the model prediction distribution is too scattered, a definitive diagnosis is not output, but a prompt for review is output. Through this mechanism, the present invention is more in line with real clinical workflow and avoids generating overly definitive automatic reports when there is insufficient evidence.
[0172] For example, the final output of this invention includes patient-level carotid ultrasound diagnostic conclusions, lesion distribution on the left and right sides and in each vessel segment, plaque attributes, degree of stenosis, plaque vulnerability level, patient-level multi-vessel segment lesion burden, model confidence, knowledge graph evidence path, and structured diagnostic report file.
[0173] In summary, through the above steps, this invention realizes an end-to-end intelligent processing flow from carotid ultrasound images, report text, and medical knowledge graphs to patient-level interpretable diagnostic results.
[0174] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0175] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0176] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A carotid artery-assisted diagnostic method based on ultrasound visual language, characterized in that, Includes the following steps: Using the constructed carotid artery ultrasound multi-vessel segment image and text dataset, a carotid artery diagnostic knowledge graph was built, and after embedding the knowledge graph using TransE, a patient-level vascular heterogeneous instance graph was constructed. Based on the patient-level vascular heterogeneous instance map, the HGT Heterogeneous Graph Transformer is used for relation modeling. After performing image-text-atlas trimodal pre-training, the downstream task of carotid ultrasound is transferred to generate an interpretable diagnostic path for carotid artery assisted diagnosis.
2. The carotid artery-assisted diagnosis method based on ultrasound visual language according to claim 1, characterized in that: When constructing the carotid artery ultrasound multi-segment image and text dataset, the original DICOM video stream, B-mode keyframes, color Doppler keyframes, spectral Doppler images, doctor's diagnostic reports, structured measurements, and expert annotation information were collected. Using the patient's unique identifier as the primary key, all images and reports from the same patient and the same examination were bound as patient-level samples. The left and right sides, vessel segments, section types, acquisition timestamps, image modes, and device sources were recorded to construct the carotid artery ultrasound multi-segment image and text dataset.
3. The carotid artery-assisted diagnosis method based on ultrasound visual language according to claim 2, characterized in that: When constructing the carotid ultrasound multi-segment image dataset, for the video stream, a joint screening based on inter-frame differences and image sharpness is performed to remove redundant similar frames and obviously blurred frames.
4. The carotid artery-assisted diagnosis method based on ultrasound visual language according to claim 3, characterized in that: When constructing the carotid artery diagnostic knowledge graph, anatomical structure nodes, left and right side nodes, section nodes, image pattern nodes, lesion nodes, ultrasound sign nodes, measurement parameter nodes, risk level nodes, and diagnostic conclusion nodes are used as entity nodes of the graph; belonging to, located at, corresponding section, manifested as, supported by parameters, indicating risk, supporting diagnosis, and follow-up changes are used as relationship types of the graph to construct the carotid artery diagnostic knowledge graph.
5. The carotid artery-assisted diagnostic method based on ultrasound visual language according to claim 1, characterized in that: When using TransE for knowledge graph embedding, the TransE model is used to vectorize entity nodes and relation edges in the knowledge graph. For any triple, TransE is used to constrain the head entity vector, relation vector, and tail entity vector.
6. The carotid artery-assisted diagnostic method based on ultrasound visual language according to claim 1, characterized in that: When constructing a patient-level heterogeneous vascular instance graph, the patient-level heterogeneous vascular instance graph is constructed based on the patient's node set, internal relation edge set, node type and edge type set, and node feature set. The node types include patient nodes, left and right side nodes, vascular segment nodes, section nodes, keyframe nodes, lesion nodes, attribute nodes, measurement parameter nodes, and diagnostic conclusion candidate nodes.
7. The carotid artery-assisted diagnosis method based on ultrasound visual language according to claim 1, characterized in that: When performing relational modeling, a four-layer HGT heterogeneous graph Transformer is used to model the relationships of patient-level vascular instance graphs. HGT is used to set independent projection matrices for different node types and different relational types to handle heterogeneous relationships between patient nodes, vascular segment nodes, keyframe nodes, lesion nodes, attribute nodes and measurement parameter nodes.
8. The carotid artery-assisted diagnostic method based on ultrasound visual language according to claim 1, characterized in that: During trimodal pre-training, the DINOv2 ViT-B / 14 visual encoder, Chinese-MedBERT-base text encoder, TransE knowledge graph embedding, and HGT patient-level vascular graph neural network are used in combination to map image representation, text representation, knowledge graph entity representation, and patient-level graph representation to a unified 512-dimensional common space.
9. The carotid artery-assisted diagnostic method based on ultrasound visual language according to claim 1, characterized in that: When performing downstream task transfer for carotid ultrasound, the visual encoder, TransE knowledge embedding, and HGT graph neural network are transferred to the downstream task of carotid ultrasound. For image-level tasks, joint prediction is performed using keyframe image representation, representation of the vessel segment to which it belongs, and patient-level graph representation. For the segmentation task, the multi-layer visual features output by DINOv2 ViT-B / 14 are used to construct the FPN decoder, which outputs patch segmentation masks and intima-media complex segmentation masks. For patient-level multi-segment lesion burden assessment, attention-weighted convergence of the node representations of each segment on the left and right sides is performed to obtain the patient-level lesion burden vector. In the plaque vulnerability prediction task, lesion node representation, corresponding keyframe visual representation, plaque segmentation morphological features and knowledge graph attributes are embedded into the common input vulnerability task head. The task head is used to output the probability of stable plaques, suspected unstable plaques and high-risk vulnerable plaques, and at the same time output the main basis nodes. In the stenosis grading task, the morphological stenosis rate, lumen diameter change and spectral Doppler velocity parameters are input into the stenosis grading task header. If the image morphology indicates mild stenosis but the spectral velocity indicates moderate stenosis, it is marked by the conflict edge in the patient-level instance image, and the doctor is prompted to review it in conjunction with the original spectral image during the report generation stage.
10. A carotid artery-assisted diagnostic system based on ultrasound visual language, used to perform the carotid artery-assisted diagnostic method based on ultrasound visual language as described in claim 1, characterized in that, include: The front-end preprocessing module is used to construct a carotid artery diagnostic knowledge graph using the constructed carotid ultrasound multi-vessel segment image and text dataset, and then construct a patient-level vascular heterogeneous instance graph after embedding the knowledge graph using TransE. The downstream auxiliary diagnostic module is used to perform relationship modeling based on the patient-level vascular heterogeneous instance map using HGT Heterogeneous Graph Transformer, and after performing image-text-atlas trimodal pre-training, to perform the transfer of carotid ultrasound downstream tasks to generate interpretable diagnostic paths for carotid artery auxiliary diagnosis.