Axillary venogram image synthesis method and system based on multi-modal condition generation

CN122820463APending Publication Date: 2026-09-25SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611307396.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但在血管成像方向存在以下核心问题:仅依靠统计模型或隐式特征学习,忽略了骨与静脉的空间对应关系;没有考虑患者个体化生理特征,难以适配不同患者的解剖变异;单阶段生成架构要求同一模型同时解决血管定位和边界精度两个不同问题,在医学场景高质量配对数据稀缺的情况下,联合训练收敛难度大,过拟合风险也更高

Benefits of technology

本发明通过骨性解剖与血管生成的结构化耦合,区别于现有生成式血管合成方案依赖统计血管模型的隐式约束或像素空间的端到端黑箱映射,本发明首次将临床医生完成腋静脉盲穿刺所依赖的核心认知“以骨推断血管”转化为可计算、可嵌入生成网络的结构化条件:通过显式分割四类骨骼掩码并将其空间分布编码为条件张量的骨骼几何通道,以骨架化计算锁骨-第一肋交点作为显式空间锚点,使得生成血管的走行路径被骨骼参照系锚定,提升了生成结果的解剖合理性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820463A_ABST
    Figure CN122820463A_ABST
Patent Text Reader

Abstract

The application provides an axillary vein angiography image synthesis method and system based on multi-modal condition generation, and belongs to the technical field of medical image intelligent processing. The method comprises the following steps: acquiring a non-contrast X-ray chest film and preprocessing; inputting the X-ray chest film into a bone segmentation network, extracting a bone structure mask, and calculating a clavicle-first rib intersection point as a bony anatomical anchor point; generating a bony anatomical guided blood vessel probability hot zone based on the bone structure mask and the bony anatomical anchor point; constructing a multi-channel anatomical-physiological condition tensor; inputting the X-ray chest film and the multi-channel anatomical-physiological condition tensor into a double-branch cascaded generation network to obtain a fine blood vessel mask, and synthesizing an angiography image based on the fine blood vessel mask and the X-ray chest film. The application significantly improves the accuracy of axillary vein visualization image synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent medical image processing technology, and in particular relates to a method and system for synthesizing axillary vein angiography images based on multimodal condition generation. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Axillary vein puncture is a crucial procedure in surgeries such as the implantation of cardiac implantable electronic devices, whose electrode leads need to be inserted into the heart chambers via a vein. Among the three main venous access routes, the axillary vein, because it runs outside the thoracic cavity, reduces the risk of long-term lead breakage due to pleural injury and "clavicular compression syndrome," and has a high success rate and good safety profile, and has gradually become the preferred approach for these surgeries.

[0004] However, axillary vein puncture presents challenges: the axillary vein is soft tissue and not visualized under conventional X-ray fluoroscopy. Operators must rely on bony landmarks to infer the vein's course, but the correspondence between bone and vein positions varies greatly among individuals, with even more pronounced variations in certain populations. If the initial blind puncture fails, contrast-enhanced venography is usually required to visualize the vessel. Therefore, the limitation of traditional fluoroscopic guidance is the inability to directly display the target puncture structure. While venography can improve the success rate, it can lead to contrast-induced nephropathy, allergic reactions, increased radiation exposure, and prolonged procedure time. Ultrasound guidance can visualize the vessel in real time and has a higher success rate, but it is limited by equipment costs, training requirements, and the technical skill required for operation.

[0005] In recent years, generative artificial intelligence technology has made groundbreaking progress in the field of medical imaging. However, the following core problems exist in vascular imaging: relying solely on statistical models or implicit feature learning ignores the spatial correspondence between bones and veins; it does not consider the individualized physiological characteristics of patients, making it difficult to adapt to the anatomical variations of different patients; the single-stage generative architecture requires the same model to solve two different problems simultaneously: vascular localization and boundary accuracy. In medical scenarios where high-quality paired data is scarce, joint training convergence is difficult, and the risk of overfitting is also higher. Summary of the Invention

[0006] To overcome the shortcomings of the existing technology, this invention proposes a method and system for synthesizing axillary vein angiography images based on multimodal condition generation. This method can directly synthesize anatomically accurate and vascularly continuous axillary vein imaging images from conventional non-contrast X-ray fluoroscopy chest radiographs without contrast agents or additional radiation. This invention explicitly encodes bony anatomical structures as spatial conditions, injects patient physiological characteristics as personalized constraints, and decouples coarse localization and fine generation through architectural design to adapt to the realistic constraints of small sample training in medicine.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, the present invention discloses a method for synthesizing axillary vein angiography images based on multimodal condition generation, comprising: Acquire and preprocess chest X-rays without contrast. The chest X-ray is input into a bone segmentation network, which extracts the bone structure mask through a self-configured encoder-decoder bone segmentation architecture and calculates the clavicle-first rib intersection as a bony anatomical anchor point. Based on the bone structure mask and bony anatomical anchor points, a bony anatomical-guided vascular probability hotspot is generated. A multi-channel anatomy-physiology condition tensor is constructed based on the skeletal structure mask, vascular probability hotspots, and patient physiological information. The chest X-ray and multi-channel anatomical-physiological condition tensor are input into a two-branch cascaded generation network to obtain a fine vascular mask. Angiographic images are synthesized based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, while the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

[0008] Secondly, this invention discloses a system for synthesizing axillary vein angiography images based on multimodal condition generation, comprising: The image acquisition module is configured to acquire and preprocess uncontrast-free chest X-rays. The bone segmentation module is configured to: input the chest X-ray into the bone segmentation network, the bone segmentation network extracts the bone structure mask through a self-configured encoder-decoder bone segmentation architecture, and calculates the clavicle-first rib intersection as the bony anatomical anchor point; A heat zone generation module is configured to generate a vascular probability heat zone guided by bony anatomy based on the bone structure mask and bony anatomical anchor points. The condition generation module is configured to construct a multi-channel anatomical-physiological condition tensor based on the skeletal structure mask, vascular probability hotspots, and patient physiological information. The image synthesis module is configured to: input the chest X-ray and multi-channel anatomical-physiological condition tensor into a two-branch cascaded generation network to obtain a fine vascular mask; and synthesize an angiographic image based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, and the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

[0009] Thirdly, the present invention discloses an electronic device, including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when run by the processor, complete the steps of the above-mentioned method for synthesizing axillary vein angiography images based on multimodal conditions.

[0010] Fourthly, the present invention discloses a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the above-described method for synthesizing axillary vein angiography images based on multimodal conditions.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention, through the structured coupling of bony anatomy and angiogenesis, differs from existing generative angiogenesis schemes that rely on implicit constraints of statistical vascular models or end-to-end black-box mapping of pixel space. For the first time, this invention transforms the core understanding of "inferring blood vessels from bones" that clinicians rely on to perform blind axillary vein puncture into a computable and embeddable structured condition for generative networks: by explicitly segmenting four types of skeletal masks and encoding their spatial distribution into skeletal geometric channels of conditional tensors, and using the skeletalized calculation of the clavicle-first rib intersection as an explicit spatial anchor point, the path of the generated blood vessels is anchored by the skeletal reference system, thereby improving the anatomical rationality of the generated results.

[0012] Simultaneously, this invention achieves individualized anatomical adaptation for patients. Gender and body mass index (BMI) are injected as physiological prior conditions into the generation process. Gender is mapped using a learnable embedding table, and BMI is mapped using Fourier features and then embedded into the physiological feature channels of the conditional tensor. This allows the same model to produce adaptive vascular course results based on the body characteristics of different patients. For subgroups with significant anatomical variations, such as obese patients (high BMI leading to subcutaneous fat thickening and chest wall deformation) and female patients (gender differences in chest wall morphology affecting the relative position of blood vessels and bones), the model automatically adjusts the offset of blood vessels relative to the skeletal reference frame, avoiding the problem of degradation in generation accuracy in specific subgroups caused by a "one-size-fits-all" approach of general models.

[0013] This invention significantly reduces the risk of overfitting under small medical sample conditions through a coarse-fine decoupled cascade architecture. The angiogenesis task is explicitly decoupled into two sub-problems: "Where is the region?" and "How accurate are the edges?", using a phased frozen training strategy. The first branch is trained until convergence and then frozen; the second branch is then trained under the frozen prior conditions, effectively avoiding gradient conflicts in multi-objective joint optimization using a single-stage model. Furthermore, the vessel continuity regularization term in the second branch explicitly suppresses common failure modes of diffusion models under small medical sample conditions at the loss function level.

[0014] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0015] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0016] Figure 1 This is a flowchart of the axillary vein angiography image synthesis method based on multimodal conditions as described in Embodiment 1 of the present invention.

[0017] Figure 2 This is a schematic diagram of the architecture of the dual-branch cascaded generation network described in Embodiment 1 of the present invention.

[0018] Figure 3 This is a diagram illustrating the results of the experiment described in Embodiment 1 of the present invention; wherein, as shown... Figure 3 (a) shows the base map, as shown in Figure 1. Figure 3 Image (b) shows the contrast image, as shown in Figure 1. Figure 3 (c) in the figure shows the generated graph. Detailed Implementation

[0019] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0020] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0021] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0022] Example 1 In one or more embodiments, a method for synthesizing axillary vein angiography images based on multimodal conditional generation is disclosed, such as... Figure 1 As shown, it includes the following steps: Step S1: Obtain and preprocess chest X-rays without contrast.

[0023] Acquire intraoperative chest radiographs without contrast under fluoroscopy. ,in This refers to the pixel dimensions of a chest X-ray image (anteroposterior view). The image is a single-channel grayscale image, with pixel values ​​normalized to [value missing]. .

[0024] To enhance the contrast between bony landmarks (clavicle, edge of the first rib) and surrounding soft tissues to facilitate subsequent bone segmentation, conventional contrast enhancement and edge sharpening preprocessing were performed on the chest radiographs to obtain enhanced chest radiographs for use in subsequent modules.

[0025] As an example, the contrast enhancement and edge sharpening preprocessing includes a weighted fusion of contrast-limited adaptive histogram equalization and Laplacian sharpening, each with a weight of 0.5.

[0026] Simultaneously, the patient's gender category is extracted from the medical imaging (DICOM) header or electronic medical record. , ( For women, (for men) and body mass index (Unit: kg / m²), used as physiological information input for subsequent conditional coding modules.

[0027] This invention uses conventional non-contrast X-ray fluoroscopy chest radiographs as the sole image input, requiring only two physiological parameters—sex and BMI—obtainable from electronic medical records. The entire process requires no injection of iodine contrast agents and involves no additional X-ray scans or three-dimensional volumetric acquisition, achieving non-invasiveness and zero additional radiation exposure. It fundamentally eliminates the risks of contrast agent nephrotoxicity, allergic reactions, and increased radiation exposure for both doctors and patients. Compared to DSA / CTA-based vascular reconstruction methods (which require contrast-injected angiography data), this invention represents a paradigm shift from "modeling after contrast enhancement" to "contrast-free generation." Compared to ultrasound-guided methods (which require specialized probes, sterile consumables, and operator training), this invention runs purely in software on existing fluoroscopy equipment, without increasing hardware costs or sterilization procedures.

[0028] Step S2: Input the chest X-ray into the bone segmentation network. The bone segmentation network extracts four types of bone structure masks—clavicle, first rib, outer contour of sternum, and humerus—through a self-configured encoder-decoder bone segmentation architecture, and calculates the clavicle-first rib intersection as a bony anatomical anchor point.

[0029] In this embodiment, the skeletal segmentation network adopts a self-configured encoder-decoder skeletal segmentation architecture, including an encoder and a decoder.

[0030] The encoder part uses stacked depthwise convolutions to extract hierarchical feature maps at four scales through progressive downsampling:

[0031] in, For scale i The hierarchical feature map, whose spatial resolution ranges from Decrease step by step to The number of channels increases progressively, for example from 64 to 128 to 256 and then to 512, thus forming a feature pyramid; These are the learnable parameters of the encoder; For encoder.

[0032] Among them, deep convolution stacking can use ResNet-50 or Vision Transformer as the feature extraction backbone.

[0033] The decoder section restores the multi-scale feature maps to their original resolution through step-by-step upsampling and skip connections, outputting binary masks for four types of skeletal structures. These correspond to the clavicle, the first rib, the outer contour of the sternum, and the humerus, respectively.

[0034] As one implementation method, the skeleton segmentation network employs the nnU-Net v2 model (a self-configuring encoder-decoder segmentation framework). Specifically: The framework trained four independent nnU-Net models for four types of bones, including the clavicle, first rib, outer contour of the sternum, and humerus, with each model being a 2D configuration.

[0035] The encoder backbone can be ResNet-50 or Vision Transformer as the feature extraction backbone, and hierarchical feature maps of four scales are extracted by progressive downsampling.

[0036] The decoder restores the original resolution through step-by-step upsampling and skip connections, and outputs binary masks for four types of bones.

[0037] The training objective of the skeletal segmentation network is a combination of Dice similarity coefficient loss and weighted cross-entropy loss:

[0038] in, For the first Artificially labeled mask for skeleton-like structures. To smooth out small quantities (take) ), For the category weights of various skeletons, For standard cross-entropy loss, These are spatial pixel coordinates. Preferably, the category weights for each type of bone are: clavicle. First rib Because its boundaries are less defined, it requires higher weighting; thoracic cavity humerus .

[0039] It should be understood that the training data must be annotated pixel by pixel on an angiographic chest radiograph by a radiologist or a trained annotator to define the contour regions of the four types of bones.

[0040] Preferably, for each type of skeletal mask in the segmentation output, the maximum connected component is selected to remove isolated noise fragments.

[0041] Subsequently, the clavicle mask was applied separately. and the first rib mask Perform morphological skeletonization. Iteratively peel away boundary pixels until a skeleton curve with a single pixel width is obtained:

[0042] in, For skeleton curves, This is a skeletalized process.

[0043] Bony anatomical anchor point coordinates Defined as the point on the clavicular side among the nearest pair of points between two skeleton point sets:

[0044] in, The pixel coordinates of the clavicle mask; The pixel coordinates of the first rib mask; The skeleton curve of the clavicle mask; The skeleton curve of the first rib mask.

[0045] The clinical anatomical basis for this anchor point is that, under standard anteroposterior or right anterior oblique X-ray fluoroscopy, the axillary vein typically runs through the intersection of the clavicle and the first rib, with its approximate orientation relative to the intersection point being approximately 4 o'clock. This pattern has been repeatedly validated in multiple independent clinical cohorts. Encoding this empirical knowledge into computable anchor point coordinates allows the search space for vascular region localization to be expanded from the entire map (…). The search space is compressed from pixels to a limited neighborhood centered on the anchor point (with a radius of about 100-200 pixels), reducing the search space by more than an order of magnitude.

[0046] Step S3: Generate a vascular probability hotspot guided by the skeletal structure mask and bony anatomical anchor points. The vascular probability hotspot compresses the vascular localization search space from the entire map to a reasonably anatomically appropriate region.

[0047] After obtaining the skeleton mask and bony anatomical anchor points, a vascular probability heatmap is generated. This is used to characterize the probability of the axillary vein appearing at various spatial locations on a chest radiograph. The probabilistic hotspot generation model employs an encoder-decoder segmentation architecture, with the enhanced chest radiograph as its input. And auxiliary channels constructed from skeleton masks and anchor points:

[0048] in, For the anchor point distance map, the value at each spatial location is defined as the distance from that point to the anchor point. The normalized value of the Euclidean distance after Gaussian kernel decay:

[0049] In the formula, This is a spatial scale parameter that controls the spatial scale of the influence range of the anchor point; typical value. Pixels. This distance map provides the network with explicit spatial attention guidance on "looking for veins near anchor points".

[0050] In this embodiment, the probabilistic hotspot generation model can adopt network models such as the U-Net series or its variants.

[0051] The training objective of the probabilistic hotspot generation model is binary cross-entropy loss. With Dice loss Weighted combination:

[0052] in, For true labeling, the axillary vein mask, manually annotated in contrast-enhanced venography images, is morphologically dilated (dilution kernel radius). (Pixels) are generated. The dilated labels cover the blood vessels and the surrounding area within a certain tolerance range, forming a "coarse but highly recall" training target.

[0053] After training, the blood vessel probability hotspots output during the inference phase are analyzed. Perform threshold binarization:

[0054] in, Preset threshold (recommended value) ); The rough mask for blood vessels can serve as the attention range in which the subsequent fine generative network should "fill" the blood vessels.

[0055] In addition to the learning-based methods mentioned above, the following alternative strategies can be used during the training phase to generate rough vascular masks: the static elliptical region method, which uses the anchor point as the center and fits a statistical ellipse of the vascular distribution in the training set (with the major axis along the typical course of the axillary vein and a major-to-minor axis ratio of approximately 3:1) as the vascular region; and the manual annotation dilation method, which directly performs morphological dilation on the manually annotated precise vascular mask to obtain a rough mask.

[0056] In actual deployment, the inference phase prioritizes the use of pre-trained probabilistic hotspot generation models, as they can adapt to the anatomical structures of different patients without the need for manual annotation.

[0057] Step S4: Construct a multi-channel anatomy-physiology condition tensor based on the bone structure mask (bone geometry information), vascular probability hotspot (vascular region information) and patient physiological information.

[0058] The three types of heterogeneous information are uniformly encoded into a structured conditional tensor that can be consumed by the downstream generating network. The three types of information and their encoding paths are as follows: (a) Skeletal geometry information: four types of skeleton masks First via the skeletal encoder Multi-scale spatial feature extraction is performed. (Skeleton encoder) This is a small convolutional network (without using downsampling to maintain full resolution), accepting multi-channel stacked binary masks as input, with an output channel count of [number missing]. Skeletal geometric feature map This feature map encodes geometric information that constrains the course of blood vessels, such as the curvature of the lower edge of the clavicle, the inclination angle of the first rib, and the width of the clavicle-rib gap.

[0059] (b) Vascular region information: probability hotspots thermal compression mapping Output Spatial attention feature map of the channel This indicates the spatial extent in which blood vessels may be present. Among these, thermal zone compression mapping... It includes 1×1 convolutional layer dimensionality reduction and the Sigmoid activation function.

[0060] (c) Patient physiological information: gender Learnable Embedded Tables The gender embedding vector is obtained by looking up the table. :

[0061] in, For learnable embedded table dimensions.

[0062] Body Mass Index (BMI scalar) Fourier feature mapping is used to embed into a high-dimensional space to capture high-frequency variation patterns of blood vessel course corresponding to different body types.

[0063] in, This is the BMI mapping vector. The frequency basis is random Gaussian sampling. ,, To control the frequency scale, For mapping dimensions, preferably .

[0064] In this embodiment, by using Fourier feature mapping to project a low-dimensional scalar onto a high-dimensional space through a random Fourier basis, subsequent linear layers can more effectively learn the high-frequency change patterns of the input, which is particularly crucial for modeling the nonlinear effects of BMI changes on vein course.

[0065] Furthermore, the gender embedding vector and the BMI mapping vector are fused using a physiological feature fusion function. The images are concatenated and broadcast to the entire image space, with an output channel count of [number missing]. Physiological characteristics diagram The physiological feature fusion function includes fully connected layers and spatial broadcasting.

[0066] Spatial Broadcasting copies the physiological feature vectors output by the fully connected layer to all spatial locations in the entire image, generating a physiological feature map. This ensures that each pixel carries the same individualized physiological information of the patient, which can be accessed by the subsequent convolutional operations of the generative network at each spatial location.

[0067] Finally, the three types of feature maps are concatenated along the channel dimension to obtain a multi-channel anatomical-physiological condition tensor. :

[0068] in, This represents the total number of channels, and Typical configuration is .

[0069] This conditional tensor The core design concept is as follows: the skeletal mask is a two-dimensional geometric shape (encoding the "skeleton reference frame"), the probability hotspot is a two-dimensional probability distribution (encoding "where it was generated"), and the physiological parameters are one-dimensional scalars (encoding "for whom it was generated"). These three differ significantly in mathematical properties and are not suitable for simple summation or averaging. This step achieves a unified representation of heterogeneous information in the feature space by designing suitable embedding paths for various types of information (convolutional coding for geometric shapes, dimensionality reduction compression for probability distributions, and Fourier mapping + embedding tables for scalars / categories). The two branches of the subsequent bi-branch cascaded generation network will be fed in simultaneously to ensure that coarse localization and fine generation share the same set of anatomical-physiological constraint space.

[0070] Step S5: Input the chest X-ray and multi-channel anatomical-physiological condition tensor into a two-branch cascaded generation network to obtain a fine vascular mask. Synthesize an angiographic image based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, and the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

[0071] like Figure 2 As shown, with enhanced chest X-ray and multichannel anatomy-physiology condition tensor As input, a high-precision axillary vein mask is generated through a two-branch cascaded generative network. This network decouples the difficult task of "generating a vascular mask from a non-angiographic chest X-ray" into two sub-problems, each handled by a separate branch: First branch – Regional coarse localization subnetwork: The first branch uses a pre-trained generative prior diffusion model as its foundation, and adds a conditional injection layer on top of its self-configured encoder-decoder skeletal segmentation architecture. The foundation is a generative diffusion model pre-trained on natural images or large-scale medical image datasets, which has the ability to cross-modal map from pixel space to semantic space.

[0072] Specifically, conditional tensors The generation process is injected at various scale levels of the decoder through multiple conditional injection layers. The first branch outputs a rough mask of blood vessels. .

[0073] Specifically, the conditional injection layer can employ a spatial adaptive normalization layer or a cross-attention layer. This conditional injection layer is present at various scale levels of the decoder; specifically, it can be connected to each convolutional layer of the decoder, achieving full-scale conditional guidance from coarse to fine. The decoder starts with the deepest features, upsampling level by level. At each level, conditional information is injected through the spatial adaptive normalization layer / cross-attention layer, and the decoder outputs a coarse mask of blood vessels.

[0074] In this process, the spatial adaptive normalization layer normalizes the decoder feature map channel by channel, and then uses the scaleγ and biasβ parameters learned from the conditional tensor through a small convolutional network to perform an affine transformation, thereby injecting anatomical-physiological conditional information into the generation process.

[0075] The cross-attention layer uses the decoder features as the query and the spatial features of the condition tensor C as the key and value, and fuses the conditional information through an attention mechanism.

[0076] The training objective of the first branch is Tversky loss, which focuses on recall, meaning it's better to overreport (tolerance for false positives) than to miss (severe punishment for false negatives).

[0077] in, The loss function for the first branch; For precise vessel masks manually annotated from contrast-enhanced intravenous imaging images, The false positive penalty weight for the first branch, The weight for the false negative penalty in the first branch, and , , .because The model's penalty for missed vascular pixels (false negatives) is approximately equal to that for false positives of non-vascular pixels (false positives). This multiplies the output of the first branch, resulting in a wide range of outputs and a high recall rate in rough vascular regions.

[0078] The first branch is trained for approximately 100-150 epochs until the Dice coefficients converge on the validation set. Then, all parameters are frozen, and the second branch is started.

[0079] Second branch – Anatomically constrained fine-grained generative subnetwork: The second branch has the same structure as the first branch (it copies the backbone weights of the first branch as initialization), but the input channels and training strategies are different.

[0080] The input to the second branch is the enhanced chest X-ray. Conditional tensor The coarse mask output by the first branch and four types of skeleton masks The stack, total One input channel. As a spatial attention prior, the network is informed that "blood vessels are roughly located in this area," while the skeleton mask serves as a hard anatomical constraint, informing the network that "the skeleton region is not a blood vessel."

[0081] The training parameters for the second branch are performed while the parameters for the first branch and the pre-trained prior diffusion model are frozen. Its training objectives include three aspects, specifically:

[0082] in, The loss function for the second branch; For Tversky's loss, For fine masking of blood vessels, For balance coefficient, To regularize vascular continuity; The intensity coefficient of the continuous penalty; The false positive penalty weight for the second branch; The false negative penalty weight for the second branch; The first term is the Tversky loss. In contrast to the first branch, the second branch focuses on accuracy, imposing a higher penalty on false positives (the generation of blood vessels in non-vascular areas) and causing the boundary to shrink to the precise location.

[0083] The second item is inter-branch consistency regularization. The L1 distance constrains the fine mask to not deviate too far from the coarse mask (to prevent the second branch from over-contracting and causing blood vessel rupture when optimizing accuracy).

[0084] The third item is vascular continuity regularization. Defined as a second-order smoothness constraint along the direction of the vessel centerline:

[0085] in, For the mask The centerline curve of the blood vessels extracted after skeletonization. For arc length parameters, The arc length on the center line is The spatial coordinates of the location. The regularization term penalizes drastic discontinuous changes in the mask along the direction of the blood vessel. If the generated mask experiences a rupture in the lumen at some point (the mask value on the center line suddenly jumps from 1 to 0), its second derivative will produce a maximum value, thus being significantly penalized by the regularization term.

[0086] After the second branch is trained to convergence, the final high-precision fine-grained blood vessel mask is obtained. .

[0087] The dual-branch design of this invention is not a simple model ensemble, but is based on the following observation: Under extremely small sample conditions with only a few dozen pairs of paired training data, when a single generative model simultaneously optimizes "where are the blood vessels roughly?" (global localization, a low spatial frequency task) and "where are the blood vessel boundaries precisely?" (local boundaries, a high spatial frequency task), the gradients of the two tasks conflict in the shared parameter space. High recall requires the model to output a large, loose region, while high precision requires the model to output a compact, precise contour. By assigning the two subtasks to two independently tunable branches and employing staged freeze training, this architecture effectively avoids multi-objective gradient conflicts without increasing the amount of training data, allowing each branch to converge stably under its own dedicated loss function hyperparameter configuration.

[0088] Furthermore, fine vascular masking Rendered as a semi-transparent color layer using a preset color mapping scheme, and overlaid on the original enhanced chest X-ray. Above, generate synthetic imaging images. :

[0089] in, To display the threshold, the preferred value is... Pixels with a low probability of being below this value are considered non-vascular and are not superimposed; Adjust the opacity of color blending; These are color mapping parameters; To preset the color for displaying blood vessels, magenta RGB (255, 0, 128) is preferably used to create high contrast with the grayscale background of X-rays.

[0090] Synthetic imaging The video stream is pushed to the fluoroscopic monitor in real time and can be displayed side by side with the original fluoroscopic image or in a picture-in-picture format, allowing the surgeon to directly refer to and observe the color-marked axillary vein path during the puncture process, and complete the puncture approach planning and needle insertion operation under direct vision.

[0091] The target computation latency for a single-frame pipeline is within 300 ms (including image preprocessing, skeleton segmentation, hotspot generation, conditional coding, bi-branch inference, and overlay rendering). Typical hardware configuration is an NVIDIA RTX 3090 / 4090 or equivalent GPU accelerator card. At a perspective frame rate of 7.5-15 fps, a frame skipping strategy can be adopted (execute a complete pipeline every 2-3 frames, and reuse the most recently generated mask in intermediate frames after rigid body registration and fine-tuning) to balance computational load and display continuity.

[0092] Preferably, this embodiment performs a multi-dimensional anatomical consistency quality assessment on the synthesized contrast images.

[0093] To further provide a quantitative confidence reference for the generated results, the following three-dimensional consistency evaluation indicators can be added (this evaluation module does not increase inference latency and is calculated asynchronously after the generated results are output): (a) Consistency score of vascular diameter:

[0094] in, and The values ​​represent the average lumen diameter of the axillary vein at key anatomical locations (lower border of the clavicle, first costal cross, and deep to the pectoralis minor muscle) in the generated image and the reference template, respectively. This is the normalized tolerance coefficient.

[0095] (b) Hausdorff distance from the vascular centerline:

[0096] in, To generate a spatial coordinate point on the center line of the blood vessel; This is a spatial coordinate point on the center line of the blood vessel in the reference template; The generated vessel centerline curve is extracted by skeletonizing the fine mask; The reference template vascular centerline curve is derived from a population statistical anatomical prior template; This is to find the minimum distance from a point to all points on curve C, i.e., to calculate the distance between the generated curve and the reference curve. This is the maximum value after traversing all points on the curve, that is, the maximum value among the shortest distances from all points on the input curve to the target curve.

[0097] To measure the spatial deviation between the generated blood vessel course and the anatomical reference template.

[0098] (c) Deviation between bony anchor point and vessel relative position:

[0099] bony anatomical anchor points The nearest edge point to the generated blood vessel The Euclidean distance between them directly measures whether the generated result conforms to the clinical anatomical prior that "the axillary vein is located at approximately the 4 o'clock position near the intersection of the clavicle and the first rib".

[0100] The overall quality score is defined as:

[0101] in, As the indicator weight, and ; and The normalization threshold can take values ​​of: , If and only if At that time, the generated contrast images were determined to meet the usable quality standards for clinical puncture planning.

[0102] Through the aforementioned multi-dimensional and quantifiable anatomical consistency quality assessment, this invention establishes a three-dimensional quality assessment system encompassing vessel diameter, topological course, and spatial deviation between bony anchor points and vessels. This provides a transparent, objective, and reproducible evaluation standard for the clinical usability of generated images. The calculation process of this assessment module does not rely on real angiographic images; the reference template required for assessment is a population statistical prior, which can run asynchronously during the inference phase without increasing the latency of the main pipeline path.

[0103] This invention is deployed as a software module on the image processing workstation or external GPU computing unit of the fluoroscopy equipment. During the operation, only fluoroscopic frames need to be extracted from the DICOM video stream, and a composite image with color-coded vascular annotations is output and sent back to the monitor. The entire process requires no change to the existing operating room layout, disinfection procedures, or surgeon's operating habits, and is seamlessly integrated with the existing fluoroscopy workflow. This "zero hardware barrier" characteristic is particularly crucial for its widespread adoption in grassroots hospitals with limited medical resources.

[0104] Experiments were conducted to verify the beneficial effects of this scheme. The experimental results are as follows: Figure 3 As shown, Figure 3 (a) shows the base image, which refers to the input chest X-ray, such as... Figure 3 Image (b) shows an angiography image, which refers to the angiography of an actual patient, such as... Figure 3 (c) shows the generated image, which refers to the vascular result image generated by this method. It can be seen that the results generated by this method are very close to the actual angiography, demonstrating high accuracy.

[0105] Table 1 Comparative Experimental Results

[0106] Furthermore, a comparative experiment was conducted on this method. The method was compared with the basic diffusion model, GPT-Image2 (multimodal large model), on an independent dataset of 25 individuals. The dataset used was legal and compliant, and the consent of the data subjects was obtained. The generation results of the three methods are shown in Table 1. It can be seen that the overlap measure of the generated blood vessels and the real blood vessels by this method, DICE, is significantly improved compared with the comparison methods.

[0107] This invention can be applied to X-ray fluoroscopic guidance for axillary vein puncture during cardiac implantable electronic devices (CIED) implantation, as well as clinical scenarios requiring the location of invisible vascular structures under fluoroscopy, such as central venous catheterization and hemodialysis access establishment.

[0108] Example 2 In one or more embodiments, a system for synthesizing axillary vein angiography images based on multimodal condition generation is disclosed, specifically including: The image acquisition module is configured to acquire and preprocess uncontrast-free chest X-rays. The bone segmentation module is configured to: input the chest X-ray into the bone segmentation network, the bone segmentation network extracts the bone structure mask through a self-configured encoder-decoder bone segmentation architecture, and calculates the clavicle-first rib intersection as the bony anatomical anchor point; A heat zone generation module is configured to generate a vascular probability heat zone guided by bony anatomy based on the bone structure mask and bony anatomical anchor points. The condition generation module is configured to construct a multi-channel anatomical-physiological condition tensor based on the skeletal structure mask, vascular probability hotspots, and patient physiological information. The image synthesis module is configured to: input the chest X-ray and multi-channel anatomical-physiological condition tensor into a two-branch cascaded generation network to obtain a fine vascular mask; and synthesize an angiographic image based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, and the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

[0109] Example 3 This embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the computer instructions are executed by the processor, they complete the steps of the above-described method for synthesizing axillary vein angiography images based on multimodal conditions.

[0110] Example 4 This embodiment provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the above-described method for synthesizing axillary vein angiography images based on multimodal conditions.

[0111] This invention is described with reference to flowchart illustrations and / or block diagrams of methods and apparatus (systems) according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer instructions. These computer instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer instructions can also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0115] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for synthesizing axillary vein angiography images based on multimodal condition generation, characterized in that, include: Acquire and preprocess chest X-rays without contrast. The chest X-ray is input into a bone segmentation network, which extracts the bone structure mask through a self-configured encoder-decoder bone segmentation architecture and calculates the clavicle-first rib intersection as a bony anatomical anchor point. Based on the bone structure mask and bony anatomical anchor points, a bony anatomical-guided vascular probability hotspot is generated. A multi-channel anatomy-physiology condition tensor is constructed based on the skeletal structure mask, vascular probability hotspots, and patient physiological information. The chest X-ray and multi-channel anatomical-physiological condition tensor are input into a two-branch cascaded generation network to obtain a fine vascular mask. Angiographic images are synthesized based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, while the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

2. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 1, characterized in that, An encoder-decoder segmentation architecture is used to generate the vascular probability hotspot guided by the bony anatomy. The input to the encoder-decoder segmentation architecture is an enhanced chest X-ray and an auxiliary channel constructed from a bone mask and anchor points. The auxiliary channel includes a stitched chest X-ray, a bone mask, and an anchor point distance map. The value at each spatial location of the anchor point distance map is defined as the normalized value of the Euclidean distance from that point to the anchor point after Gaussian kernel attenuation.

3. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 1, characterized in that, The first branch of the dual-branch cascaded generative network is based on a pre-trained prior diffusion model, with a conditional injection layer added on its encoder-decoder backbone. The conditional tensor is injected into the generation process at various scale levels of the decoder through multiple conditional injection layers. Specifically, the conditional injection layers adopt spatial adaptive normalization layers or cross-attention layers, and the conditional injection layers are connected to each convolutional layer of the decoder.

4. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 1, characterized in that, The second branch of the bi-branch cascaded generative network has the same structure as the first branch, and its input consists of the enhanced chest X-ray, the conditional tensor, the coarse mask output by the first branch, and the stack of four types of bone masks.

5. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 4, characterized in that, The training parameters for the second branch are performed under the condition that both the first branch and the pre-trained prior diffusion model are frozen. Specifically: in, The loss function for the second branch; For Tversky's loss, For fine masking of blood vessels, For balance coefficient, To regularize vascular continuity; The intensity coefficient of the continuous penalty; Precise vascular masks manually annotated from contrast-enhanced intravenous imaging images; The false positive penalty weight for the second branch; The false negative penalty weight for the second branch; This is a coarse mask.

6. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 5, characterized in that, A vascular continuity regularization is introduced to suppress luminal fragmentation artifacts. The vascular continuity regularization is defined as a second-order smoothing constraint along the vascular centerline direction. in, For the mask The centerline curve of the blood vessels extracted after skeletonization. For arc length parameters, The arc length on the center line is Spatial coordinates at the location.

7. The method for synthesizing axillary vein angiography images based on multimodal condition generation as described in claim 1, characterized in that, The process of synthesizing an angiographic image based on the fine vascular mask and chest X-ray involves: rendering the fine vascular mask as a semi-transparent color layer using a preset color mapping scheme, and overlaying it onto the original enhanced chest X-ray. Above, generate synthetic imaging images. : in, To display the threshold; Adjust the opacity of color blending; These are color mapping parameters; The preset color is used to display blood vessels; For fine vascular masking; These are pixel coordinates; It is a chest X-ray.

8. A system for synthesizing axillary vein angiography images based on multimodal condition generation, characterized in that, include: The image acquisition module is configured to acquire and preprocess uncontrast-free chest X-rays. The bone segmentation module is configured to: input the chest X-ray into the bone segmentation network, the bone segmentation network extracts the bone structure mask through a self-configured encoder-decoder bone segmentation architecture, and calculates the clavicle-first rib intersection as the bony anatomical anchor point; A heat zone generation module is configured to generate a vascular probability heat zone guided by bony anatomy based on the bone structure mask and bony anatomical anchor points. The condition generation module is configured to construct a multi-channel anatomical-physiological condition tensor based on the skeletal structure mask, vascular probability hotspots, and patient physiological information. The image synthesis module is configured to: input the chest X-ray and multi-channel anatomical-physiological condition tensor into a two-branch cascaded generation network to obtain a fine vascular mask; and synthesize an angiographic image based on the fine vascular mask and the chest X-ray. The first branch of the two-branch cascaded generation network performs coarse region localization with recall as the optimization objective, and the second branch performs fine generation of anatomical constraints with precision as the optimization objective under the condition of frozen prior, and introduces vascular continuity regularization to suppress lumen rupture artifacts.

9. An electronic device, characterized in that, The method includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the axillary vein angiography image synthesis method based on multimodal condition generation as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, complete the axillary vein angiography image synthesis method based on multimodal condition generation as described in any one of claims 1-7.