Rendering processing method, system and equipment based on medical image intelligent segmentation and medium

By employing an image processing method based on a semantic interaction model, combined with natural language tasks and image features, rapid and accurate segmentation and rendering of cardiovascular and cerebrovascular disease images were achieved. This addresses the shortcomings of existing image processing methods and improves image readability and the accuracy of segmentation results.

CN122023441APending Publication Date: 2026-05-12BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2026-01-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing medical image processing methods are insufficient for quickly and accurately identifying the details of cardiovascular and cerebrovascular diseases, especially in the emergency, surgical planning, and follow-up assessment stages, and cannot meet physicians' needs for rapid and accurate interpretation of images.

Method used

An image processing method based on a semantic interaction model is adopted. By associating natural language task instructions with image feature vectors to generate joint feature vectors, clarifying questions are actively initiated to update segmentation parameters. Combined with graph attention network and dynamic graph convolutional network, blood vessel segmentation and rendering are performed to generate hemodynamic parameters and achieve fine rendering of the three-dimensional voxel structure of blood vessels.

Benefits of technology

It enables rapid and accurate segmentation and rendering of cardiovascular and cerebrovascular disease images, improving image readability, reducing the user's interpretation burden, and meeting clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023441A_ABST
    Figure CN122023441A_ABST
Patent Text Reader

Abstract

The invention provides a rendering processing method, system and device based on medical image intelligent segmentation and a medium. The rendering processing method comprises the following steps: acquiring semantic confidence distribution of each voxel in a target image by utilizing a semantic interaction model; defining, for each voxel, a multi-modal input signal representing grayscale or intensity information from a different image modality; and determining the semantic weight of each voxel according to the semantic interaction model, and determining the illumination intensity and transparency corresponding to each voxel according to the multi-modal input signal and the semantic weight of each voxel to render the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and more specifically, to rendering processing methods, systems, devices, and media based on medical image segmentation. Background Technology

[0002] The diagnosis and treatment of cardiovascular and cerebrovascular diseases usually require precise localization of lesion details. Taking aortic dissection as an example, the location of the tear and the extent of the false lumen directly determine the surgical plan. This is consistent with the logic of how the width and location of the neck of a cerebral aneurysm affect the surgery. Both require imaging technology to clearly show the relationship between the lesion and the surrounding structures.

[0003] Image reconstruction and refined rendering form the "data foundation" for refined rendering of cardiovascular and cerebrovascular images. The degree of structuring directly determines the clinical value of the rendering, permeating the entire diagnostic and treatment process and influencing physician interpretation habits. For example, in the emergency phase, image reconstruction needs to convert two-dimensional images into three-dimensional structured models, clearly marking lesion targets (such as aortic ruptures and cerebral aneurysms) to help emergency physicians quickly locate critical lesions. Furthermore, in the surgical planning phase, reconstruction needs to preserve details such as vessel wall thickness and branch opening locations to support transparent rendering. For instance, coronary CTA reconstruction can reveal the relationship between plaques and the vessel lumen by making the vessel wall transparent, while aortic reconstruction can make the true lumen transparent to determine the distance between stents and branch vessels. In addition, during follow-up evaluation, reconstruction needs to generate standardized morphological parameters (such as aortic diameter and infarct volume) to facilitate comparison of rendering results with historical data.

[0004] In clinical practice, physicians need to quickly identify key information about lesions, such as locating aortic ruptures or occluded vessels in cerebral infarction during emergencies. Existing image post-processing methods are insufficient to meet the requirements for "rapid and accurate" interpretation. Therefore, a new rendering processing method based on medical image segmentation is needed to improve the aforementioned technical problems. Summary of the Invention

[0005] To address the aforementioned issues, this disclosure provides an image processing method, system, device, and medium based on a semantic interaction model.

[0006] According to one aspect of this disclosure, a rendering processing method based on medical image segmentation is provided, comprising: associating semantic feature vectors obtained according to natural language task instructions. Image feature vectors in the target image To generate joint feature vectors ,in, Represents the image voxel coordinates; based on the semantic feature vector with the highest confidence. The system generates initial values ​​for a set of configuration parameters based on the corresponding joint feature vectors, wherein the set of configuration parameters includes one or more segmentation parameters. Based on the consistency between the values ​​of each segmentation parameter in the set of configuration parameters and the image feature reference values, a clarification question is raised to the user to obtain updated values ​​for each segmentation parameter in the set of configuration parameters. Segmentation processing is performed on the target image based on the determined updated values ​​of each segmentation parameter in the set of configuration parameters to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of the blood vessel. Key regions in the three-dimensional voxel structure of the blood vessel are determined based on the natural language task instructions and the user's feedback to the clarification question; a physically based volume rendering method is used to render the three-dimensional voxel structure of the blood vessel and enhance the rendering of the key regions. Hemodynamic parameters are determined based on the blood vessel segmentation result, and the three-dimensional voxel structure of the blood vessel is rendered according to the cardiac cycle based on the hemodynamic parameters, wherein the hemodynamic parameters include at least the streamline data and spatial coordinates of the blood flow streamline nodes.

[0007] According to one aspect of this disclosure, determining key regions in the vascular 3D voxel structure based on the natural language task instructions and user feedback on the clarification question includes: determining semantic nodes representing key rendering requirements from the user based on the natural language task instructions and user feedback on the clarification question; and associating the semantic nodes representing the key requirements with the vascular 3D voxel structure using a graph attention network (GAT), wherein in the graph attention network, the attention coefficient between each vascular 3D voxel node in the vascular 3D voxel structure and its semantic nodes in the anatomical structure is... As shown in the following formula:

[0008]

[0009] in and These represent the feature vectors of a 3D voxel node and a semantic node, respectively. For learnable weight matrix, The attention parameter vector is defined; and the key regions in the three-dimensional blood vessel voxel structure are determined based on the attention coefficients corresponding to each three-dimensional blood vessel voxel node in the three-dimensional blood vessel voxel structure.

[0010] According to one aspect of this disclosure, performing segmentation processing on the target image based on the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a blood vessel segmentation result includes: performing a pre-segmentation model on the target image based on the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a pre-segmentation result; and using a first computing network to obtain a complete segmentation result for the target image based on the pre-segmentation result, and using a second computing network to obtain a straightened blood vessel model based on the complete segmentation result to obtain a simplified segmentation result.

[0011] According to one aspect of this disclosure, the simplified segmentation result is represented as a three-dimensional point cloud. Extracting hemodynamic parameters based on the simplified segmentation result includes: inputting the internal point cloud of the simplified segmentation result's three-dimensional point cloud into a first channel of a Dynamic Graph Convolutional Network (DGCNN); inputting the surface point cloud of the simplified segmentation result's three-dimensional point cloud into a second channel of the DGCNN, wherein the first channel is independent of the second channel; the first channel and the second channel respectively encode features of the input point cloud, wherein features of the surface point cloud in the second channel are aggregated to form a global feature tensor; feature representations of the internal point cloud in the first channel are preserved point-by-point; feature fusion is performed on the feature encoding results of the first channel and the second channel, wherein the global features are concatenated with the point-by-point features of the internal point cloud through a broadcast mechanism to form a comprehensive feature representation that simultaneously includes local geometric details and global structural awareness.

[0012] According to one aspect of this disclosure, determining hemodynamic parameters based on the vessel segmentation results, and rendering the three-dimensional voxel structure of the vessel according to the hemodynamic parameters in accordance with the cardiac cycle, includes: drawing the blood flow streamline nodes on the three-dimensional voxel structure in topological order according to the spatial coordinates of the blood flow streamline nodes; generating an independent data structure representing a set of lines for each frame in the cardiac cycle, wherein each data structure representing a set of lines is composed of the blood flow streamline nodes in the three-dimensional voxel structure.

[0013] According to one aspect of this disclosure, the method of obtaining the semantic feature vector obtained according to natural language task instructions through a large language model interface (API) further includes: comparing one or more of the following: a first rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel, a second rendering result obtained by enhancing the rendering of the key region, and a third rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel according to the hemodynamic parameters according to the cardiac cycle, with a corresponding reference rendering result to obtain rendering loss parameters; and training the large language model interface based on the rendering loss parameters.

[0014] According to one aspect of this disclosure, the rendering processing method based on medical image segmentation further includes: comparing one or more of the following: a first rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel, a second rendering result obtained by enhancing the rendering of the key region, and a third rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel according to the hemodynamic parameters according to the cardiac cycle, with a corresponding reference rendering result to obtain rendering loss parameters; and training the first network and the second network according to the rendering loss parameters.

[0015] According to one aspect of this disclosure, a rendering processing system based on medical image segmentation is provided, comprising: a data management module configured to connect to an image database to retrieve a corresponding target image from the image database according to an input instruction; and an interactive reasoning-based large language model module configured to associate semantic feature vectors obtained according to natural language task instructions. Image feature vectors in the target image To generate joint feature vectors Based on the semantic feature vector with the highest confidence level The corresponding joint feature vector generates the initial value of the configuration parameter set, and based on the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values, a clarification question is raised to the user to obtain the updated values ​​of each segmentation parameter in the configuration parameter set. The system comprises: an image segmentation module, configured to perform segmentation processing on the target image based on the updated values ​​of each segmentation parameter in the determined configuration parameter set to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of the blood vessel; an image rendering module, configured to determine key regions in the three-dimensional voxel structure of the blood vessel based on the natural language task instructions and the user's clarification question, render the three-dimensional voxel structure of the blood vessel using a physically based volume rendering method, and enhance the rendering of the key regions; and a parameter extraction module, configured to determine hemodynamic parameters based on the blood vessel segmentation result, wherein the hemodynamic parameters include at least streamline data and spatial coordinates of blood flow streamline nodes, and the image rendering module is further configured to render the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters determined by the parameter extraction module.

[0016] According to one aspect of this disclosure, an electronic device is provided, characterized in that it includes: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the above-described rendering processing method based on medical image segmentation.

[0017] According to one aspect of this disclosure, a computer-readable storage medium is provided, characterized in that a computer program is stored on the computer-readable storage medium, the computer program being executed by a processor to perform the steps of the above-described rendering processing method based on medical image segmentation.

[0018] Therefore, the image processing method, system, device, and medium based on the semantic interaction model according to the embodiments of this disclosure can conveniently manage large amounts of patient multimodal image data, initiate clarification questions to users through semantic understanding, thereby quickly and accurately calling up images of cases, dimensions, and modalities, more accurately and intelligently segmenting the regions of interest of related images, calculating and outputting the required mechanical result parameters, and performing refined rendering and targeted display of the segmentation results, improving the readability of the segmentation results and reducing the user's interpretation burden. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some exemplary embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0020] Figure 1 This is a schematic diagram illustrating a rendering processing method based on medical image segmentation according to an embodiment of the present disclosure;

[0021] Figure 2 This is a schematic diagram illustrating a target image acquired from a PACS according to an embodiment of the present disclosure;

[0022] Figure 3 This is a schematic diagram illustrating the complete segmentation result of a blood vessel according to an embodiment of the present disclosure;

[0023] Figure 4 This is a schematic diagram showing a simplified segmentation result of a blood vessel according to an embodiment of the present disclosure;

[0024] Figure 5 This is a flowchart illustrating an example according to the present disclosure, using a second computational network to obtain a straightened blood vessel model based on the complete segmentation result to obtain a simplified segmentation result.

[0025] Figure 6 This illustrates an example of the predictive performance of transient hemodynamic parameters using a DGCNN structure, according to this disclosure.

[0026] Figure 7 A block diagram of a medical image segmentation-based rendering processing system according to an embodiment of the present disclosure is shown.

[0027] Figure 8 This is a schematic diagram illustrating the hardware structure of a device according to an embodiment of the present disclosure. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0029] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described object changes. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of some known functions and components are omitted.

[0030] This disclosure uses flowcharts to illustrate the steps of a method according to embodiments of this disclosure. It should be understood that the preceding or following steps are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0031] In the specification and drawings of this disclosure, elements are described in singular or plural forms according to embodiments. However, the singular and plural forms are suitably chosen for the presented cases merely for ease of explanation and are not intended to limit the disclosure thereto. Thus, a singular form may include a plural form, and a plural form may include a singular form, unless the context clearly indicates otherwise.

[0032] The rendering processing method, system, device, and medium based on medical image segmentation provided in this disclosure will now be described in detail with reference to the accompanying drawings. The semantic interaction model described in the embodiments of this disclosure can be deployed and run in medical imaging centers, research institutions, or cloud-based intelligent platforms, employing an architecture that combines local computing with cloud-based inference to meet the privacy protection requirements of medical data and the computational needs of large-scale model inference.

[0033] Figure 1 This is a schematic diagram illustrating a rendering processing method based on medical image segmentation according to an embodiment of the present disclosure. Figure 1 As shown, in step S101, the semantic feature vector obtained according to the natural language task instructions is associated. Image feature vectors in the target image To generate joint feature vectors ,in, This represents the coordinates of image voxels. For example, the semantic feature vector obtained by encoding instructions from a natural language task can be used first. With image feature vector The image features are projected onto a unified feature space to achieve dimensional and scale alignment. Then, semantic features are broadcast spatially, mapping them one-to-one with image features at the pixel or voxel level. Next, the semantic features are used to weight and enhance the image features, focusing the model on image regions highly relevant to the task instructions. Finally, the modulated image features and semantic features are fused using concatenation or weighting to generate a joint feature vector h.

[0034] For example, the semantic interaction model described in this embodiment can be connected to a hospital image archiving and communication system (PACS) or a research image database. The image feature vector of the target image in step S101 can be the image feature vector of an image stored in the PACS or research image database. ,in, This represents the voxel coordinates of the image. The target image is the medical image to be segmented.

[0035] According to one example of this disclosure, images stored in a PACS or scientific image database may include images of different modalities, such as computed tomography (CT), magnetic resonance imaging (MRI), magnetic resonance angiography (MRA), intravascular ultrasound (IVUS), optical coherence tomography (OCT), etc. Figure 2 This is a schematic diagram illustrating a target image acquired from a PACS according to an embodiment of the present disclosure.

[0036] Optionally, isotropic resampling can be performed on the input target medical image to ensure consistency in spatial resolution and intensity distribution, thereby improving the model's robustness to different scanning devices and protocols. Isotropic resampling unifies the original voxel size to a fixed interval through linear interpolation. This ensures that the scales of each vector are balanced. Furthermore, HU window normalization can be performed on medical images. HU normalization will adjust the pixel intensity... Mapping to standard range The data is then normalized to ensure that the input data meets the distribution requirements during model training.

[0037] In the example of this disclosure, semantic feature vector Can be combined with image feature vectors By integrating spatial parameters and anatomical attributes from images of different modalities, the system drives semantic reasoning. Thus, the semantic decision-making process not only relies on user language input but also dynamically references structural features from image data, achieving self-consistent reasoning between semantic understanding and image content. Through the fusion of spatial parameters and anatomical attributes from images of different modalities, the system can detect images in real time during semantic reasoning, thereby generating parameter configurations that best match the task.

[0038] In step S102, based on the semantic feature vector with the highest confidence... The corresponding joint feature vector generates initial values ​​for a set of configuration parameters, which includes one or more segmentation parameters. In one example of this disclosure, initial values ​​can be generated from one or more standard medical knowledge graphs, such as SNOMED CT, the Basic Anatomical Model (FMA), or the RadLex dictionary of radiology. This method identifies standardized mappings of nouns, disease entities, and functional parameters from natural language task instructions, thereby transforming the natural language task instruction representation into a structured semantic entity set.

[0039]

[0040] in, Indicates the first The identified semantic entities. Semantic feature vectors. It can be a set of semantic entities The continuous vectorization representation is obtained by jointly encoding one or more entities and their relationships in the entity set through a semantic encoder (such as a Transformer-based text encoding model or a knowledge-enhanced embedding model) to obtain a high-dimensional feature vector that can reflect the overall semantics of the task.

[0041] Based on this, conditional probability inference models can be constructed. For example, the input semantics and candidate configuration categories can be calculated. similarity between The probability distribution of the configuration parameters is obtained through the following normalization function (1):

[0042] ... (1)

[0043] in, This indicates that, given a semantic feature vector and knowledge graph Under the condition, configuration category The posterior probability; This represents the matching score between the semantic input and the configuration category; This represents the set of all candidate configuration categories.

[0044] The posterior probability of each candidate parameter set is calculated using a conditional probability model. And based on the semantic feature vector with the highest confidence. The corresponding joint feature vector generates the initial values ​​of the segmentation parameters in the configuration parameter set.

[0045] For example, when a user inputs the task command "Please help me divide the aortic arch", the semantic feature vector can be first identified according to step S101. For the "aortic arch," the standard anatomical entity ID corresponding to the target object "aortic arch" is identified in the semantic space, and an image region index associated with it is established. Subsequently, according to step S102, a set of configuration parameters is generated based on historical interaction experience and the joint feature vector h.

[0046]

[0047] in, Indicates the first Each segmentation parameter (such as segmentation range, true / false cavity detection, branch preservation threshold, output format, etc.) is configured, and the initial values ​​of each segmentation parameter in the configuration parameter set are determined.

[0048] In step S103, based on the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values, a clarification question can be proactively initiated to the user to obtain the updated values ​​of each segmentation parameter in the configuration parameter set.

[0049] For example, in step S103, it can be determined whether the consistency between the initial value of each segmentation parameter in the configuration parameter set and the image feature reference value meets a predetermined threshold. When the consistency between the value of each segmentation parameter in the configuration parameter set and the image feature reference value does not meet the predetermined threshold, a clarification question is actively initiated to the user, and the value of the parameter in the configuration parameter set is updated according to the user's feedback on the clarification question, until the consistency between the value of each segmentation parameter in the configuration parameter set and the image feature reference value meets the predetermined threshold.

[0050] According to one example of this disclosure, it is possible to configure a set of parameters. Each segmentation parameter The value of is consistent with the corresponding image feature reference value. For example, the consistency detection function can be shown in the following formula (2):

[0051] ... (2)

[0052] in, These are the weighting coefficients for each parameter. For the segmentation parameters, These are image feature reference values ​​recommended by the automatic image detection module or knowledge graph. For example, image feature reference values ​​could be reference indicators obtained from real images, such as local gray-level histograms, gradient distribution at the lumen edge, and vessel wall continuity. When necessary, the system can proactively ask the user for clarification, prompting them to input confirmation, correction, or supplementary information. The system will then update the settings based on the user's input, until the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values ​​meets a predetermined threshold.

[0053] Optionally, according to another example disclosed herein, in step S103, it may also be determined whether the values ​​of at least a portion of the segmentation parameters (e.g., key segmentation parameters) in the configuration parameter set exceed a confidence threshold. When the values ​​of at least a portion of the segmentation parameters (e.g., key segmentation parameters) in the configuration parameter set do not exceed the confidence threshold, a clarification question is proactively initiated to the user, and the configuration parameter set is updated based on the user's feedback on the clarification question, until the values ​​of at least a portion of the segmentation parameters (e.g., key segmentation parameters) in the configuration parameter set exceed the confidence threshold.

[0054] For example, each splitting parameter in the configuration parameter set can be configured. Calculate its confidence level If satisfied

[0055]

[0056] (in If a confidence threshold is set, then semantic understanding is deemed to be uncertain, and a clarification question is generated. For example, if the user does not specify a distinction between true and false cavities, a clarification question can be generated: "Is it necessary to distinguish between true and false cavities?" Then, the segmentation parameters can be updated based on user feedback.

[0057]

[0058] in, This indicates the amount of parameter updates caused by user feedback.

[0059] For example, when updating the segmentation parameters based on user input in each round, the confidence level of the segmentation parameters can be determined first, and then the consistency detection function can be determined, and vice versa.

[0060] According to another example of this disclosure, the question-and-answer sequence of each user interaction, the final configuration, and segmentation performance metrics can also be stored during runtime to build an experience memory:

[0061]

[0062] in, Represents the semantic vector of user input. For the corresponding parameter set, The system scores the quality of task execution. When a new task arrives, the system uses a similarity retrieval function:

[0063]

[0064] By referencing historical best configurations It also generates an initial candidate set, thereby reducing the number of user interaction rounds.

[0065] Furthermore, by utilizing spatial confidence maps from multimodal images (such as CT grayscale distribution, MRI tissue contrast, and IVUS parietal reflection features), regions with low local confidence are enhanced, increasing their priority in the next round of semantic questioning. This mechanism is able to proactively raise questions based on uncertain regions in real images, such as "Should this region contain arterial wall calcification?", thereby continuously improving segmentation accuracy during interaction.

[0066] In step S104, the target image is segmented according to the updated values ​​of each segmentation parameter in the configuration parameter set determined in step S103 to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of blood vessels.

[0067] According to one example of this disclosure, in step S104, a pre-segmentation model may first be executed based on the updated values ​​of each segmentation parameter in the set of configuration parameters determined in step S103. The segmentation results of the pre-segmentation of the target image can be quickly previewed at low resolution and provide preliminary information for fine segmentation at high resolution.

[0068] Optionally, according to an example of this disclosure, the pre-segmentation result can be sampled multiple times for uncertainty detection so that when the uncertainty of the pre-segmentation result does not meet a predetermined threshold, the user is prompted to further limit the input conditions or specify an image region.

[0069] For example, to ensure the reliability of the results, an uncertainty estimation algorithm based on Monte Carlo Dropout can be used to perform N forward samplings on the current configuration parameters to obtain the predicted output sequence.

[0070] And calculate the output variance:

[0071]

[0072] in, This indicates the uncertainty of the prediction result. Indicates the first The predicted output of the next sample. This is the average predicted value. When the variance... If the value is too large, it indicates that the model confidence is low, and the user can be prompted to further limit the input conditions or specify a specific image region in the target image.

[0073] Alternatively, in another example of this disclosure, the information entropy can be applied to each pre-segmented region. This generates an uncertainty heatmap, which visually represents the area most difficult for the model to determine and displays it to the user, where x represents a specific spatial location in the image, i.e., a voxel; It describes the magnitude of the uncertainty in the prediction result at that voxel; Indicated in voxels At this location, the model predicts that the voxel belongs to the first position. The probability values ​​for each category are displayed. These areas typically correspond to the edges of the vessel wall, distal small branches, or vessel inlets. Users can view the preview results and uncertainty distribution in real time on a 3D visualization interface, and interactively provide feedback or focus on key areas.

[0074] Then in step S104, the first computational network can be used to obtain a complete segmentation result for the target image based on the obtained pre-segmentation result; and the second computational network can be used to obtain a straightened blood vessel model based on the complete segmentation result to obtain a simplified segmentation result. Figure 3 This is a schematic diagram illustrating the complete segmentation result of a blood vessel according to an embodiment of the present disclosure. Figure 4 This is a schematic diagram showing a simplified segmentation result of a blood vessel according to an embodiment of the present disclosure.

[0075] According to the example of this disclosure, in step S104, a multi-scale cascaded segmentation network can be constructed through a multi-stage segmentation computational network, and the transition stage between the two algorithm networks can be optimized. Based on the complete segmentation result output by the first computational network, the curvature information of the blood vessels is extracted and stored as skeleton lines. Then, the blood vessel model is discretized and recombined based on the skeleton structure to obtain a straightened blood vessel model. The second computational network is then trained and tested based on the straightened blood vessel model. Preferably, the horizontal resolution of the image should not be lower than 96*96, and the number of layers is the same as the number of skeleton points.

[0076] Figure 5 This is a flowchart illustrating an example according to the present disclosure, using a second computational network to obtain a straightened blood vessel model based on the complete segmentation result to obtain a simplified segmentation result. (See flowchart for example.) Figure 5 As shown, in step S501, the skeleton line calculation of the blood vessel can be performed on the voxelized structure generated by the first computational network segmentation to form a discrete skeleton line with a step size of the blood vessel. In step S502, the discrete skeleton line of the blood vessel can be smoothed based on a loss function. For example, skeleton line smoothing involves three loss functions, the first of which constrains the magnitude of the smoothed movement of the skeleton points: (4)

[0077] in Indicates the first The moving weights of each skeleton point.

[0078] The second control can be the distance difference between different points.

[0079] (5)

[0080] in For the first The coordinate vectors of the skeleton points .

[0081] The third loss function controls the smoothness of the centerline:

[0082] (6)

[0083] Optionally, a comprehensive loss function can be established by linearly weighting the three loss functions:

[0084] (7)

[0085] in , They represent the first and the The coordinates of the skeleton points. These are hyperparameters used to balance the second loss term. . For each The weighting coefficients usually follow... Increase and decrease to control the smoothing intensity under different time lengths.

[0086] In step S503, a cross-section is extracted at the center point of each smoothed skeleton line of the blood vessel, and a 2D coordinate system is randomly established on the cross-section. Then, in step S504, all the discrete cross-sections are combined, and the randomly generated 2D coordinate system is rotated.

[0087] For example, in step S504, the following loss function can be used for rotation adjustment:

[0088] (8)

[0089] in This represents the loss value based on a plane or cross-section. It is a set of parameterized discrete cross-sectional point sequences. This represents the rotation parameter or other optimization parameter. By minimizing this loss function, the coordinate system can be fine-tuned by rotation, making the point sequence closer to a straight line. : Offset or step size parameter, used to select interval points in the point sequence and calculate the second difference under different step sizes. : Sequence of points for each discrete cross-section The length of the sequence is the number of points in the sequence. A sequence of single discrete cross-sectional points, belonging to An element in a set. Each It is a point in an ordered sequence, denoted as Each of them It is a 2D coordinate point. : Index in a point sequence, with values ​​ranging from arrive Used to iterate through possible combinations of three points in a sequence (start point, middle point, and end point). : point sequence The first in Each point is a 2D vector. Similarly, and They represent the first and the One point. The square of the Euclidean norm is used to calculate the square length of a vector in 2D space. In the formula, it measures three points. , , The degree of collinearity. If the three points are collinear, the value is zero; otherwise, a positive value indicates the degree of curvature or deviation from a straight line.

[0090] In step S505, the blood vessel is straightened and mapped based on the combined and rotated cross-section to generate a straightened image. The straightened image is then segmented and inversely transformed to obtain a simplified segmentation result. Because the forward sampling of the straightened image is sparsity, the segmentation result will have structural pores during inverse mapping. Therefore, according to an example of this disclosure, during the inverse transformation, the segmentation result can first be binarized and inverted, and a distance transform can be calculated for each point within the pores to obtain the filling value of the nearest foreground voxel.

[0091] Furthermore, according to an example of this disclosure, a large-scale smoothing term corresponding to the entire centerline segment can be added to the straightening algorithm. This large-scale smoothing term enhances the smoothing effect on centerline segments with high curvature, preventing the cross-section from extending into the blood vessel due to abnormal morphology of local lesions, thus avoiding adverse straightening results. Alternatively, the coefficients of smoothing constraints acting on the local neighborhood or short scale of the centerline segment can be retained to control the smoothing coefficients of the large-scale smoothing term in the proximal, distal, and high-curvature regions. The coefficients of the small smoothing term can primarily be used to maintain the local geometric continuity and detail consistency of the centerline.

[0092] In addition, the coefficient formula shown in Equation (9) can be used to establish the trend of unimodal coefficient variation, and the sampling scheme can be adaptively adjusted according to the number of points on the center line.

[0093] (9)

[0094] Among them, the length l of the center line can control the selection of k. The longer the length, the lower the global coefficient decay, the greater the influence of the large-scale control terms, and the stronger the smoothing force. However, when the length exceeds the prior threshold of 700, the decay force is 0. The length l can also control γ. This parameter can change the relative difference of the control coefficients in different regions of the center line. As the center line extends, the relative difference of the coefficients on both sides centered on the peak point will gradually decrease.

[0095] Furthermore, according to one example of this disclosure, the centerline can also be extended in the aortic sinus region, and the direction of the smoothed centerline extension can be controlled by the first principal component direction of the sinus region voxel coordinates.

[0096] In an example according to this disclosure, the second computational network augments the data by randomly rotating the image in a plane, compensating for the data reduction caused by full-image training.

[0097] According to an example of this disclosure, after step S505, the sensitivity of the simplified segmentation result to different coefficient orders can be calculated to evaluate the accuracy of the simplified segmentation result. The output of the simplified segmentation result is only used as the segmentation result when the accuracy meets a predetermined value. For example, the sensitivity of the normal vessel segmentation result to the straightening parameter can be evaluated using two optimization strategies: abnormal centerline path and abnormal aortic sinus direction. When the coefficient verification result shows significant sensitivity, the centerline smoothing calculation, straightening mapping, and inverse transformation calculation can be re-performed.

[0098] According to another example of this disclosure, the accuracy of the simplified segmentation results can be evaluated based on the ResBlock network structure and U-Net, SegNet, and DeepLab models. Furthermore, segmentation accuracy can be represented by the Dice coefficient. Using the Dice coefficient as an evaluation metric, a number of randomly selected cases are used as a test set to verify the segmentation accuracy of the lesion region in preoperative imaging, the segmentation accuracy of lesion regions with small postoperative structural volume, and the overall segmentation accuracy of blood vessels before and after surgery. Optionally, the selected cases may be no fewer than 200. Additionally, when any Dice coefficient is less than 0.9, the multi-stage segmentation calculation is repeated.

[0099] In this example, by using semantic understanding and proactively initiating multiple rounds of interactive probing, the corresponding cases and / or images can be quickly and accurately invoked and preprocessed, and the regions of interest in the relevant images can be segmented more accurately.

[0100] In step S105, key regions in the three-dimensional voxel structure of the blood vessel are determined according to the natural language task instructions and the user's clarification question. The three-dimensional voxel structure of the blood vessel is rendered using a physical volume rendering method, and the key regions are enhanced.

[0101] For example, in step S105, the multi-head attention system of a large language model based on the Transformer architecture can be used to capture and integrate the key needs from the user's natural language task instructions in step S101 and the user's feedback on the clarifying question in step S103. Then, based on a semantic-spatial alignment algorithm using a graph attention network (GAT), the semantic nodes in the key needs of the text description are associated with the three-dimensional voxel structure of the blood vessels obtained in step S104. For example, the attention coefficient between each voxel node and the semantic nodes in its anatomical structure. As shown in the following formula (10):

[0102] (10)

[0103] in and These represent the feature vectors of a 3D voxel node and a semantic node, respectively. For learnable weight matrix, This is the attention parameter vector. This enables semantic-to-voxel structure alignment, accurately identifying key regions based on the key needs of both the user's natural language task instructions and the user's feedback to clarify questions. In other words, it can be based on attention coefficients. To identify key areas. For example, The larger the value, the more critical the corresponding 3D voxel node. By calculating the attention coefficient between the 3D voxel node and the semantic node, the distribution of voxel importance under semantic guidance can be obtained. Voxel nodes with higher attention coefficients indicate that they are more relevant to the current semantic task. Based on the distribution of these attention coefficients, the system can identify key spatial regions that are highly relevant to the natural language task instructions, thereby guiding subsequent segmentation, analysis, or interactive operations.

[0104] According to an example of this disclosure, before the physically based volume rendering method renders the three-dimensional voxel structure of the blood vessel in step S105, the segmentation mask can be smoothed by anisotropic diffusion filtering to suppress noise and preserve important anatomical boundaries.

[0105] According to another example of this disclosure, key areas indicated by the user are enhanced by adjusting the transfer function to assign higher opacity and a more pronounced color gamut to the key structures. Alternatively, the boundaries of the key areas can be determined by calculating the dot product of the viewpoint direction and the normal, facilitating real-time contour detection and enhanced rendering of the boundaries of the key areas. For example, the viewpoint direction can be the direction vector from the current viewing point to a point on the surface of the structure, and the normal can be the normal direction at that surface point. By analyzing the relationship between the viewpoint direction and the normal, the areas where the structure forms a contour under the current viewpoint can be identified. Furthermore, dark highlights can be applied to the contour lines, and anti-aliasing can be performed.

[0106] According to yet another example of this disclosure, in step S105, the surface visual details of the three-dimensional voxel structure of the blood vessel can also be enhanced by real-time calculation of the illumination model:

[0107] (11)

[0108] Where I_α, I_d, and I_s represent the ambient light, diffuse light, and specular light components, respectively, and k_α, k_d, and k_s are the corresponding reflection coefficients. It is the normal vector. The direction of the reflected light. The direction of the line of sight.

[0109] Therefore, in step S105, by mapping the semantics from the user to the rendering parameters, highly personalized and interpretable post-processing results of three-dimensional medical images can be provided while ensuring the accuracy of the anatomy.

[0110] In step S106, hemodynamic parameters are determined based on the blood vessel segmentation results, and the three-dimensional voxel structure of the blood vessel is rendered according to the cardiac cycle based on the hemodynamic parameters, wherein the hemodynamic parameters include at least the streamline data and spatial coordinates of the blood flow streamline nodes.

[0111] For example, a Dynamic Graph Neural Network (DGCNN) can be constructed from the simplified 3D point cloud of the segmentation result to extract hemodynamic parameters. For instance, a DGCNN can be constructed with two parallel and independent channels: one channel for the interior point cloud, denoted as {Ii | i = 1, 2, ..., N1}; and the other channel for the wall point cloud, denoted as {Wj | j = 1, 2, ..., N2}, where i and j represent the indices of the interior and surface point clouds, respectively, and N1 and N2 represent the number of interior and surface points in the network input, corresponding to the number of neurons in the input layer. The simplified 3D point cloud of the segmentation result can be used as input data for the DGCNN.

[0112] The constructed DGCNN can predict the instantaneous velocity and pressure field distribution inside blood vessels based on the input data. For example, the output tensor of DGCNN can have four channels of components, where the components of the four channels correspond to the velocity components in the x, y, and z directions and the pressure scalar value, respectively.

[0113] According to an example of this disclosure, DGCNN encodes features from point clouds input in two channels separately. For example, in each channel, the input point cloud is processed sequentially through multiple Edge Convolution (EdgeConv) and Multi-Layer Perceptron (MLP) modules. EdgeConv constructs the connectivity of local point clouds based on an adjacency graph structure, extracting geometric and topological information from the local neighborhood by combining node and edge features, and can adaptively update the graph structure during training to enhance feature representation. The MLP module consists of a multi-layer fully connected network and a non-linear activation function, performing high-dimensional mapping and non-linear transformation on point-by-point features, thereby further improving the discriminative power of point cloud features. Through the layer-by-layer stacking of multiple levels of EdgeConv and MLP, the network can fuse local and global features to achieve multi-scale information extraction. Ultimately, the feature dimensions of both the internal and surface point cloud channels can be expanded to at least 4096. Among them, the surface point cloud channel aggregates features through global max pooling to obtain rotation and translation invariance and generates a global feature tensor with a shape that is 1 times the feature dimension; the internal point cloud channel retains the point-by-point feature representation with a shape that is N1 times the feature dimension to ensure that each internal point can obtain global context information.

[0114] Then, DGCNN performs feature fusion on the feature encoding results of the two channels. The global features of the surface point cloud are concatenated with the point-by-point features of the internal point cloud through a broadcast mechanism to form a comprehensive feature representation that simultaneously includes local geometric details and global structural awareness. This fused feature is then input into a regression module composed of multiple MLPs, which outputs the velocity three components and pressure prediction for each internal point. It is worth noting that all MLP layers in the network use one-dimensional convolutions with a kernel size of 1 to achieve efficient point-by-point feature processing and reduce computational overhead. Finally, the network outputs the velocity field and pressure field prediction results corresponding to each sampling point inside the blood vessel. For example, the loss function of the network shown in the following formula (12) can be used to ensure that the prediction results converge to the Convolutional Feature Descriptor (CFD). Simulation results:

[0115] (12)

[0116] Furthermore, a normalization can be performed based on formula (12) to evaluate the performance metrics of the model, as shown in the following formula (13):

[0117] (13)

[0118] In the above formulas (12) and (13), i is the index of a sample point, N represents the number of point clouds, and P_i and P̂_i represent the CFD result and the neural network prediction result, respectively.

[0119] The DGCNN structure described in the above embodiments exhibits superior performance in predicting transient hemodynamic parameters. In addition to transient hemodynamic parameters, the time-averaged parameter field of the aorta also has significant clinical reference value in aortic disease risk assessment. Unlike transient results, the mean field reflects the hemodynamic characteristics throughout a complete cardiac cycle, avoiding the influence of transient fluctuations, and thus better reveals the long-term stress state and flow field distribution characteristics of the blood vessel.

[0120] Alternatively, CFD simulations of the entire cardiac cycle can be used, and the distributions of the mean pressure field, mean velocity field, and mean wall shear stress (TAWSS) can be obtained through time integration and normalization. The DGCNN model's prediction performance on this task is as follows: Figure 6 As shown, the qualitative comparison results demonstrate that the model can accurately reproduce the overall distribution trend and local detailed features of the CFD simulation, maintaining high consistency, especially in complex blood flow areas such as the aortic arch and branch vessel inlets. The normalized mean absolute error (NMAE) of the mean pressure field prediction was calculated to evaluate the accuracy and robustness of DGCNN in capturing global steady-state flow patterns during the cardiac cycle, providing a reliable computational tool for long-term risk assessment of aortic disease.

[0121] According to another example of this disclosure, in step S106, streamline nodes can be drawn on a three-dimensional voxel structure according to their topological order based on their spatial coordinates. For simulation results covering the entire cardiac cycle, streamline data typically contains multiple time frames. An independent data structure representing a set of lines (e.g., a LineSystem in graphics and computer-aided design (CAD)) can be generated for each frame, where the data structure representing the set of lines can be composed of individual streamline nodes. When displaying the rendering results, the rendered data structures representing the set of lines can be displayed sequentially according to the time order within the cardiac cycle, allowing the observer to clearly perceive the evolution of blood flow patterns within the cardiac cycle. For example, based on the HSL (Hue-Saturation-Luminance) color space, velocity values ​​can be linearly mapped to hue, allowing streamline segments with different velocities to exhibit continuous, gradual color changes, thereby achieving intuitive visualization of velocity information. As another example, to enhance streamline visualization and facilitate observation of internal blood flow structures, the transparency of the aortic geometry model can be increased, allowing users to visually view the flow path inside the blood vessel from the outside.

[0122] Given the large size of the original pipeline data files, potentially reaching hundreds of megabytes, direct storage would result in significant loading latency and cloud storage overhead. According to an example disclosed herein, the geometric coordinates and color attributes of each node in each frame of the LineSystem can be extracted, encoded, and stored. The frame index and pipeline number of each node are recorded as a two-dimensional array in the metadata portion of the stored file, enabling rapid data recovery during decoding. This allows the approximately 100 MB original file to be compressed to approximately 100 KB, significantly reducing storage space usage while improving the efficiency of subsequent data loading and network transmission.

[0123] According to one example of this disclosure, when training a large language model interface (API) to obtain semantic feature vectors and a computer network for training blood vessel segmentation and rendering, a joint optimization strategy can be used to achieve synergistic convergence of semantic extraction and image processing.

[0124] For example, the consistency optimization of semantic extraction, segmentation and rendering results can be achieved through the joint loss function shown in the following formula (14):

[0125] L_joint=L_seg+μ_1 L_sem+μ_2 L_rend (14)

[0126] Where L_seg is the segmentation loss function (such as Dice Loss or CrossEntropy Loss), used to optimize the contour accuracy of the target region; L_sem is the semantic consistency loss, used to maintain the coherence of semantic decisions after multiple rounds of interaction; L_rend is the rendering consistency loss, used to ensure visual continuity under different modalities or viewpoints; μ_1 and μ_2 are loss balance coefficients. The segmentation loss can be used to constrain the accuracy of structural boundaries, the semantic consistency loss is used to constrain the stability of semantic decisions under multiple rounds of interaction and multiple viewpoints, and the rendering consistency loss is used to constrain the visual continuity of rendering results under different modalities or viewing conditions.

[0127] By optimizing the joint loss function, the system can reverse the semantic feature extraction process based on the segmentation or rendering results, and can also guide the segmentation and rendering processes based on semantic features, thereby achieving synergistic convergence between semantic understanding and image processing.

[0128] The rendering consistency loss can be represented by the following formula (15):

[0129] (15)

[0130] Where R_pred(x) is the intensity of the predicted rendered image generated by the model, and R_gt(x) is the reference rendered result. The intensity of the rendered image can be the brightness or color intensity value of a pixel or voxel generated in the display space after volume rendering or surface rendering processes, used to characterize the visual representation of the structure in the final rendered result, rather than the sampling intensity of the original medical image.

[0131] The above combination Figure 1-6 The rendering processing method based on medical image segmentation provided in this disclosure is described below, and will be combined with... Figure 7 The rendering and processing system based on medical image segmentation provided in this disclosure. Because... Figure 7 The rendering and processing system 700 based on medical image segmentation shown above is combined with the above. Figure 1-6 The description corresponds to the rendering processing method based on medical image segmentation, so for simplicity, a detailed description of the same content is omitted here.

[0132] Figure 7 A block diagram of a rendering processing system based on medical image segmentation according to an embodiment of the present disclosure is shown. Figure 7 As shown, the medical image segmentation-based rendering processing system 700 according to an embodiment of this disclosure may include a data management module 710, an interactive reasoning large language model module 720, an image segmentation module 730, an image rendering module 740, and a parameter extraction module 750.

[0133] The data management module 710 can connect to a PACS or scientific image database to retrieve the corresponding target image based on input commands from the user or the interactive reasoning large language model module 720. According to one example of this disclosure, the images retrieved by the data management module 710 from the PACS or scientific image database may include images of different modalities, such as computed tomography (CT), magnetic resonance imaging (MRI), magnetic resonance angiography (MRA), intravascular ultrasound (IVUS), optical coherence tomography (OCT), etc.

[0134] The interactive reasoning-based large language model module 720 can associate semantic feature vectors obtained from natural language task instructions. Image feature vectors in the target image To generate joint feature vectors ,in, This represents the image voxel coordinates. For example, the interactive reasoning large language model module 720 described in this embodiment can be connected to a PACS or scientific image database via the data management module 710.

[0135] In the example of this disclosure, semantic feature vector Can be combined with image feature vectors By integrating spatial parameters and anatomical attributes from images of different modalities, the system drives semantic reasoning. Thus, the semantic decision-making process not only relies on user language input but also dynamically references structural features from image data, achieving self-consistent reasoning between semantic understanding and image content. Through the fusion of spatial parameters and anatomical attributes from images of different modalities, the system can detect images in real time during semantic reasoning, thereby generating parameter configurations that best match the task.

[0136] The interactive reasoning large language model module 720 can also base its model on the semantic feature vector with the highest confidence level. The corresponding joint feature vector generates initial values ​​for a set of configuration parameters, which includes one or more segmentation parameters. In one example of this disclosure, initial values ​​can be generated from one or more standard medical knowledge graphs, such as SNOMED CT, the Basic Anatomical Model (FMA), or the RadLex dictionary of radiology. This method identifies standardized mappings of nouns, disease entities, and functional parameters from natural language task instructions, thereby transforming the natural language task instruction representation into a structured semantic entity set.

[0137]

[0138] in, Indicates the first The identified semantic entities.

[0139] Based on this, conditional probability inference models can be constructed. For example, the input semantics and candidate configuration categories can be calculated. similarity between The probability distribution of the configuration parameters is obtained through the following normalization function (1):

[0140] ... (1)

[0141] in, This indicates that, given a semantic feature vector and knowledge graph Under the condition, configuration category The posterior probability; This represents the matching score between the semantic input and the configuration category; This represents the set of all candidate configuration categories.

[0142] The posterior probability of each candidate parameter set is calculated using a conditional probability model. And based on the semantic feature vector with the highest confidence. The corresponding joint feature vector generates the initial values ​​of the segmentation parameters in the configuration parameter set.

[0143] For example, when a user inputs the task command "Please help me divide the aortic arch", the interactive reasoning large language model module 720 can recognize semantic feature vectors. For the term "aortic arch," the standard anatomical entity ID corresponding to the target object "aortic arch" is identified in the semantic space, and an image region index associated with it is established. Subsequently, the interactive reasoning-based large language model module 720 can generate a set of configuration parameters based on historical interaction experience and the joint feature vector h.

[0144]

[0145] in, Indicates the first Each segmentation parameter (such as segmentation range, true / false cavity detection, branch preservation threshold, output format, etc.) is configured, and the initial values ​​of each segmentation parameter in the configuration parameter set are determined.

[0146] In addition, the interactive reasoning large language model module 720 can proactively ask the user clarification questions based on the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values, so as to obtain the updated values ​​of each segmentation parameter in the configuration parameter set.

[0147] For example, the interactive reasoning large language model module 720 can determine whether the consistency between the initial value of each segmentation parameter in the configuration parameter set and the image feature reference value meets a predetermined threshold, and when the consistency between the value of each segmentation parameter in the configuration parameter set and the image feature reference value does not meet the predetermined threshold, it can proactively initiate a clarification question to the user, and update the value of the parameter in the configuration parameter set according to the user's feedback on the clarification question, until the consistency between the value of each segmentation parameter in the configuration parameter set and the image feature reference value meets the predetermined threshold.

[0148] According to one example of this disclosure, the interactive reasoning large language model module 720 can be configured based on a set of parameters. Each segmentation parameter The value of is consistent with the corresponding image feature reference value. For example, the consistency detection function can be as shown in formula (2) above.

[0149] For example, image feature reference values ​​can be reference indicators obtained from real images, such as local gray-level histograms, gradient distribution at the lumen edge, and vessel wall continuity. When necessary, the system can proactively ask the user for clarification, prompting them to input confirmation, correction, or supplementary information. The system will then update the settings based on the user's input, until the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values ​​meets a predetermined threshold.

[0150] The image segmentation module 730 can perform segmentation processing on the target image according to the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of the blood vessel.

[0151] According to one example of this disclosure, the image segmentation module 730 may first perform a pre-segmentation model on the target image based on the updated values ​​of each segmentation parameter in the determined set of configuration parameters. The segmentation results of the pre-segmentation of the target image can be quickly previewed at low resolution and provide preliminary information for fine-grained segmentation at high resolution.

[0152] Optionally, according to an example of this disclosure, the pre-segmentation result can be sampled multiple times for uncertainty detection so that when the uncertainty of the pre-segmentation result does not meet a predetermined threshold, the user is prompted to further limit the input conditions or specify an image region.

[0153] Furthermore, in another example of this disclosure, each pre-segmented region can be determined based on information entropy. This generates an uncertainty heatmap, which visually represents and displays the areas most difficult to determine in the model. These areas typically correspond to the edges of the vessel walls, distal small branches, or vessel inlets. Users can view the preview results and uncertainty distribution in real time on the 3D visualization interface and interactively provide feedback or focus on key areas.

[0154] Then, the image segmentation module 730 can use the first computational network to obtain a complete segmentation result for the target image based on the obtained pre-segmentation result; and use the second computational network to obtain a straightened blood vessel model based on the complete segmentation result to obtain a simplified segmentation result. Figure 3 This is a schematic diagram illustrating the complete segmentation result of a blood vessel according to an embodiment of the present disclosure. Figure 4 This is a schematic diagram showing a simplified segmentation result of a blood vessel according to an embodiment of the present disclosure.

[0155] According to the examples of this disclosure, the image segmentation module 730 can construct a multi-scale cascaded segmentation network through a multi-stage segmentation computational network and optimize the transition stage between the two algorithm networks. Based on the complete segmentation result output by the first computational network, the curvature information of the blood vessels is extracted and stored as skeleton lines. Then, based on the skeleton structure, the blood vessel model is discretized and recombined to obtain a straightened blood vessel model. The second computational network is then trained and tested based on the straightened blood vessel model. Preferably, the horizontal resolution of the image should not be lower than 96*96, and the number of layers is the same as the number of skeleton points. For example, the image segmentation module 730 can, according to... Figure 5 The method shown uses a second computational network to obtain a straightened blood vessel model based on the complete segmentation results, resulting in a flowchart of the simplified segmentation results.

[0156] Optionally, according to an example of this disclosure, system 700 may further include an accuracy verification module. The accuracy verification module calculates the sensitivity of simplified segmentation results to different coefficient orders to evaluate the accuracy of the calculated simplified segmentation results, and only uses the output of the simplified segmentation results as the segmentation result when the accuracy meets a predetermined value. For example, the sensitivity of normal vessel segmentation results to straightening parameters can be evaluated using two optimization strategies: abnormal centerline path and abnormal aortic sinus direction. When the coefficient verification results show significant sensitivity, centerline smoothing calculation, straightening mapping, and inverse transformation calculation can be re-performed.

[0157] Furthermore, the accuracy verification module can evaluate the accuracy of simplified segmentation results based on the ResBlock network structure and U-Net, SegNet, and DeepLab models. Additionally, segmentation accuracy can be represented by the Dice coefficient. Using the Dice coefficient as the evaluation metric, several randomly selected cases were used as a test set to verify the segmentation accuracy of lesion regions in preoperative imaging, the segmentation accuracy of lesion regions with small postoperative structural volume, and the overall segmentation accuracy of blood vessels both preoperatively and postoperatively.

[0158] The image rendering module 740 can determine the key regions in the three-dimensional voxel structure of the blood vessels according to the natural language task instructions and the user's clarification questions, render the three-dimensional voxel structure of the blood vessels using a physical volume rendering method, and enhance the rendering of the key regions.

[0159] For example, the image rendering module 740 can utilize the multi-head attention system of a large language model based on the Transformer architecture to capture and integrate key needs from the user's natural language task instructions and the user's feedback on clarifying questions. Then, based on a semantic-spatial alignment algorithm using a graph attention network (GAT), the anatomical structures in the key needs of the text description are associated with the three-dimensional voxel structures of blood vessels obtained by the image segmentation module 730.

[0160] According to one example of this disclosure, before rendering the three-dimensional voxel structure of the blood vessel using a physically based volume rendering method, the image rendering module 740 can smooth the segmentation mask by anisotropic diffusion filtering to suppress noise and preserve important anatomical boundaries.

[0161] According to another example of this disclosure, key regions indicated by the user are enhanced by adjusting the transfer function to impart higher opacity and a more pronounced color gamut to the key structures. Alternatively, the boundaries of the key regions can be determined by calculating the dot product of the viewpoint direction and the normal, facilitating real-time contour detection and enhanced rendering of the boundaries of the key regions. For example, dark highlights can be applied to the contours, and anti-aliasing can be applied.

[0162] According to another example of this disclosure, the image rendering module 740 can also enhance the surface visual details of the three-dimensional voxel structure of blood vessels by, for example, by calculating a lighting model in real time as shown in formula (11) above. Thus, the image rendering module 740 can provide highly personalized and interpretable three-dimensional medical image post-processing results while ensuring anatomical accuracy by mapping semantics from the user to the rendering parameters.

[0163] The parameter extraction module 750 can determine hemodynamic parameters based on the blood vessel segmentation results, wherein the hemodynamic parameters include at least the streamline data and spatial coordinates of the blood flow streamline nodes. The image rendering module 740 can render the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters determined by the parameter extraction module 750.

[0164] For example, the parameter extraction module 750 can construct a Dynamic Graph Neural Network (DGCNN) for the simplified segmentation result's 3D point cloud to extract hemodynamic parameters. For instance, the DGCNN can be constructed with two parallel and independent channels: one channel for the interior point cloud, denoted as {Ii | i = 1, 2, ..., N1}; and the other channel for the wall point cloud, denoted as {Wj | j = 1, 2, ..., N2}, where i and j represent the indices of the interior and surface point clouds, respectively, and N1 and N2 represent the number of interior and surface points in the network input, corresponding to the number of neurons in the input layer. The simplified segmentation result's 3D point cloud can be used as input data for the DGCNN.

[0165] The DGCNN constructed by the parameter extraction module 750 can predict the instantaneous velocity and pressure field distribution inside the blood vessel based on the input data. For example, the output tensor of the DGCNN can have four channel components, where the components of the four channels correspond to the velocity components in the x, y, and z directions and the pressure scalar value, respectively.

[0166] According to an example of this disclosure, DGCNN encodes features from point clouds input in two channels separately. For example, in each channel, the input point cloud is processed sequentially through multiple Edge Convolution (EdgeConv) and Multi-Layer Perceptron (MLP) modules. EdgeConv constructs the connectivity of local point clouds based on an adjacency graph structure, extracting geometric and topological information from the local neighborhood by combining node and edge features, and can adaptively update the graph structure during training to enhance feature representation. The MLP module consists of a multi-layer fully connected network and a non-linear activation function, performing high-dimensional mapping and non-linear transformation on point-by-point features, thereby further improving the discriminative power of point cloud features. Through the layer-by-layer stacking of multiple levels of EdgeConv and MLP, the network can fuse local and global features to achieve multi-scale information extraction. Ultimately, the feature dimensions of both the internal and surface point cloud channels can be expanded to at least 4096. Among them, the surface point cloud channel aggregates features through global max pooling to obtain rotation and translation invariance and generates a global feature tensor with a shape that is 1 times the feature dimension; the internal point cloud channel retains the point-by-point feature representation with a shape that is N1 times the feature dimension to ensure that each internal point can obtain global context information.

[0167] Then, the parameter extraction module 750 performs feature fusion on the feature encoding results of the two channels. The global features of the surface point cloud are concatenated with the point-by-point features of the internal point cloud through a broadcast mechanism to form a comprehensive feature representation that simultaneously includes local geometric details and global structure awareness. This fused feature is then input into a regression module composed of multiple MLP layers, which outputs the velocity three components and pressure predictions for each internal point. It is worth noting that all MLP layers in the network use one-dimensional convolutions with a kernel size of 1 to achieve efficient point-by-point feature processing and reduce computational overhead. Finally, the network outputs the velocity and pressure field predictions for each sampling point inside the blood vessel.

[0168] The DGCNN structure described in the above embodiments exhibits superior performance in predicting transient hemodynamic parameters. In addition to transient hemodynamic parameters, the time-averaged parameter field of the aorta also has significant clinical reference value in aortic disease risk assessment. Unlike transient results, the mean field reflects the hemodynamic characteristics throughout a complete cardiac cycle, avoiding the influence of transient fluctuations, and thus better reveals the long-term stress state and flow field distribution characteristics of the blood vessel.

[0169] Optionally, the parameter extraction module 750 can use CFD simulation results of the complete cardiac cycle and obtain the distributions of mean pressure field, mean velocity field, and mean wall shear stress (TAWSS) through time integration and normalization. The DGCNN model's prediction performance on this task is as follows: Figure 6 As shown, the qualitative comparison results demonstrate that the model can accurately reproduce the overall distribution trend and local detailed features of the CFD simulation, maintaining high consistency, especially in complex blood flow areas such as the aortic arch and branch vessel inlets. By calculating the normalized mean absolute error (NMAE) of the mean pressure field prediction, the accuracy and robustness of DGCNN in capturing global steady-state flow patterns during the cardiac cycle are evaluated, providing a reliable computational tool for long-term risk assessment of aortic disease.

[0170] The image rendering module 740 can draw the streamline nodes on a three-dimensional voxel structure according to their topological order based on the spatial coordinates of the blood flow streamline nodes obtained by the parameter extraction module 750. For simulation results covering the entire cardiac cycle, streamline data typically contains multiple time frames. An independent data structure representing a set of lines (e.g., a LineSystem in graphics and computer-aided design (CAD)) can be generated for each frame, where the data structure representing the set of lines can be composed of individual streamline nodes. When displaying the rendering results, the rendered data structures representing the set of lines can be displayed sequentially according to the time order within the cardiac cycle, allowing the observer to clearly perceive the evolution of blood flow patterns within the cardiac cycle. For example, based on the HSL (Hue-Saturation-Luminance) color space, the flow velocity value can be linearly mapped to hue, allowing streamline segments with different flow velocities to exhibit continuous and gradual color changes, thereby achieving intuitive visualization of velocity information. Furthermore, to enhance streamline visualization and facilitate observation of internal blood flow structures, the transparency of the aortic geometry model can be increased, allowing users to visually observe the flow path inside the blood vessel from the outside.

[0171] Given the large size of the original streamline data file, potentially reaching hundreds of megabytes, direct storage would result in significant loading latency and cloud storage overhead. According to an example disclosed herein, the image rendering module 740 can extract the geometric coordinates and color attributes of each node in each frame of the LineSystem, encode and store them, and record the frame index and streamline number of the nodes in the metadata portion of the stored file as a two-dimensional array, thereby enabling rapid data recovery during decoding. This allows the approximately 100 MB original file to be compressed to approximately 100 KB, significantly reducing storage space usage while improving the efficiency of subsequent data loading and network transmission.

[0172] Furthermore, the block diagrams used in the above embodiments illustrate blocks based on functionality. These functional blocks (structural units) are implemented through any combination of hardware and / or software. Moreover, the means of implementing each functional block are not particularly limited. That is, each functional block can be implemented using a single device that is physically and / or logically combined, or it can be implemented using multiple devices by directly and / or indirectly (e.g., via wired and / or wireless) connecting two or more physically and / or logically separate devices.

[0173] For example, the device of one embodiment of this disclosure can function as a computer performing the image processing method based on the semantic interaction model of this disclosure. Figure 8 This is a schematic diagram illustrating the hardware structure of a device 800 according to an embodiment of the present disclosure. The device 800 described above can be configured as a computer device that physically includes a processor 810, a memory 820, an input / output device 830, a bus 840, etc.

[0174] Additionally, in the following description, the word "device" can be replaced with circuit, device, unit, etc. The hardware structure of the terminal may include one or more of the devices shown in the figures, or may not include some of the devices.

[0175] For example, only one processor 810 is shown, but there can also be multiple processors. Furthermore, processing can be performed by a single processor, or by more than one processor simultaneously, sequentially, or using other methods. Additionally, processor 810 can be mounted on more than one chip.

[0176] The functions of device 800 are implemented, for example, by reading instructions (programs) stored in memory 820 into hardware such as processor 810, thereby enabling processor 810 to perform operations, control input / output device 830, and control the reading and / or writing of data in memory 820.

[0177] The processor 810, for example, enables the operating system to operate, thereby controlling the computer as a whole. The processor 810 may be composed of a central processing unit (CPU) that includes interfaces with peripheral devices, control devices, arithmetic devices, registers, etc. For example, the aforementioned processing units can be implemented by the processor 810.

[0178] Furthermore, the processor 810 reads programs (program code), data, etc., from the memory 820 and performs various processes accordingly. The program can be one that causes the computer to perform at least a portion of the actions described in the above embodiments. For example, the method executed by the first AP can be implemented using a control program stored in the memory 820 and operated by the processor 810. Similarly, the method executed by the second AP can also be implemented using a control program stored in the memory 820 and operated by the processor 810.

[0179] The memory 820 may be a computer-readable recording medium, such as at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically programmable read-only memory (EEPROM), a random access memory (RAM), or other suitable storage media. The memory 820 may include registers, caches, main memory (main storage device), etc. The memory 820 may store executable programs (program code), software modules, etc., for implementing the methods according to an embodiment of this disclosure.

[0180] In addition, the memory 820 may also include, for example, a computer-readable recording medium consisting of at least one of a flexible disk, a floppy disk, a magneto-optical disk (e.g., a read-only optical disc (CD-ROM, etc.), a digital universal optical disc, a Blu-ray disc), a removable disk, a hard disk, a smart card, a flash memory device (e.g., a card, a stick, a key driver), a magnetic stripe, a database, a server, or other suitable storage media.

[0181] The input / output device 830 may include an input unit (e.g., keyboard, mouse, microphone, switch, button, sensor, etc.) that accepts input from the outside, and an output unit (e.g., display, speaker, light-emitting diode (LED) lamp, etc.) that performs output to the outside. Alternatively, the input unit and the output unit may be integrated into one structure (e.g., a touch panel).

[0182] Furthermore, the processor 810, memory 820, input / output device 830, and other devices are connected via a bus 840 for communication of information. The bus 840 can consist of a single bus or different buses between devices.

[0183] Furthermore, device 800 may include hardware such as a microprocessor, digital signal processor (DSP), application-specific integrated circuit (ASIC), programmable logic device (PLD), and field-programmable gate array (FPGA), which can be used to implement some or all of the functional blocks. For example, processor 810 can be installed using at least one of these hardware components.

[0184] The present disclosure has been described in detail above, but it will be apparent to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modifications and variations without departing from the spirit and scope of the invention as defined by the claims. Therefore, the description in this disclosure is for illustrative purposes and is not intended to be restrictive.

[0185] In this disclosure, whether the software is referred to as software, firmware, middleware, microcode, hardware description language, or any other name, it should be broadly interpreted as meaning instruction, instruction set, code, code segment, program code, program, subroutine, software module, application, software application, software package, routine, subroutine, object, executable file, execution thread, procedure, function, etc.

[0186] The information, signals, etc., described in this disclosure can also be represented using one of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be mentioned throughout the above description, can also be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or photons, or any combination thereof.

[0187] Furthermore, the terms used in this disclosure and those necessary for understanding this disclosure may be replaced with terms that have the same or similar meanings. For example, a signal may also be a message or signaling.

[0188] Furthermore, the information, parameters, etc., described in this disclosure can be represented using absolute values, relative values ​​with respect to a specific value, or other corresponding information. For example, wireless resources can also be indicated by an index.

[0189] The term "determine" as used in this disclosure sometimes encompasses a variety of operations. For example, "determine" or "obtain" may include actions such as making a judgment, calculating, deriving, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), confirming, etc. Furthermore, "determine" or "obtain" may include actions such as receiving (e.g., receiving information), sending (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory) as "judging" or "deciding," etc.

[0190] Unless otherwise expressly stated, the use of the word "based on" in this disclosure does not imply "based on only". In other words, the use of the word "based on" implies both "based on only" and "based on at least".

[0191] In this disclosure, the term "unit" in the structure of the above-mentioned devices may also be replaced with "circuit", "device", etc.

Claims

1. A rendering processing method based on medical image segmentation, comprising: Associate semantic feature vectors obtained from natural language task instructions Image feature vectors in the target image To generate joint feature vectors ,in, Indicates the coordinates of the image voxel; Based on the semantic feature vector with the highest confidence level The corresponding joint feature vector generates the initial value of the configuration parameter set, wherein the configuration parameter set includes one or more segmentation parameters; Based on the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values, a clarification question is initiated to the user to obtain the updated values ​​of each segmentation parameter in the configuration parameter set. The target image is segmented according to the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of blood vessels. Based on the natural language task instructions and user feedback on the clarification questions, key regions in the 3D voxel structure of the blood vessels are determined. A physically based volume rendering method is used to render the 3D voxel structure of the blood vessels, and the key regions are enhanced. Based on the blood vessel segmentation results, hemodynamic parameters are determined to render the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters, wherein the hemodynamic parameters include at least the streamline data and spatial coordinates of the blood flow streamline nodes.

2. The method according to claim 1, wherein determining the key regions in the three-dimensional voxel structure of the blood vessels based on the natural language task instructions and the user's feedback to the clarification question comprises: Based on the natural language task instructions and the user's feedback on the clarification questions, determine the semantic nodes of the user's key rendering requirements; The semantic nodes of the key requirements are associated with the three-dimensional vascular voxel structure using a graph attention network (GAT), wherein the attention coefficient between each three-dimensional vascular voxel node in the vascular voxel structure and the semantic nodes in its anatomical structure is defined in the graph attention network. As shown in the following formula: in and These represent the feature vectors of the 3D voxel node and the semantic node, respectively. For learnable weight matrix, For the attention parameter vector; and The key regions in the three-dimensional vascular voxel structure are determined based on the attention coefficient corresponding to each three-dimensional vascular voxel node in the structure.

3. The method according to claim 1 or 2, wherein performing segmentation processing on the target image based on the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a blood vessel segmentation result includes: The pre-segmentation model is performed on the target image shown based on the updated values ​​of each segmentation parameter in the determined set of configuration parameters to obtain the pre-segmentation result; as well as Using a first computational network, a complete segmentation result for the target image is obtained based on the pre-segmentation result shown, and using a second computational network, a straightened blood vessel model is obtained based on the complete segmentation result shown to obtain a simplified segmentation result.

4. The method according to claim 3, wherein The simplified segmentation result is represented as a three-dimensional point cloud. Extracting hemodynamic parameters based on the simplified segmentation results includes: The internal point cloud of the simplified segmentation result 3D point cloud is input into the first channel of the Dynamic Graph Convolutional Network (DGCNN), and the surface point cloud of the simplified segmentation result 3D point cloud is input into the second channel of DGCNN, wherein the first channel is independent of the second channel; The first channel and the second channel respectively encode the features of the input point cloud. Specifically, the surface point cloud in the second channel is aggregated to form a global feature tensor; the internal point cloud in the first channel retains the feature representation point by point. Feature fusion is performed on the feature encoding results of the first channel and the second channel, wherein the global features are spliced ​​with the point-by-point features of the internal point cloud through a broadcast mechanism to form a comprehensive feature representation that simultaneously includes local geometric details and global structural awareness.

5. The method according to claim 1 or 2, wherein determining hemodynamic parameters based on the vessel segmentation result, and rendering the three-dimensional voxel structure of the vessel according to the cardiac cycle based on the hemodynamic parameters, comprises: Based on the spatial coordinates of the blood flow streamline nodes, the blood flow streamline nodes are drawn on a three-dimensional voxel structure in their topological order. For each frame in the cardiac cycle, an independent data structure representing a set of lines is generated, wherein each data structure representing a set of lines is composed of the blood flow streamline nodes in the three-dimensional voxel structure.

6. The method according to claim 1, wherein The semantic feature vector is obtained through the Large Language Model Interface (API) based on the natural language task instructions. The method further includes: One or more of the following rendering results are compared with corresponding reference rendering results: a first rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel, a second rendering result obtained by enhancing the rendering of the key region, and a third rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters; and a rendering loss parameter is obtained. The large language model interface is trained based on the rendering loss parameters.

7. The method according to claim 3, further comprising: One or more of the following are compared with the corresponding reference rendering results: a first rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel, a second rendering result obtained by enhancing the rendering of the key region, and a third rendering result obtained by rendering the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters; to obtain the rendering loss parameters. as well as The first network and the second network are trained based on the rendering loss parameters.

8. A rendering processing system based on medical image segmentation, comprising: The data management module is configured to connect to the image database to retrieve the corresponding target image from the image database according to the input command; The interactive reasoning-based large language model module is configured to associate semantic feature vectors obtained from natural language task instructions. Image feature vectors in the target image To generate joint feature vectors Based on the semantic feature vector with the highest confidence level The corresponding joint feature vector generates the initial value of the configuration parameter set, and based on the consistency between the values ​​of each segmentation parameter in the configuration parameter set and the image feature reference values, a clarification question is raised to the user to obtain the updated values ​​of each segmentation parameter in the configuration parameter set. The set of configuration parameters, representing image voxel coordinates, includes one or more segmentation parameters. The image segmentation module is configured to perform segmentation processing on the target image based on the update values ​​of each segmentation parameter in the determined set of configuration parameters to obtain a blood vessel segmentation result, wherein the blood vessel segmentation result includes a three-dimensional voxel structure of blood vessels. The image rendering module is configured to determine key regions in the 3D voxel structure of the blood vessels based on the natural language task instructions and the user's clarification question, render the 3D voxel structure of the blood vessels using a physically based volume rendering method, and enhance the rendering of the key regions; and The parameter extraction module is configured to determine hemodynamic parameters based on the vessel segmentation results, wherein the hemodynamic parameters include at least streamline data and spatial coordinates of the blood flow streamline nodes. The image rendering module is further configured to render the three-dimensional voxel structure of the blood vessel according to the cardiac cycle based on the hemodynamic parameters determined by the parameter extraction module.

9. An electronic device, comprising: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the rendering processing method based on medical image segmentation as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the rendering processing method based on medical image segmentation as described in any one of claims 1 to 7.