Object sketching method and device and electronic equipment
By integrating material distribution and boundary information through a deep learning model with edge linear attention, the method enhances the precision and robustness of medical image segmentation, overcoming errors from image registration and improving clinical accuracy.
Patent Information
- Application Number
- CN202510414260.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When existing medical imaging technology handles organ boundaries, complex lesion morphology and poor image quality, the automatic outline method has problems with insufficient segmentation accuracy and robustness, especially in the errors introduced during CT and MRI image registration.
By combining the target substance distribution information and boundary information, a deep learning model is used to fusion of features, and an encoder-decoder structure and edge linear attention module are used to generate accurate sub-object images.
It significantly improves outline accuracy and robustness, can accurately segment organ boundaries in complex structures and low signal-to-noise ratios, reduces clinical workload and improves the reliability of treatment plans.
Smart Images

Figure CN120318266A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the medical field and the field of deep learning, and more particularly, to an object delineation method, apparatus, and electronic device. Background Art
[0002] Existing medical imaging technologies have been widely applied in clinical diagnosis and treatment. Especially in the field of tumor radiotherapy, accurately delineating human organs and lesion areas is of crucial significance for the formulation of treatment plans. Currently, manual delineation is still the mainstream method. However, this process is time-consuming and laborious, and highly dependent on the doctor's experience level, resulting in significant differences in the subjectivity and consistency of the delineation results. For this reason, the automatic delineation technology based on deep learning has gradually become a research hotspot, which can quickly process a large amount of imaging data, thereby improving efficiency and accuracy. Nevertheless, the existing technologies still show great limitations when facing situations such as blurred organ boundaries, complex lesion morphologies, and poor image quality, and cannot fully meet the clinical needs. Especially for the automatic segmentation technology based on multi-modal medical imaging data such as CT and MRI, it is usually necessary to first perform image registration from MRI to CT, and the errors introduced in the registration process significantly affect the segmentation accuracy. However, due to the large errors introduced during the image registration of CT and MRI images, the existing methods still lack sufficient segmentation accuracy and robustness when dealing with blurred organ boundaries, complex lesion morphologies, and imaging quality affected by noise. Summary of the Invention
[0003] In view of this, the present disclosure provides an object delineation method and an electronic device.
[0004] One aspect of the present disclosure provides an object delineation method, including: obtaining target imaging data, where the target imaging data at least represents an image generated based on a target ray for a target object, and the target object includes a plurality of sub-objects; generating first information and second information based on the target imaging data, where the first information represents the distribution information of a target substance in the target object, and the second information represents the boundary information of each sub-object in the target object; and delineating the images of at least one sub-object according to the first information and the second information.
[0005] According to an embodiment of the present disclosure, the distribution information of the target substance includes the distribution state information of the target substance and / or the density information of the target substance.
[0006] According to an embodiment of the present disclosure, the target substance includes at least one of carbon element and oxygen element.
[0007] According to an embodiment of the present disclosure, generating first information includes: determining at least one test characteristic parameter of the target object based on target image data; determining at least one test material type of the target object based on the test characteristic parameters; determining material information of at least one target substance in the target object based on material distribution characteristics corresponding to the test material type and each test characteristic parameter of the target object; generating first information based on the material information corresponding to at least one substance.
[0008] According to an embodiment of the present disclosure, an image of at least one sub-object is outlined based on the first information and the second information, including: fusing the first information and the second information to obtain target information; inputting the target information into a target model to generate an image of at least one sub-object of the target.
[0009] According to an embodiment of the present disclosure, the target model includes an up-sampling module, a down-sampling module and an edge linear attention module; the up-sampling module is used for deep features in the target information; the down-sampling module is used to extract semantic information in the deep features and generate an image of the sub-object; the edge linear attention module is used to generate boundary information based on the feature maps generated by the up-sampling module and the down-sampling module and transmit it to the up-sampling module, so that the up-sampling module generates an image of the sub-object based on the boundary information.
[0010] According to an embodiment of the present disclosure, the training process of the target model includes: acquiring sample first information and sample second information, the sample first information characterizing the sample material distribution state information in the sample object, and the sample second information characterizing the sample boundary information of each sub-object in the target object; training the target model according to the sample first information, the sample second information and the target loss function.
[0011] According to an embodiment of the present disclosure, the target loss function includes at least a first loss function, and the first loss function is used to calculate the error between the image boundary of the sub-object and the second information of the sample in the output of the model.
[0012] Another aspect of the present disclosure provides an object delineation device, including: a first acquisition module, used to acquire target image data, the target image data at least represents an image generated based on target rays for a target object, the target object including multiple sub-objects; a first generation module, used to generate first information and second information based on the target image data, the first information represents target material distribution information in the target object, and the second information represents boundary information of each sub-object in the target object; and a delineation module, used to delineate an image of at least one sub-object based on the first information and the second information.
[0013] Another aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the object delineation method of any one of the foregoing embodiments.
[0014] Another aspect of the present disclosure provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the object delineation method according to any one of the foregoing embodiments.
[0015] Another aspect of the present disclosure provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the operations of the object delineation method of any one of the foregoing embodiments are implemented.
[0016] According to the embodiments of the present disclosure, the object delineation method provided by the present disclosure has at least one of the following beneficial effects: combining the distribution state information of the target substance and the sub-object boundary information, automatically delineating the sub-objects in the target image on the basis of integrating different types of features, and significantly improving the accuracy and robustness of delineation. Among them, the target substance distribution information can reflect the physical or chemical composition of different tissues, which helps to distinguish tissue regions with similar shapes but different compositions; the boundary information enhances the clarity of the sub-object contour, and still has strong recognition ability especially in the case of blurred tissue boundaries or low image signal-to-noise ratio. By fusing these two types of information and inputting them into the target model, accurate segmentation of complex structures can be achieved, effectively overcoming the problems of misjudgment and missed detection caused by single features in the existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0018] Figure 1 Schematically shows a flowchart of the object delineation method according to an embodiment of the present disclosure;
[0019] Figure 2 Schematically shows a flowchart of generating the first information in the object delineation method according to an embodiment of the present disclosure;
[0020] Figure 3 Schematically shows a flowchart of delineating sub-objects in the object delineation method according to an embodiment of the present disclosure;
[0021] Figure 4 Schematically shows a structural diagram of the target model in the object delineation method according to an embodiment of the present disclosure;
[0022] Figure 5Schematically shows a training flowchart of a target model in an object delineation method according to an embodiment of the present disclosure;
[0023] Figure 6 Schematically shows a block diagram of an object delineation apparatus according to an embodiment of the present disclosure; and
[0024] Figure 7 Schematically shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure. Detailed implementation manners
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0026] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0029] In the embodiments of the present disclosure, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard user personal information security, network security and national security.
[0030] Embodiments of the present disclosure provide an object delineation method, including: obtaining target image data, where the target image data at least represents an image generated based on target rays for a target object, and the target object includes multiple sub-objects; generating first information and second information based on the target image data, where the first information represents the distribution information of target substances in the target object, and the second information represents the boundary information of each sub-object in the target object; delineating the images of at least one sub-object according to the first information and the second information.
[0031] Figure 1 Schematically shows a flowchart of the object delineation method according to an embodiment of the present disclosure.
[0032] As Figure 1 shown, the object delineation method may at least include operations S110 to S130.
[0033] In operation S110, obtain target image data, where the target image data at least represents an image generated based on target rays for a target object, and the target object includes multiple sub-objects.
[0034] The target image data can be image data obtained by a medical imaging device (such as a CT scanner, an MRI device, etc.). The target rays include, but are not limited to, X-rays, γ-rays, or other forms of electromagnetic waves, and the rays can be used to penetrate the internal structure of the target object. The target object is usually a human body or other object with a certain organizational structure.
[0035] The target object includes multiple sub-objects, which can be different organs inside the human body (such as the heart, liver, lungs, etc.), or specific object parts that need to be analyzed. Each sub-object has different physical and chemical properties, and their shapes and positions usually show different characteristics in the image.
[0036] Specifically, the target image data can be generated by obtaining ray data at multiple projection angles and performing reconstruction to generate three-dimensional image data, or further analyzed through an image composed of two-dimensional image layers.
[0037] For example, the target image data can be image data from a CT scan, where each pixel value represents density information at a specific position, or image data from an MRI scan, and each pixel value represents the water molecule density information of different tissues.
[0038] In operation S120, generate first information and second information based on the target image data, where the first information represents the distribution information of target substances in the target object, and the second information represents the boundary information of each sub-object in the target object.
[0039] Specifically, the first information is generated by performing elemental decomposition or other substance analysis methods on the target image data, and the analysis methods can extract the distribution of different tissues or substances. This first information can reflect the density and composition distribution of various substances in the target object, and even include the physical or chemical element information of different tissues or substances. For example, through elemental decomposition technology, information related to different substance components can be separated from CT image data, so as to more accurately obtain the substance distribution characteristics of each organ.
[0040] The second information can be extracted by a boundary detection algorithm, which can clearly define the outer contours of each sub-object in the target object. Especially in areas with blurred boundaries or poor image quality, the boundary detection technology can still provide effective boundary information. Common boundary detection methods include the Sobel operator, Canny edge detection, etc., which can detect the edges of the target object based on image gradients or pixel value changes.
[0041] In operation S130, according to the first information and the second information, the image of at least one sub-object is outlined. Specifically, based on the generated first information and second information, an image segmentation technology is adopted to extract the image area of the target sub-object by fusing the substance distribution state information and the boundary information.
[0042] According to the embodiments of the present disclosure, by combining the distribution state information of the target substance and the sub-object boundary information, and automatically outlining the sub-objects in the target image on the basis of fusing different types of features, the accuracy and robustness of the outlining can be significantly improved. Among them, the target substance distribution information can reflect the physical or chemical composition of different tissues, which helps to distinguish tissue regions with similar shapes but different compositions; the boundary information enhances the clarity of the sub-object contour, and still has strong recognition ability especially in cases where the tissue boundary is blurred or the image signal-to-noise ratio is low. By fusing these two types of information and inputting them into the target model, accurate segmentation of complex structures can be achieved, effectively overcoming the problems of misjudgment and missed detection caused by single features in existing methods.
[0043] According to the embodiments of the present disclosure, the target substance distribution information includes the distribution state information of the target substance, and / or the density information of the target substance.
[0044] The target substance can be a physical or chemical component with representative significance in the target object, such as common human tissue composition elements such as carbon element, oxygen element, hydrogen element, calcium element, etc., or a composite substance that can be identified by ray response characteristics, such as soft tissue, bone tissue, adipose tissue, lesion tissue, etc.
[0045] The distribution state information can be used to describe the spatial position relationship of the target substance in the target image, that is, the regional existence situation of the substance inside the target object. Its form can be a grayscale image, a pseudo-color image, a label image, etc. of a two-dimensional or three-dimensional image, and is usually obtained through an element decomposition algorithm or an image reconstruction technique based on physical modeling. This type of information can reflect the compositional boundaries between different tissues and helps to enhance the distinguishability between sub-objects.
[0046] The density information is a numerical representation of the physical concentration, attenuation coefficient, or equivalent density of the target substance in a unit space, and is usually obtained by calculating the CT value (Hounsfield Units) or the MRI signal intensity. This information can further refine the material characteristics of the tissue and helps to improve the discriminative ability of the model for tissue types.
[0047] According to an embodiment of the present disclosure, in addition to the distribution state information and the density information, the target substance distribution information may further include relative proportion information of the target substance, spectral response characteristic information, ionization radiation absorption coefficient information, or magnetic resonance relaxation characteristic information, etc. The relative proportion information is used to describe the proportion of a specific substance in a local area and can be used for tissue heterogeneity analysis; the spectral response characteristic information can reflect the attenuation differences of the substance under multi-energy conditions; the ionization radiation absorption coefficient information is used to quantify the absorption ability of the substance for X-rays of different energies; the magnetic resonance relaxation characteristic information can be used to distinguish the T1 and T2 relaxation time characteristics of different types of soft tissues. All of these information can be used as fusion features to be input into a neural network model to improve the overall performance and accuracy of the automatic delineation algorithm.
[0048] Figure 2 Schematically shows a flowchart of generating the first information in the object delineation method according to an embodiment of the present disclosure.
[0049] As Figure 2 shown, on the basis of the foregoing embodiment, S120 may include operations S210 to S240.
[0050] In operation S210, at least one test feature parameter of the target object is determined according to the target image data.
[0051] Specifically, the target image data can be image data obtained by medical imaging devices (such as CT, MRI, or dual-energy CT). In this process, first, image acquisition is performed on the target object, and through different energy rays or scanning modes, image information reflecting different physical or chemical characteristics of the target object is obtained. The test characteristic parameters include, but are not limited to, the HU value, pixel gray value, image density, relative electron density, effective atomic number (Zeff), etc. of the target object. These parameters can reflect the physical properties of the target object under different imaging modes. For example, the HU value is usually used in CT images to represent the attenuation degree of each part in the image and can be used to infer the density of substances; while the relative electron density and effective atomic number Zeff can provide the relative properties of different substances in the image, which helps in further classification and substance identification. For example, in dual-energy CT images, the HU values under two different energy spectra obtained by using two different energy ray sources can be used to calculate the relative electron density and effective atomic number of the target object, providing necessary test characteristic parameters for the subsequent analysis of the target object.
[0052] In operation S220, at least one test material type of the target object is determined according to the test characteristic parameters.
[0053] Specifically, based on the test characteristic parameters extracted from the target image data, such as relative electron density and effective atomic number Zeff, the target object can be classified by material. The test material type refers to dividing at least part of the target object into different categories according to the physical properties, elemental composition of the target object, and the manifestation of these properties in the image data. For example, by comparing the effective atomic number Zeff of the target object with a predetermined threshold, different tissue types can be distinguished. When Zeff is greater than a certain preset threshold, the corresponding part of the target object can be determined to be bone tissue; if Zeff is less than this threshold, it can be judged as soft tissue. This process helps to establish an accurate classification model for different types of tissues or substances, ensuring the efficient analysis of image data. For example, in dual-energy CT scans, bone tissue usually contains a large amount of calcium elements, resulting in a relatively high effective atomic number, while soft tissue is mainly composed of carbon, oxygen, and hydrogen elements, with a lower Zeff. Based on this characteristic, the test characteristic parameters can be used to quickly and accurately identify bone tissue and soft tissue.
[0054] In operation S230, according to the substance distribution characteristics corresponding to the test material type and each test characteristic parameter of the target object, determine the substance distribution state information of at least one target substance in the target object. Specifically, after determining the material types of the target object, the substance distribution characteristics of this type of material can be further analyzed. For example, the elemental composition of bone tissue includes elements such as calcium, phosphorus, carbon, oxygen, and nitrogen, and its distribution characteristics in different regions have significant regularity. By analyzing the test characteristic parameters of the target object and combining the known material element distribution laws, the specific physical or chemical composition and its distribution of each substance in the target object can be inferred.
[0055] The substance distribution state information refers to the spatial distribution characteristics of the target substance within the target object, and this information can reflect the existence probability, relative concentration, or mass fraction of a specific substance at different positions. Generally, this information can be constructed as a two-dimensional image matrix or a three-dimensional voxel matrix, where the value of each pixel or voxel corresponds to the relative content or density level of the substance at that location, thereby realizing the visual expression of the substance in space.
[0056] For example, after determining that the test material type of the target object includes bone tissue, based on the elemental composition of bone tissue (such as high concentrations of calcium and phosphorus, and moderate concentrations of carbon and oxygen) and the test characteristic parameters, the spatial distribution of calcium or carbon elements in this bone tissue can be further inferred, and then the distribution image of this element can be constructed. This image constitutes the substance distribution state information of the target substance and can be used for subsequent tasks such as image processing, automatic contouring, or feature extraction.
[0057] In operation S240, generate the first information according to the substance information corresponding to at least one substance. Specifically, based on the distribution state information, density information, and elemental composition of the target substance, the first information can be generated, and this information can be further used as the algorithm input for image segmentation or organ contouring. The first information can include data such as the spatial distribution characteristics, elemental composition, and physical density of the target substance, providing more accurate basic data for subsequent automated processing or diagnosis. For example, in the analysis of bone tissue, based on the distribution and physical properties of the target substance (such as calcium element, phosphorus element, etc.), the generated first information can be a set of three-dimensional matrices, and each element in the matrix represents the relative density or mass distribution of calcium element in a certain region of the bone tissue. In this way, the first information provides a more detailed and accurate physical basis for subsequent contouring and segmentation.
[0058] Figure 3 Schematically shows a flowchart of delineating a sub-object in an object delineation method according to an embodiment of the present disclosure.
[0059] As Figure 3 shown, on the basis of the foregoing embodiments, S130 may include operations S310 to S350.
[0060] In operation S310, the first information and the second information are data-fused to obtain target information.
[0061] Specifically, the first information represents the distribution characteristics of the target substance in the target object, and the second information represents the boundary contour information of each sub-object, and the two have different semantic feature dimensions. In order to make full use of the complementarity of the two in material composition and morphological boundaries, the first information and the second information are fused in the spatial domain or feature domain through feature fusion operations to form target information containing rich structural and compositional features. The fusion operation may include pixel-level fusion, feature-level splicing, weighted fusion guided by attention mechanisms, etc., to enhance the expression ability of subsequent models for complex structures. The first information and the second information can be spliced into a multi-dimensional input tensor according to the channel dimension, or the two can be input into different encoding branches respectively, fused in the middle layer, and the fused target information is uniformly output, which will simultaneously reflect the boundary morphology and material distribution status of the sub-objects.
[0062] In operation S320, the target information is input into the target model to generate an image of at least one sub-object of the target. Specifically, the target model can be a trained deep learning neural network model, which can receive the fused target information, automatically extract multi-level information such as semantic features, spatial features and boundary features, and generate a prediction result for outlining the target sub-object. The generated sub-object image can be a label map, a mask map or a pseudo-color map, which identifies the precise area corresponding to the sub-object in the target image. For example, the target model can adopt a network architecture with an encoder-decoder structure, extract multi-scale features in the encoding stage, and perform refined prediction in combination with shallow boundary information in the decoding stage, so as to output a sub-object image result with higher accuracy, that is, it can outline organs with higher accuracy.
[0063] Figure 4 The structure diagram of the target model in the object delineation method according to the embodiment of the present disclosure is schematically shown.
[0064] like Figure 4 As shown, the target model includes an up-sampling module, a down-sampling module and an edge linear attention module. The down-sampling module is used for deep features in the target information. The up-sampling module is used to extract semantic information from the deep features and generate images of sub-objects; the edge linear attention module is used to generate boundary information based on the feature maps generated by the up-sampling module and the down-sampling module and transmit it to the up-sampling module, so that the up-sampling module generates images of sub-objects based on the boundary information.
[0065] like Figure 4As shown, the overall target model adopts an encoder-decoder structure. On the left is the encoder part. The image data is first input into the Edge Encoder to extract edge-related features, and then further transformed into a compressed feature representation (deep features) by the Mamba Encoder, constituting the described downsampling module. In the downsampling module, feature maps (Encoder Feature Map) are extracted layer by layer, and through the multi-level downsampling operations represented by arrows, the spatial resolution is continuously compressed and deep features are extracted.
[0066] After downsampling is completed, the feature maps are input into the decoder path for gradual restoration, corresponding to the upsampling module in the figure. This module restores the low-resolution feature maps to high-resolution features through the upsampling operations shown by the arrows. Each layer of downsampling generates corresponding outputs (Decoder Feature Map), which are transmitted layer by layer until a sub-object image with a complete structure and accurate distribution is generated. The upsampling module not only performs the spatial restoration task but also is responsible for fusing semantic and boundary information to generate a sub-object image with a complete structure and accurate distribution.
[0067] Between the upsampling and downsampling paths, multiple Edge Linear Attention modules are set up to establish boundary enhancement connections between different scale layers, represented by cubes in the figure. This module simultaneously receives the feature maps from the downsampling module (encoder path) and the upsampling module (decoder path) and generates boundary information (Edge Feature Map) accordingly.
[0068] As Figure 4 shown in the upper right corner, first, the feature map from the encoder path (the upper left yellow feature block in the figure) is input into a Softmax module to generate an attention distribution for the encoder features; at the same time, the feature map from the decoder path is also input into another independent Softmax module to generate an attention distribution on the decoder side.
[0069] Subsequently, the feature map from the decoder path and the attention weights generated by the encoder path perform the first matrix multiplication operation (Matrix Multiply, the first X module in the figure) to form a preliminary fusion feature; immediately afterwards, this preliminary fusion feature will perform a second matrix multiplication operation (Matrix Multiply, the second X module in the figure) with the feature map from the encoder path to further model the deep correlation between contexts and generate an intermediate edge feature map containing edge representations.
[0070] The intermediate edge feature map enters the feature fusion module (Concat, module C in the figure), is concatenated and fused with the feature map of the decoder path, and finally forms an edge-aware feature map with enhanced structure (the blue-gray fusion feature block in the lower right of the figure), which is used as one of the inputs for the subsequent upsampling module to generate a sub-object image with clear boundaries.
[0071] Figure 5 Schematically shows a training flowchart of a target model in an object delineation method according to an embodiment of the present disclosure.
[0072] As Figure 5 shown, based on the foregoing embodiments, the training process of the target model may include operations S510 to S520.
[0073] In operation S510, sample first information and sample second information are obtained. The sample first information represents the sample substance distribution state information in the sample object, and the sample second information represents the sample boundary information of each sub-object in the target object.
[0074] Specifically, the sample first information is the substance distribution state information extracted based on the physical characteristics of the sample object after medical imaging scanning (such as CT, MRI, or dual-energy CT) of the sample object. The sample substance distribution state information may include the spatial distribution characteristics of specific substances (such as calcium, carbon, adipose tissue, etc.) in the sample object, usually represented in the form of a two-dimensional or three-dimensional matrix, where each element of the matrix corresponds to the relative concentration or mass fraction of the specific substance in the sample object. The sample second information is the boundary contour information extracted from the image data of the target object through an edge detection algorithm or a pre-annotation method, representing the spatial boundary positions of each sub-object (such as organs, bones, lesions, etc.) in the target object, and can be presented in the form of a mask image or an edge label image. For example, for a bone sample object, the sample first information may be a three-dimensional distribution matrix of calcium in the bone, reflecting the concentration change of calcium in different regions; the sample second information may be a mask image of the bone edge, marking the boundary between the bone and the surrounding soft tissue, providing boundary supervision information for subsequent model training.
[0075] At operation S520, the target model is trained according to the first sample information, the second sample information, and the target loss function. The target model can be a deep learning network with an encoder-decoder structure. The training process takes the first sample information and the second sample information as inputs, fuses the features of the two types of information to optimize the model parameters. During the training process, first, data preprocessing is performed on the first sample information and the second sample information. For example, in the way of feature splicing or weighted fusion, the substance distribution state information and the boundary information are integrated into comprehensive target input features. Subsequently, the target input features are input into the target model, and the model generates prediction results through forward propagation. The target loss function is used to measure the difference between the model prediction results and the true annotations, and updates the model parameters through the backpropagation algorithm, so that the model gradually converges to the optimal state. The training process can adopt batch gradient descent or its variants (such as the Adam optimizer) to accelerate convergence, and at the same time prevent overfitting by setting appropriate learning rates and regularization strategies (such as Dropout). For example, in a medical image segmentation task, the first sample information (calcium element distribution matrix) and the second sample information (bone edge mask image) are input into the target model. The model learns the association between the substance distribution and the boundary features through multiple rounds of iteration, and finally generates a sub-object image that can accurately delineate the bone region.
[0076] According to an embodiment of the present disclosure, the target loss function at least includes a first loss function, and the first loss function is used to calculate the error between the content output by the edge linear attention module in the output of the target model and the second sample information.
[0077] Specifically, the edge linear attention module is mainly used to capture the boundary features of each sub-object in the target object, and output an edge feature map to characterize its spatial boundary structure. The role of the first loss function is to supervise the edge feature map output by this module and the true edge label map in the second sample information, so as to optimize the response ability of the attention mechanism to the edge region. Through this supervision mechanism, the model can explicitly learn the attention distribution related to the boundary during the training process, making the generated feature map closer to the actual edge region, thereby improving the structural integrity and edge alignment accuracy of the entire segmentation result.
[0078] For example, let the edge feature map output by the edge linear attention module be , where represents the edge response intensity at the i-th pixel; the edge label map in the second sample information is , where represents whether the i-th pixel is an edge pixel, and N is the total number of pixels.
[0079] Then the first loss function can adopt the binary cross entropy (BCE) loss function Ledge , and its expression can be as follows
[0080]
[0081] In some embodiments, the loss function can also be extended to an edge-weighted cross-entropy loss, which assigns higher weights to regions near the true boundary pixels to enhance the model's sensitivity to blurred boundaries or fine-grained structures. This loss function can be optimized in parallel with the main segmentation loss of the target model to achieve an end-to-end boundary-guided training framework.
[0082] It should be noted that the above loss function is only an example. Those skilled in the art can design and select other loss functions that conform to this solution according to the solution provided by the present disclosure, which should be included within the scope protected by the present disclosure.
[0083] According to the embodiments of the present disclosure, by combining the element decomposition technology and boundary detection technology of CT images and using a deep learning model for multi-feature fusion, the accuracy and robustness of organ delineation are significantly improved. Compared with traditional methods, the present invention uses the element decomposition results of CT images to reflect the boundary information of organs, avoiding the errors caused by multi-modal image registration. At the same time, accurate boundary features are obtained through an edge extraction algorithm and input into the network together with the original CT gray information to achieve "end-to-end" feature extraction and fusion; the Mamba encoder is used to efficiently capture multi-level image information, and the BoundaryLinearAttention module is innovatively introduced to achieve efficient fusion of multi-scale features through a linear attention mechanism, further enhancing the ability to represent boundary details. In addition, the introduction of the boundary auxiliary loss function strengthens the model's recognition ability for boundary regions, making the segmentation results show higher accuracy and adaptability in complex structures and low-contrast regions, thus providing more reliable and efficient technical support for the formulation of tumor radiotherapy plans, significantly reducing the clinical workload and improving the treatment effect.
[0084] Figure 6 A block diagram of an object delineation device according to an embodiment of the present disclosure is schematically shown.
[0085] As Figure 6 shown, the object delineation device 600 may include a first acquisition module 610, a first generation module 620, and a delineation module 630.
[0086] The first acquisition module 610 is configured to acquire target image data, and the target image data at least represents an image generated based on a target ray for a target object, and the target object includes multiple sub-objects. In some embodiments, the first acquisition module 610 may be configured to perform the operation S110 in the above object delineation method, which will not be elaborated here.
[0087] The first generation module 620 is configured to generate first information and second information based on target image data, where the first information represents the distribution information of the target substance in the target object, and the second information represents the boundary information of each sub-object in the target object. In some embodiments, the first generation module 620 may be configured to perform the operation S120 in the above object delineation method, which will not be elaborated herein.
[0088] The delineation module 620 is configured to delineate the images of at least one sub-object according to the first information and the second information. In some embodiments, the delineation module 620 may be configured to perform the operation S130 in the above object delineation method, which will not be elaborated herein.
[0089] According to embodiments of the present disclosure, any plurality of modules, sub-modules, units, and sub-units, or at least part of the functions of any of them may be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0090] For example, any combination of the first acquisition module 610, the first generation module 620, and the drawing module 630 may be implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units may be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units may be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first acquisition module 610, the first generation module 620, and the drawing module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation modes of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the first acquisition module 610, the first generation module 620, and the drawing module 630 may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.
[0091] It should be noted that the data processing system part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. For a specific description of the data processing system part, reference may be made to the data processing method part, and details are not described herein again.
[0092] Figure 7 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 7 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0093] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0094] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present disclosure by executing programs in the ROM 702 and / or the RAM 703. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing programs stored in the one or more memories.
[0095] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input part 706 including a keyboard, a mouse, etc.; an output part 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 708 including a hard disk, etc.; and a communication part 709 including a network interface card such as a LAN card, a modem, etc. The communication part 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage part 708 as needed.
[0096] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-mentioned functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0097] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.
[0098] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0099] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.
[0100] Embodiments of the present disclosure further include a computer program product, which includes a computer program that contains program code for executing the methods provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the control methods provided by the embodiments of the present disclosure.
[0101] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0102] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 709, and / or installed from the removable medium 711. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above. According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0104] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. An object delineation method, comprising: Obtaining target image data, where the target image data at least represents an image generated based on a target ray for a target object, and the target object includes a plurality of sub-objects; Generating first information and second information based on the target image data, where the first information represents the distribution information of a target substance in the target object, and the second information represents the boundary information of each of the sub-objects in the target object; Delineating the images of at least one of the sub-objects according to the first information and the second information.
2. The method according to claim 1, wherein The target substance distribution information includes the distribution state information of the target substance, and / or the density information of the target substance.
3. The method according to claim 2, wherein, The target substance includes at least one of carbon element, oxygen element, hydrogen element, and calcium element.
4. The method according to claim 1, wherein, The generating of the first information includes: Determining at least one test characteristic parameter of the target object according to the target image data; Determining at least one test material type of the target object according to the test characteristic parameter; Determining the substance information of at least one of the target substances in the target object according to the substance distribution characteristics corresponding to the test material type and each test characteristic parameter of the target object; Generating the first information according to the substance information corresponding to at least one of the substances.
5. The method according to claim 1, wherein, Delineating the images of at least one of the sub-objects according to the first information and the second information includes: Performing data fusion on the first information and the second information to obtain target information; Inputting the target information into a target model to generate the images of at least one of the sub-objects of the target.
6. The method according to claim 5, wherein, The target model includes an upsampling module, a downsampling module, and an edge linear attention module; The downsampling module is used for deep features in the target information; The upsampling module is used to extract semantic information from the deep features to generate the images of the sub-objects; The edge linear attention module is used to generate boundary information according to the feature maps generated by the upsampling module and the downsampling module and transmit it to the upsampling module, so that the upsampling module generates the images of the sub-objects according to the boundary information.
7. The method according to claim 5, wherein The training process of the target model includes: Obtaining sample first information and sample second information, where the sample first information represents the sample substance distribution state information in the sample object, and the sample second information represents the sample boundary information of each sub-object in the target object; Training the target model according to the sample first information, the sample second information, and a target loss function.
8. The method according to claim 7, wherein, The target loss function at least includes a first loss function, and the first loss function is used to calculate the error between the image boundary of the sub-object in the output of the model and the sample second information.
9. An object delineation device, comprising: A first acquisition module, configured to acquire target image data, where the target image data at least represents an image generated based on a target ray for a target object, and the target object includes a plurality of sub-objects; A first generation module, configured to generate first information and second information based on the target image data, where the first information represents the distribution information of the target substance in the target object, and the second information represents the boundary information of each of the sub-objects in the target object; And A delineation module, configured to delineate an image of at least one of the sub-objects according to the first information and the second information.
10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Methods and system for linking geometry obtained from images
CA2938077A1
Image segmentation method, model and equipment based on modal specificity and storage medium
CN115187618A
Target segmentation method and device based on boundary attention and distance transformation
CN115546239A
Multi-modal constitution component marking and analyzing method based on artificial intelligence technology
CN116228624A