Image processing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610637726.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]本发明实施例提供了一种图像处理方法、装置、电子设备及存储介质,以至少解决图像处理准确性较低的技术问题
[0020]In this embodiment of the invention, medical images of a target organ are received; the medical images are input into a target model; and the target model outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This solves the technical problem of low image processing accuracy in the prior art.
Smart Images

Figure CN122657545A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more specifically, to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Pancreatic ductal adenocarcinoma is a type of malignant pancreatic tumor, and it often presents as insidious on subclinical CT scans (computed tomography) and lacks specific symptoms. Among related technologies, manual screening methods struggle to detect subtle textural heterogeneity on non-contrast CT, while artificial intelligence models lack robustness in complex real-world scenarios. Therefore, the accuracy of CT-based early detection of pancreatic ductal adenocarcinoma is relatively low.
[0003] There is currently no effective solution to the problem of low image processing accuracy mentioned above. Summary of the Invention
[0004] The present invention provides an image processing method, apparatus, electronic device, and storage medium to at least solve the technical problem of low image processing accuracy.
[0005] To achieve the above objectives, according to one aspect of this application, an image processing method is provided, the method comprising: receiving a medical image of a target organ; inputting the medical image into a target model; and outputting multiple detection results using the target model, wherein a first stage of the target model is used to locate the target region where the target organ is located based on the medical image, and a second stage of the target model is used to output multiple detection results based on the target region image, the target region image being an image of the target region in the medical image, and the multiple detection results including lesion segmentation results, lesion subtype results, and surrounding structure segmentation results, the surrounding structure segmentation results being segmentation results of tissue structures within a preset range from the target organ.
[0006] Furthermore, the medical images are input into the target model, and the target model outputs multiple detection results, including: inputting the medical images into the first-stage sub-model of the target model, and using the first-stage sub-model to output a three-dimensional mask image based on the medical images; cropping the target region image from the three-dimensional mask image in the target model, wherein the target region image is used to represent the image information of the target region; inputting the target region image into the second-stage sub-model of the target model, and using the second-stage sub-model to simultaneously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, and determining the detection results based on the lesion segmentation results, lesion subtype results, and surrounding structure segmentation results.
[0007] Furthermore, the second-stage sub-model synchronously outputs lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, including: using the segmentation path network in the second-stage sub-model to extract lesion segmentation results and surrounding structure segmentation results based on the target region image, wherein the segmentation path network includes downsampling paths and upsampling paths, with the downsampling paths and upsampling paths being skip connections; extracting spatial features with location encoding corresponding to the feature maps from multiple feature maps of the upsampling paths to obtain multiple spatial features; and using the memory path network in the second-stage sub-model to generate lesion subtype results based on multiple spatial features and a preset memory matrix.
[0008] Furthermore, the generation of lesion subtype results based on multiple spatial features and a preset memory matrix using the memory path network in the second-stage sub-model includes: using multiple attention interaction layers in the memory path network to encode multiple spatial features and the preset memory matrix to obtain discriminative feature vectors, where the discriminative feature vectors are used to encode the texture pattern and spatial structure information of the lesions, and the attention interaction layers have a one-to-one correspondence with the spatial features. The first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vectors; the discriminative feature vectors are input into the classification layer in the memory path network, and the classification layer is used to output the lesion subtype results.
[0009] Furthermore, the training steps of the target model include: obtaining a first training dataset and training a first preset neural network model using the first training dataset to obtain a first-stage sub-model; obtaining a second training dataset and training a second preset neural network model using the second training dataset to obtain a second-stage sub-model; and determining the target model based on the first-stage sub-model and the second-stage sub-model.
[0010] Furthermore, the first training samples in the first training dataset include a union mask of healthy tissue of the target organ and potential lesions. Training the first preset neural network model using the first training dataset to obtain the first stage sub-model includes: inputting the first training samples into the first preset neural network model; updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask; and determining the first stage sub-model based on the updated first preset neural network model.
[0011] Furthermore, the second training samples in the second training dataset include lesion segmentation annotation information, lesion subtype annotation information, and surrounding structure segmentation annotation information corresponding to the target organ. The second training dataset is used to train the second preset neural network model to obtain the second-stage sub-model, which includes: inputting the second training samples into the second preset neural network model and simultaneously training multiple tasks, including lesion segmentation, lesion subtype classification, and surrounding structure segmentation; determining the sum of errors between the prediction results of the multiple tasks and the corresponding annotation information, and updating the parameters of the second preset neural network model based on the sum of errors; and determining the second-stage sub-model based on the updated second preset neural network model.
[0012] Furthermore, determining the target model based on the first-stage sub-model and the second-stage sub-model includes: extracting target false positive samples from the target log, wherein the target log is a log of medical images of the target organ being detected according to a preset model, and the target false positive samples are samples in which the target organ is normal but the surrounding structures of the target organ are abnormal; and performing incremental learning on the first-stage sub-model and the second-stage sub-model based on the target false positive samples to obtain the target model.
[0013] To achieve the above objectives, according to another aspect of this application, a computer-aided diagnostic method based on medical images is provided. The method includes: receiving medical images of a target organ, wherein the medical images are used to detect whether the target organ has any abnormalities; inputting the medical images into a target model, and using the target model to output multiple detection results, wherein a first stage of the target model is used to locate the target region where the target organ is located based on the medical images, and a second stage of the target model is used to output multiple detection results based on the target region image, wherein the target region image is an image of the target region in the medical images, and the multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results, wherein the surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0014] To achieve the above objectives, according to another aspect of this application, a computer-aided diagnostic method for pancreatic cancer based on medical images is provided. The method includes: receiving medical images of the pancreas, wherein the medical images are used to detect whether any abnormalities exist in the pancreas; inputting the medical images into a target model, and using the target model to output multiple detection results, wherein a first stage of the target model is used to locate the target region where the pancreas is located based on the medical images, and a second stage of the target model is used to output multiple detection results based on pancreatic region images, wherein the pancreatic region images are images of the target region in the medical images, and the multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results, wherein the surrounding structure segmentation results are segmentation results of tissue structures within a preset range from the pancreas.
[0015] To achieve the above objectives, according to another aspect of this application, a computer-aided diagnostic system based on medical images is provided, comprising a client and a server, wherein: the client is used to upload medical images of a target organ to the server, wherein the medical images are used to detect whether there are abnormalities in the target organ; the server is used to input the medical images into a target model and output multiple detection results using the target model, wherein the first stage of the target model is used to locate the target region where the target organ is located based on the medical images, and the second stage of the target model is used to output multiple detection results based on the target region image, wherein the target region image is the image of the target region in the medical images, and the multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results, wherein the surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0016] To achieve the above objectives, according to another aspect of this application, an image processing apparatus is provided, comprising: a data receiving unit for receiving medical images of a target organ; and a result detection unit for inputting the medical images into a target model and outputting multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images, and the second stage of the target model is used to output multiple detection results based on the target region image. The target region image is an image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are segmentation results of tissue structures within a preset range from the target organ.
[0017] According to another aspect of this application, a computer-readable storage medium is provided, which includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform any method.
[0018] According to another aspect of this application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for performing any one of the methods.
[0019] According to another aspect of this application, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of any of the above methods.
[0020] In this embodiment of the invention, medical images of a target organ are received; the medical images are input into a target model; and the target model outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This solves the technical problem of low image processing accuracy in the prior art.
[0021] By using a target model to locate the target region of the target organ based on medical images, the search space is compressed in the first stage, avoiding false detections in non-target regions and significantly reducing errors caused by detection confusion. Then, in the second stage, lesion segmentation results, lesion subtype results, and surrounding structure segmentation results are obtained based on the target region image. By adding the surrounding structure segmentation results to the detection results, the target model can pay more attention to clues in the tissues surrounding the target organ, improving the identification ability of surrounding tissues and avoiding detection errors caused by abnormalities in surrounding tissues. As a result, the accuracy of the target model in detecting medical images is significantly improved, thus enhancing the accuracy of image processing. Attached Figure Description
[0022] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0023] Figure 1 A hardware structure block diagram of a computer terminal for implementing an image processing method is shown.
[0024] Figure 2 This is a flowchart of an image processing method provided according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the network structure of the target model in the image processing method provided in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of an image processing apparatus provided according to an embodiment of this application;
[0027] Figure 5 This is a flowchart of a computer-aided diagnostic method based on medical images provided according to an embodiment of this application;
[0028] Figure 6 This is a flowchart of a computer-aided diagnostic method for pancreatic cancer based on medical images, according to an embodiment of this application.
[0029] Figure 7 This is a schematic diagram of a computer-aided diagnostic system based on medical images provided according to an embodiment of this application;
[0030] Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] It should be noted that all information and data involved in this application (including but not limited to data used for training and analysis) are information and data authorized by the user or fully authorized by all parties. For example, if there is an interface between this system and the relevant user or organization, before obtaining the relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information only after receiving consent from the aforementioned user or organization.
[0034] Example 1
[0035] According to an embodiment of the present invention, an embodiment of an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0036] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0037] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0038] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image processing method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned image processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0039] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0040] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0041] Under the aforementioned operating environment, this application provides the following: Figure 2 The image processing method shown. Figure 2 This is a flowchart of an image processing method according to Embodiment 1 of the present invention.
[0042] Step S201: Receive medical images of the target organ.
[0043] Optionally, the medical image in this embodiment can be a plain CT scan, i.e., a non-contrast CT scan, thereby providing a low-cost image data source for routine examinations. The medical image includes at least image information of the target organ, and may also include image information of adjacent tissues of the target organ. The target organ can be any organ in the human body, such as the pancreas, lungs, stomach, or liver.
[0044] Optionally, the aforementioned medical images can be received through a client or through a program interface.
[0045] For example, the target organ could be the pancreas, and the medical image could include image information of the pancreas, as well as image information of adjacent tissues. The medical image could be a plain abdominal CT scan.
[0046] Step S202: Input the medical image into the target model and output multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0047] Optionally, the target model in this embodiment can be a two-stage model. In other words, the target model consists of two artificial intelligence models: a first-stage model and a second-stage model. The first stage (i.e., the first-stage model) is used to locate the target organ region based on medical images, and the second stage is used to predict the final detection result based on the located target region. The input data of the first-stage model can be plain CT images, and the output data can be three-dimensional mask images (which may include multiple voxels of binary masks). Voxels with a value of 1 represent the target region, and voxels with a value of 0 represent non-target regions. Images belonging to the target region can be cropped from the three-dimensional mask images to serve as the input data of the second-stage model (i.e., the target region image). The target region may include the sum of image information of healthy tissue and potential lesions of the target organ. The second stage of the target model (i.e., the model of the second stage) can be a multi-task learning model, which outputs three results based on the target region image: lesion segmentation result, lesion subtype result, and surrounding structure segmentation result. The lesion segmentation result is used to represent the segmentation mask of suspicious lesion areas inside the target organ (e.g., the segmentation result of pancreatic lesions). The lesion subtype result is used to represent the subtype classification of potential lesions. The surrounding structure segmentation result includes the segmentation result of the tissue structure of the target organ within a preset range in each CT slice. It should be noted that the above two-stage model refers to the processing stages of two artificial intelligence models in the target model, formed sequentially, to achieve the acquisition of the above multiple detection results based on medical images. Any scheme that implements these two stages with other names (e.g., first / second sub-networks, preprocessing, and result prediction) is considered an equivalent implementation scheme.
[0048] Optionally, when the target organ is the pancreas, the preset range can be a spatial neighborhood of ±3 cm. The segmentation result of the surrounding structures can be a segmentation mask for peripancreatic structures, such as a segmentation mask for veins, arteries, and dilated common bile ducts. The training dataset for the target model can include multiple target region image samples. In addition to lesion segmentation annotation information (used to provide lesion boundary localization information) and lesion subtype annotation information (used to provide classification basis), the target region image samples are also annotated with surrounding structure segmentation annotation information, which is used to provide features of the surrounding structures of the target organ. For example, after obtaining the target region image samples, the spatial neighborhood of the preset range can be extended outward in three-dimensional space based on the pancreatic entity boundary, and the veins, arteries, and dilated common bile ducts in the spatial neighborhood can be annotated to obtain the surrounding structure segmentation annotation information. After training with the above-mentioned target region image samples to obtain the target model, the target model can directly incorporate the surrounding structure segmentation results into the detection results, improving the identification ability of peripancreatic tissues and suppressing false detections caused by surrounding structures.
[0049] In summary, by using a target model to locate the target organ's region based on medical images, the search space is compressed in the first stage, avoiding false detections in non-target areas and significantly reducing errors caused by detection confusion. Then, in the second stage, lesion segmentation results, lesion subtype results, and surrounding structure segmentation results are obtained from the target region image. By incorporating the surrounding structure segmentation results into the detection results, the target model can pay more attention to the features of the tissues surrounding the target organ, improving its ability to identify surrounding tissues and avoiding detection errors caused by abnormalities in surrounding tissues. This significantly improves the accuracy of the target model in detecting medical images and enhances the accuracy of image processing.
[0050] To improve the accuracy of image processing, optionally, medical images are input into a target model, and the target model outputs multiple detection results, including: inputting medical images into a first-stage sub-model of the target model, and using the first-stage sub-model to output a three-dimensional mask image based on the medical images; cropping a target region image from the three-dimensional mask image in the target model, wherein the target region image is used to represent the image information of the target region; inputting the target region image into a second-stage sub-model of the target model, and using the second-stage sub-model to simultaneously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, and determining the detection results based on the lesion segmentation results, lesion subtype results, and surrounding structure segmentation results.
[0051] Optionally, the target model may include a first-stage sub-model and a second-stage sub-model. The first-stage sub-model is the same as the model described in the first stage, and the second-stage sub-model is the same as the model described in the second stage. The output data of the first-stage sub-model may be a three-dimensional mask image, where voxels with a value of 1 represent the tissue region of the target organ, and voxels with a value of 0 represent the tissue region of a non-target organ (e.g., liver, stomach, spleen, bone, etc.). Based on the voxel values, images belonging to the target region can be cropped from the three-dimensional mask image to obtain the target region image. The purpose of determining the target region image based on the first-stage sub-model is to compress the processing scope of the subsequent second stage from the original medical image to the local region of the target organ, reducing computational complexity while focusing on the target region, removing interference from image information of non-target regions, and improving the prediction accuracy of the second-stage sub-model.
[0052] Optionally, the second-stage sub-model is used to simultaneously output three results: lesion segmentation result, lesion subtype result, and surrounding structure segmentation result. The lesion segmentation result is used to represent the segmentation mask of the suspicious lesion area inside the target organ (e.g., it can be binary mask data used to identify the spatial location of the lesion), and its spatial range is limited to the range of the target area image. The lesion subtype result can be a classification label (it can be 10 categories, specifically including the probability value of each label). The surrounding structure segmentation result is the segmentation result of the tissue structure within a preset range from the target organ (e.g., the segmentation mask of the peripancreatic structure). The above three results can be merged into a structured triple as the final detection result.
[0053] For example, when the target organ is the pancreas, the classification labels for lesion subtype results can be PDAC (pancreatic ductal adenocarcinoma), PNET (pancreatic neuroendocrine tumor), SPT (solid pseudopapillary tumor), IPMN (intraductal papillary mucinous tumor), MCN (mucinous cystic tumor), CP (chronic pancreatitis), AP (acute pancreatitis), SCN (serous cystic tumor), Other (other lesions, such as pancreatic lymphoma), and no lesion.
[0054] Optionally, the first-stage sub-model can be a lightweight nnU-net (an open-source neural network model employing an encoder-decoder structure). For example, the first-stage sub-model may include multiple layers of downsampling and upsampling, each layer using a 3x3x3 convolutional kernel, with standard 3D convolutions, and outputting the final prediction result after sigmoid activation. Alternatively, it can be MA-unet (an open-source neural network model including an encoder and decoder, introducing a multi-attention mechanism to improve segmentation capabilities through channel attention, spatial attention, and cross-scale information aggregation). For example, the first-stage sub-model may include 4 layers of downsampling, 4 layers of upsampling, and a multi-attention mechanism between downsampling and upsampling, with depthwise separable 3D convolutions, and outputting the final prediction result after sigmoid activation. The second-stage sub-model can be a dual-path joint learning network, which may include a segmentation path network and a memory path network. The segmentation path network outputs lesion segmentation results and surrounding structure segmentation results, while the memory path network outputs lesion subtype results. For example, a segmentation path network can be an image segmentation network that can output lesion location information and surrounding structure information (improving the interpretability of the model and making the target model focus on the features of the surrounding structure), while a memory path network can be a network based on a classification task that can output the final classification information of the lesion.
[0055] In summary, the first-stage sub-model determines the target image region, avoiding interference from image information outside the target organ. The second-stage sub-model simultaneously outputs lesion segmentation results, lesion subtype results, and surrounding structure segmentation results, thus improving the accuracy of image processing.
[0056] To improve the accuracy of image processing, optionally, the second-stage sub-model can be used to simultaneously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image. This includes: using the segmentation path network in the second-stage sub-model to extract lesion segmentation results and surrounding structure segmentation results based on the target region image, wherein the segmentation path network includes downsampling paths and upsampling paths, with the downsampling paths and upsampling paths being skipped connections; extracting spatial features with location encoding corresponding to the feature maps from multiple feature maps of the upsampling paths to obtain multiple spatial features; and using the memory path network in the second-stage sub-model to generate lesion subtype results based on the multiple spatial features and a preset memory matrix.
[0057] Optionally, the second-stage sub-model may include a segmentation path network and a memory path network. The input data of the segmentation path network may be the target region image, and the output data may be the lesion segmentation result and the surrounding structure segmentation result.
[0058] Optionally, Figure 3 This is a schematic diagram of the network structure of the target model in the image processing method provided in the embodiments of this application, with reference to... Figure 3As shown, the segmentation path network can be a multi-layer U-net, specifically including downsampling paths and upsampling paths. The downsampling path includes a preset number (e.g., 5) of 3D convolutional layers. These 3D convolutional layers can include convolutional blocks, batch normalization operations, activation operations (e.g., ReLU activation, also known as linear rectified unit activation), and max pooling operations. As the number of layers increases, the spatial resolution of each 3D convolutional layer can gradually decrease (e.g., the spatial resolution is halved with each additional layer), while the number of channels can gradually increase (e.g., the number of channels doubles with each additional layer), thereby extracting multi-level features from local texture to global semantics of the target image region. The upsampling path can include a preset number of 3D deconvolutional layers. These 3D deconvolutional layers can gradually restore the spatial resolution as the number of layers increases. Furthermore, skip connections are established between the downsampling and upsampling paths. These skip connections are used to directly transfer feature maps from the downsampling path to the corresponding layers in the upsampling path, thereby preserving high-resolution edge and detail information and significantly improving the segmentation accuracy of small lesions and structural boundaries. The segmentation path network can also embed a positional encoding. For example, a learnable positional encoding can be embedded in the feature map of each layer of the upsampled path. This positional encoding can be a tensor with the same dimension as the feature map. The encoded value is automatically optimized through backpropagation during training. The positional encoding is used to encode the three-dimensional spatial coordinates of each voxel in the target region, thereby providing the memory path network with the spatial relationship between the lesion and the surrounding structure in the target region image.
[0059] Optionally, refer to Figure 3 As shown, a memory path network can include at least a pre-defined memory matrix, which can include multiple learnable memory embedding vectors. Each memory embedding vector encodes features of a corresponding lesion subtype (i.e., features of a certain lesion subtype in terms of texture, density, morphology, and spatial distribution). The pre-defined memory matrix is part of the network parameters of the target model and is updated through backpropagation during model training. The memory path network is used for classification using spatial features with positional encoding extracted from upsampled paths (i.e., features obtained by fusing positional encoding and feature maps) and the aforementioned pre-defined memory matrix. Specifically, it can be a network structure for classification (e.g., it can include multiple hidden layers, fully connected layers, and softmax layers for classification tasks). For example, the memory path network is a dual-path transformer architecture (one path is based on the aforementioned spatial features, and the other path is the pre-defined memory matrix). Network The information from the two paths mentioned above is fused through multiple transformer layers, and the classification result is finally output through a fully connected layer and a softmax layer.
[0060] In summary, accurate semantic segmentation of target region images was achieved through downsampling and upsampling paths in the segmentation path network, yielding lesion segmentation results and surrounding structure segmentation results. By extracting spatial features with location encoding corresponding to multiple feature maps from the upsampling path, rich information on lesion features and surrounding structures was provided to the memory path network. Furthermore, accurate classification was achieved by combining the preset memory matrix, resulting in lesion subtype results and improving the accuracy of image processing.
[0061] To improve the accuracy of image processing, optionally, the memory path network in the second-stage sub-model can be used to generate lesion subtype results based on multiple spatial features and a preset memory matrix. This includes: using multiple attention interaction layers in the memory path network to encode multiple spatial features and the preset memory matrix to obtain a discriminative feature vector. The discriminative feature vector is used to encode the texture pattern and spatial structure information of the lesion. There is a one-to-one correspondence between the attention interaction layer and the spatial features. The first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vector. The discriminative feature vector is then input into the classification layer in the memory path network, and the classification layer is used to output the lesion subtype results.
[0062] Optionally, refer to Figure 3 As shown, the aforementioned spatial features are extracted from the upsampled paths of the segmentation path and are fused with learnable location encodings and feature maps. Each spatial feature corresponds to an attention interaction layer. The preset memory matrix includes multiple memory embedding vectors. The attention interaction layer is used to interact with the spatial features and the preset memory matrix. Each update can include self-attention interaction and cross-attention interaction. The classification layer can include at least one fully connected layer and a softmax layer. Its input is a discriminative feature vector, and its output is the probability of different lesion categories, thereby determining the category with the highest probability as the lesion subtype result.
[0063] For example, the input to the first attention interaction layer can be the first layer of spatial features and a preset memory matrix, and the output is the interaction result obtained by interacting with the spatial features and the preset memory matrix. The interaction result can be the updated preset memory matrix (including the updated memory embedding vector). The interaction result of the previous attention interaction layer is used as the input to the next attention interaction layer, and the next attention interaction layer interacts based on the above interaction result and the spatial features corresponding to the next attention interaction layer. After multiple interactions, the output of the last attention interaction layer is determined as the discriminative feature vector.
[0064] For example, the interaction process in each attention interaction layer can include self-attention interaction and cross-attention interaction. Taking the first attention interaction layer as an example, memory is a preset memory matrix, and pixel embedding is spatial features. A fully connected layer can perform self-attention interaction on the preset memory matrix to obtain a first self-attention interaction result (which may include query information Q, key information K, and value information V). Another fully connected layer can perform self-attention interaction on the spatial features to obtain a second self-attention interaction result (which may include query information Q, key information K, and value information V). A cross-attention mechanism formula can be used to perform cross-attention interaction based on the above first and second attention interaction results (for example, the first self-attention interaction result provides Q and the second self-attention interaction result provides K and V, or the second self-attention interaction result provides Q and the first self-attention interaction result provides K and V) to obtain the above interaction result.
[0065] In summary, by utilizing a preset memory matrix and multiple spatial features for layer-by-layer interaction, information from the preset memory matrix and spatial features is fully integrated to obtain a discriminative feature vector. Based on the discriminative feature vector, the lesion subtype result is obtained at the classification layer, thus improving the accuracy of image processing.
[0066] To improve the accuracy of image processing, the training steps of the target model may optionally include: acquiring a first training dataset and training a first preset neural network model using the first training dataset to obtain a first-stage sub-model; acquiring a second training dataset and training a second preset neural network model using the second training dataset to obtain a second-stage sub-model; and determining the target model based on the first-stage sub-model and the second-stage sub-model.
[0067] For example, the first training dataset may include multiple plain CT scan samples, each containing a voxel-level labeled target organ mask (the mask is binary data, where a voxel with a value of 1 represents the target region and a voxel with a value of 0 represents the non-target region). The target region represents the sum of healthy tissue and potential lesions. By training the first preset neural network model using the first training dataset, the model can learn how to segment target and non-target regions based on plain CT scans, thus obtaining the first-stage sub-model. The second training dataset may include multiple target region image samples, each labeled with lesion segmentation annotations, lesion subtype annotations, and surrounding structure segmentation annotations. By training the second preset neural network model using the second training dataset, the model can learn how to predict lesion segmentation annotations, lesion subtype annotations, and surrounding structure annotations based on target region images, thus obtaining the second-stage sub-model. Combining the first-stage and second-stage sub-models yields the target model.
[0068] In summary, by training the first preset neural network model using the first training dataset and the second preset neural network model using the second training dataset, a two-stage target model is obtained, thereby improving the accuracy of image processing.
[0069] To improve the accuracy of image processing, optionally, the first training samples in the first training dataset include a union mask of healthy tissue of the target organ and potential lesions. Training a first preset neural network model using the first training dataset to obtain a first-stage sub-model includes: inputting the first training samples into the first preset neural network model; updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask; and determining the first-stage sub-model based on the updated first preset neural network model.
[0070] Optionally, the target organ healthy tissue can be healthy pancreatic tissue, and the potential lesion can be the area subsequently diagnosed as a pancreatic lesion in a case (whether or not it is visible in plain CT scan), and a set mask, i.e. a three-dimensional binary voxel mask, is used to enable the first preset neural network model to learn the complete morphology of the pancreas.
[0071] Optionally, the first preset neural network model can be a lightweight nnUNet (e.g., it can include multi-layer downsampling and multi-layer upsampling, each layer can use a 3x3x3 convolutional kernel, and output the final prediction result after sigmoid activation). The first training dataset can include multiple plain CT samples, each of which includes a target organ mask annotated at the voxel level (the mask is binary data, where a voxel with a value of 1 represents the target region and a voxel with a value of 0 represents the non-target region). The input of the first preset neural network model is the plain CT sample, and its output is the mask prediction values of the target region and the non-target region. The predicted contour is the boundary contour of the target region and the non-target region obtained by dividing the output data of the first preset neural network model based on the binary mask, while the true contour refers to the contour of the target region and the non-target region obtained by dividing each plain CT sample in the first training dataset based on the pre-annotated mask. By minimizing the deviation between the predicted contour and the true contour, the parameters of the first preset neural network model are updated. When the model converges or reaches the preset number of iterations, the above-mentioned first-stage sub-model can be obtained.
[0072] In summary, by inputting the first training sample into the first preset neural network model, updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask, and determining the first stage sub-model based on the updated first preset neural network model, the accuracy of image processing is improved.
[0073] To improve the accuracy of image processing, optionally, the second training samples in the second training dataset include lesion segmentation annotation information, lesion subtype annotation information, and surrounding structure segmentation annotation information corresponding to the target organ. The second training dataset is used to train the second preset neural network model to obtain the second-stage sub-model, which includes: inputting the second training samples into the second preset neural network model and simultaneously training multiple tasks, including lesion segmentation, lesion subtype classification, and surrounding structure segmentation; determining the sum of errors between the prediction results of the multiple tasks and the corresponding annotation information, and updating the parameters of the second preset neural network model based on the sum of errors; and determining the second-stage sub-model based on the updated second preset neural network model.
[0074] Optionally, the second training dataset may include multiple target region image samples. Each target region image sample is labeled with lesion segmentation annotation information, lesion subtype annotation information, and surrounding structure segmentation annotation information. The lesion segmentation annotation information is used to provide lesion boundary localization information, the lesion subtype annotation information is used to provide classification basis, and the surrounding structure segmentation annotation information is used to provide features of the surrounding structures of the target organ.
[0075] Optionally, the second preset neural network model can be a dual-path joint learning architecture, specifically including a segmentation path (corresponding to the segmentation path network in the inference stage) and a classification path (corresponding to the memory path network in the inference stage). The segmentation path can be a multi-layer U-net (including encoder and decoder), with the input being target region image samples and the output being lesion segmentation results and surrounding structure segmentation results. The classification path can be a transformer-based classification model (for example, corresponding to the network structure of the memory path network in the inference stage), with its input being feature maps and location codes extracted from each layer of the segmentation path and its output being lesion subtype results. The training objective of the second preset neural network model can be to minimize the sum of errors between the above three outputs and the corresponding labeled information. For example, a hybrid loss function can be set, using the weighted sum of multiple different loss functions as the loss function of the second preset neural network model. When the model converges or reaches a preset number of iterations, the above-mentioned second-stage sub-model can be obtained.
[0076] In summary, by using the second training dataset to train the second preset neural network model for multi-task learning, the aforementioned second-stage sub-model was obtained, which improved the accuracy of image processing.
[0077] To improve the accuracy of image processing, optionally, determining the target model based on the first-stage sub-model and the second-stage sub-model includes: extracting target false positive samples from the target log, wherein the target log is a log of medical images of the target organ being detected according to a preset model, and the target false positive samples are samples in which the target organ is normal but the surrounding structures of the target organ are abnormal; and incrementally learning the first-stage sub-model and the second-stage sub-model based on the target false positive samples to obtain the target model.
[0078] Optionally, the preset model can be the aforementioned target model, or a three-stage artificial intelligence model for medical image detection in the prior art, or a model in other existing technologies (in this case, the collected false positive samples can help the incrementally learned target model achieve better accuracy). The target log can be a structured operational record generated by the preset model deployed in a real clinical environment for medical image detection, which may include the original medical images, detection results, and actual results. The aforementioned false positive samples can be collected from different medical institutions to enhance sample diversity.
[0079] For example, taking the pancreas as the target organ, a target false positive sample could be a sample where the pancreas is normal, but the preset model misclassifies abnormalities in peripancreatic tissues (e.g., common bile duct dilatation, duodenal diverticulum, etc.) as pancreatic abnormalities. This target false positive sample is then used to incrementally learn the first-stage and second-stage sub-models. This synergizes with the addition of surrounding structure prediction results to the detection results, further enhancing the target model's ability to identify peripancreatic tissues, avoiding misclassifications caused by peripancreatic abnormalities, and improving the accuracy of the target model. This incremental learning refers to fine-tuning and updating the parameters of the existing model using only the newly added target false positive samples without retraining the entire model. This improves the model's ability to identify new error patterns while preserving its original performance.
[0080] In summary, by acquiring false positive samples caused by abnormalities in the surrounding structure and incrementally learning the first-stage sub-model and the second-stage sub-model, the target model is obtained, avoiding misjudgments caused by abnormalities in the surrounding structure and further improving the accuracy of the target model.
[0081] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0083] Example 2
[0084] According to embodiments of the present invention, an image processing apparatus for implementing the above-described image processing method is also provided, such as... Figure 4 As shown, the device includes a data receiving unit 401 and a result detection unit 402.
[0085] Specifically, the data receiving unit 401 is used to receive medical images of the target organ;
[0086] The result detection unit 402 is used to input medical images into the target model and output multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0087] The image processing apparatus provided in this application embodiment receives medical images of a target organ through a data receiving unit 401, inputs the medical images into a target model through a result detection unit 402, and outputs multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This solves the technical problem of low image processing accuracy in the prior art and achieves the technical effect of improving image processing accuracy.
[0088] Optionally, in the image processing apparatus provided in this application embodiment, the result detection unit 402 includes: a first prediction module, used to input medical images into a first-stage sub-model in the target model, and use the first-stage sub-model to output a three-dimensional mask image based on the medical images; in the target model, to crop a target region image from the three-dimensional mask image, wherein the target region image is used to represent the image information of the target region; and a second prediction module, used to input the target region image into a second-stage sub-model in the target model, and use the second-stage sub-model to synchronously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, and to determine the detection result based on the lesion segmentation results, lesion subtype results, and surrounding structure segmentation results.
[0089] Optionally, in the image processing apparatus provided in this application embodiment, the second prediction module includes: a first path submodule, used to extract lesion segmentation results and surrounding structure segmentation results based on the target region image using the segmentation path network in the second stage sub-model, wherein the segmentation path network includes a downsampling path and an upsampling path, and the downsampling path and the upsampling path are connected in a skip connection; a feature extraction submodule, used to extract spatial features with positional encoding corresponding to the feature maps from multiple feature maps of the upsampling path, to obtain multiple spatial features; and a second path submodule, used to generate lesion subtype results based on multiple spatial features and a preset memory matrix using the memory path network in the second stage sub-model.
[0090] Optionally, in the image processing apparatus provided in this application embodiment, the second path submodule includes: a memory encoding component, used to encode multiple spatial features and a preset memory matrix using multiple attention interaction layers in the memory path network to obtain a discriminative feature vector, wherein the discriminative feature vector is used to encode the texture pattern and spatial structure information of the lesion, the attention interaction layer and the spatial feature have a one-to-one correspondence, the first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vector; and a classification component, used to input the discriminative feature vector into the classification layer in the memory path network, and use the classification layer to output the lesion subtype result.
[0091] Optionally, the image processing apparatus provided in this application embodiment further includes: a first training unit, configured to acquire a first training dataset and train a first preset neural network model using the first training dataset to obtain a first-stage sub-model; a second training unit, configured to acquire a second training dataset and train a second preset neural network model using the second training dataset to obtain a second-stage sub-model; and a model determination unit, configured to determine a target model based on the first-stage sub-model and the second-stage sub-model.
[0092] Optionally, in the image processing apparatus provided in this application embodiment, the first training samples in the first training dataset include a union mask of healthy tissue of the target organ and potential lesions. The first training unit includes: a first input module for inputting the first training samples into a first preset neural network model; a first update module for updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask; and a first determination module for determining a first-stage sub-model based on the updated first preset neural network model.
[0093] Optionally, in the image processing apparatus provided in this application embodiment, the second training samples in the second training dataset include lesion segmentation annotation information, lesion subtype annotation information, and surrounding structure segmentation annotation information corresponding to the target organ. The second training unit includes: a second input module, used to input the second training samples into a second preset neural network model and simultaneously train multiple tasks, wherein the multiple tasks include a lesion segmentation task, a lesion subtype classification task, and a surrounding structure segmentation task; a second update module, used to determine the sum of errors between the prediction results of the multiple tasks and the corresponding annotation information, and update the parameters of the second preset neural network model based on the sum of errors; and a second determination module, used to determine a second-stage sub-model based on the updated second preset neural network model.
[0094] Optionally, in the image processing apparatus provided in this application embodiment, the model determination unit includes: a sample acquisition module, used to extract target false positive samples from the target log, wherein the target log is a log of medical images of target organs detected according to a preset model, and the target false positive samples are samples in which the target organ has no abnormalities but the surrounding structures of the target organ have abnormalities; and an incremental learning module, used to perform incremental learning on the first stage sub-model and the second stage sub-model respectively based on the target false positive samples to obtain the target model.
[0095] It should be noted that the data receiving unit 401 and the result detection unit 402 mentioned above correspond to steps S201 to S202 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0096] Example 3
[0097] According to embodiments of this application, a computer-aided diagnostic method based on medical images is also provided, such as... Figure 5 As shown, the method includes:
[0098] S501, Receive medical images of the target organ, wherein the medical images are used to detect whether there are any abnormalities in the target organ;
[0099] S502, input the medical image into the target model, and use the target model to output multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0100] The computer-aided diagnostic method based on medical images provided in this application receives medical images of a target organ, which are used to detect whether the target organ has any abnormalities. The medical images are input into a target model, which then outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image, where the target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This method solves the technical problem of low accuracy in medical image-based diagnosis in existing technologies, thereby improving the accuracy of medical image-based diagnosis.
[0101] The computer-aided diagnostic method based on medical images corresponds to the method in Example 1, and will not be described again here.
[0102] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0104] Example 4
[0105] According to embodiments of this application, a computer-aided diagnostic method for pancreatic cancer based on medical imaging is also provided, such as... Figure 6 As shown, the method includes:
[0106] S601 receives medical images of the pancreas, wherein the medical images are used to detect whether there are any abnormalities in the pancreas;
[0107] S602, input the medical image into the target model, and use the target model to output multiple detection results. The first stage of the target model is used to locate the target area where the pancreas is located based on the medical image. The second stage of the target model is used to output multiple detection results based on the pancreatic area image. The pancreatic area image is the image of the target area in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the pancreas.
[0108] The computer-aided diagnosis method for pancreatic cancer based on medical images provided in this application receives medical images of the pancreas, which are used to detect the presence of abnormalities in the pancreas. The medical images are input into a target model, which then outputs multiple detection results. The first stage of the target model is used to locate the target region of the pancreas based on the medical images. The second stage of the target model outputs multiple detection results based on the pancreatic region image, where the pancreatic region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the pancreas. This method solves the technical problem of low accuracy in diagnosing pancreatic cancer based on medical images in existing technologies, thereby improving the accuracy of pancreatic cancer diagnosis based on medical images.
[0109] Optionally, the classification labels for lesion subtype results may include PDAC (pancreatic ductal adenocarcinoma), PNET (pancreatic neuroendocrine tumor), SPT (solid pseudopapillary tumor), IPMN (intraductal papillary mucinous tumor), MCN (mucinous cystic tumor), CP (chronic pancreatitis), AP (acute pancreatitis), SCN (serous cystic tumor), Other (other lesions, such as pancreatic lymphoma), and no lesion.
[0110] The computer-aided diagnostic method for pancreatic cancer based on medical images corresponds to the method in Example 1, and will not be described again here.
[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0113] Example 5
[0114] According to embodiments of this application, a computer-aided diagnostic system based on medical images is also provided, such as... Figure 7 As shown, the system includes:
[0115] Client 701 and server 702 are used to upload medical images of the target organ to the server, whereby the medical images are used to detect whether there are any abnormalities in the target organ. Server 702 is used to input the medical images into the target model and output multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0116] The computer-aided diagnostic system based on medical images provided in this application uses a client to upload medical images of a target organ to a server. The medical images are used to detect whether the target organ has any abnormalities. The server inputs the medical images into a target model and outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This solves the technical problem of low image processing accuracy in the prior art, thereby achieving the technical effect of improving the accuracy of image processing.
[0117] The computer-aided diagnostic system based on medical images corresponds to the method in Example 1, and will not be described again here.
[0118] Example 6
[0119] Embodiments of the present invention can provide a computer terminal, which can be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the computer terminal can also be replaced by a mobile terminal or other terminal device.
[0120] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0121] In this embodiment, the computer terminal described above can execute the program code for the following steps in the image processing method: receiving medical images of the target organ; inputting the medical images into a target model; and using the target model to output multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0122] Optionally, the aforementioned computer terminal can execute program code for the following steps in the image processing method: inputting medical images into a first-stage sub-model in the target model, and using the first-stage sub-model to output a three-dimensional mask image based on the medical images; cropping a target region image from the three-dimensional mask image in the target model, wherein the target region image is used to represent the image information of the target region; inputting the target region image into a second-stage sub-model in the target model, and using the second-stage sub-model to synchronously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, and determining the detection results based on the lesion segmentation results, lesion subtype results, and surrounding structure segmentation results.
[0123] Optionally, the aforementioned computer terminal can execute the program code for the following steps in the image processing method: extracting lesion segmentation results and surrounding structure segmentation results from the target region image using the segmentation path network in the second-stage sub-model, wherein the segmentation path network includes downsampling paths and upsampling paths, with the downsampling paths and upsampling paths being skipped connections; extracting spatial features with location encoding corresponding to the feature maps from multiple feature maps of the upsampling paths to obtain multiple spatial features; and generating lesion subtype results based on multiple spatial features and a preset memory matrix using the memory path network in the second-stage sub-model.
[0124] Optionally, the aforementioned computer terminal can execute the program code for the following steps in the image processing method: using multiple attention interaction layers in the memory path network to encode multiple spatial features and a preset memory matrix to obtain a discriminative feature vector, wherein the discriminative feature vector is used to encode the texture pattern and spatial structure information of the lesion, the attention interaction layer and the spatial feature have a one-to-one correspondence, the first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vector; inputting the discriminative feature vector into the classification layer in the memory path network, and using the classification layer to output the lesion subtype result.
[0125] Optionally, the computer terminal described above can execute program code for the following steps in the image processing method: obtaining a first training dataset and training a first preset neural network model using the first training dataset to obtain a first-stage sub-model; obtaining a second training dataset and training a second preset neural network model using the second training dataset to obtain a second-stage sub-model; and determining a target model based on the first-stage sub-model and the second-stage sub-model.
[0126] Optionally, the aforementioned computer terminal may execute program code for the following steps in the image processing method: inputting the first training sample into the first preset neural network model; updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask; and determining the first stage sub-model based on the updated first preset neural network model.
[0127] Optionally, the aforementioned computer terminal can execute program code for the following steps in the image processing method: inputting the second training sample into the second preset neural network model and simultaneously training multiple tasks, wherein the multiple tasks include lesion segmentation task, lesion subtype classification task, and surrounding structure segmentation task; determining the total error between the prediction results of the multiple tasks and the corresponding annotation information, and updating the parameters of the second preset neural network model based on the total error; and determining the second-stage sub-model based on the updated second preset neural network model.
[0128] Optionally, the aforementioned computer terminal can execute program code for the following steps in the image processing method: extracting target false positive samples from the target log, wherein the target log is a log of medical images of the target organ being detected based on a preset model, and the target false positive sample is a sample in which the target organ is normal but the surrounding structure of the target organ is abnormal; and incrementally learning the first-stage sub-model and the second-stage sub-model based on the target false positive samples to obtain the target model.
[0129] Optionally, Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 8 As shown, the electronic device may include: one or more ( Figure 8 (Only one is shown) processor 802, memory 804, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0130] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and apparatus in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned image processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0131] The processor can access the information and application programs stored in the memory via the transmission device to perform the above steps.
[0132] This invention provides an image processing solution. It involves receiving medical images of a target organ; inputting the medical images into a target model; and using the target model to output multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is the image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ. This solution addresses the technical problem of low image processing accuracy in existing technologies, thereby achieving the technical effect of improving image processing accuracy.
[0133] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone, tablet, handheld computer, mobile internet device (MID), PAD, or other terminal device. Figure 8 This does not limit the structure of the aforementioned electronic devices. For example, a computer terminal may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.
[0134] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0135] Example 7
[0136] Embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image processing method provided in Embodiment 1.
[0137] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0138] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: receiving medical images of a target organ; inputting the medical images into a target model; and using the target model to output multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical images. The second stage of the target model is used to output multiple detection results based on the target region image. The target region image is an image of the target region in the medical images. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
[0139] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: inputting medical images into a first-stage sub-model in the target model, and using the first-stage sub-model to output a three-dimensional mask image based on the medical images; cropping a target region image from the three-dimensional mask image in the target model, wherein the target region image is used to represent image information of the target region; inputting the target region image into a second-stage sub-model in the target model, and using the second-stage sub-model to synchronously output lesion segmentation results, lesion subtype results, and surrounding structure segmentation results based on the target region image, and determining detection results based on the lesion segmentation results, lesion subtype results, and surrounding structure segmentation results.
[0140] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: extracting lesion segmentation results and surrounding structure segmentation results from the target region image using the segmentation path network in the second-stage sub-model, wherein the segmentation path network includes downsampling paths and upsampling paths, and the downsampling paths and upsampling paths are skipped; extracting spatial features with positional encoding corresponding to the feature maps from multiple feature maps of the upsampling paths to obtain multiple spatial features; and generating lesion subtype results based on multiple spatial features and a preset memory matrix using the memory path network in the second-stage sub-model.
[0141] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: using multiple attention interaction layers in the memory path network to encode multiple spatial features and a preset memory matrix to obtain a discriminative feature vector, wherein the discriminative feature vector is used to encode the texture pattern and spatial structure information of the lesion, the attention interaction layer and the spatial feature have a one-to-one correspondence, the first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vector; the discriminative feature vector is input into the classification layer in the memory path network, and the classification layer is used to output the lesion subtype result.
[0142] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: obtaining a first training dataset and training a first preset neural network model using the first training dataset to obtain a first-stage sub-model; obtaining a second training dataset and training a second preset neural network model using the second training dataset to obtain a second-stage sub-model; and determining a target model based on the first-stage sub-model and the second-stage sub-model.
[0143] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: inputting a first training sample into a first preset neural network model; updating the parameters of the first preset neural network model by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the real contour corresponding to the union mask; and determining a first-stage sub-model based on the updated first preset neural network model.
[0144] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: inputting a second training sample into a second preset neural network model and simultaneously training multiple tasks, wherein the multiple tasks include a lesion segmentation task, a lesion subtype classification task, and a surrounding structure segmentation task; determining the sum of errors between the prediction results of the multiple tasks and the corresponding annotation information, and updating the parameters of the second preset neural network model based on the sum of errors; and determining a second-stage sub-model based on the updated second preset neural network model.
[0145] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: extracting target false positive samples from the target log, wherein the target log is a log of medical images of the target organ being detected according to a preset model, and the target false positive sample is a sample in which the target organ is normal but the surrounding structure of the target organ is abnormal; and incrementally learning the first-stage sub-model and the second-stage sub-model based on the target false positive samples to obtain the target model.
[0146] This application also provides a computer program product, which, when executed on a data processing device, is adapted to perform image processing method steps.
[0147] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0148] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0149] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0151] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0152] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0153] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that, include: Receive medical images of the target organ; The medical image is input into the target model, and the target model outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output the multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
2. The method according to claim 1, characterized in that, The medical images are input into the target model, and the target model outputs multiple detection results, including: The medical image is input into the first-stage sub-model of the target model, and the first-stage sub-model is used to output a three-dimensional mask image based on the medical image. In the target model, a target region image is cropped from the three-dimensional mask image, wherein the target region image is used to represent the image information of the target region; The target region image is input into the second-stage sub-model in the target model. The second-stage sub-model synchronously outputs the lesion segmentation result, the lesion subtype result, and the surrounding structure segmentation result based on the target region image. The multiple detection results are determined based on the lesion segmentation result, the lesion subtype result, and the surrounding structure segmentation result.
3. The method according to claim 2, characterized in that, The second-stage sub-model synchronously outputs the lesion segmentation result, the lesion subtype result, and the surrounding structure segmentation result based on the target region image, including: The segmentation path network in the second stage sub-model is used to extract the lesion segmentation result and the surrounding structure segmentation result based on the target region image. The segmentation path network includes a downsampling path and an upsampling path, and the downsampling path and the upsampling path are connected in a skip connection. Multiple spatial features are obtained by extracting spatial features with positional encoding corresponding to the feature maps of the upsampled path from multiple feature maps; The lesion subtype result is generated based on the multiple spatial features and the preset memory matrix using the memory path network in the second stage sub-model.
4. The method according to claim 3, characterized in that, The generation of the lesion subtype result using the memory path network in the second-stage sub-model based on the multiple spatial features and the preset memory matrix includes: The multiple attention interaction layers in the memory path network are used to encode the multiple spatial features and the preset memory matrix to obtain a discriminative feature vector. The discriminative feature vector is used to encode the texture pattern and spatial structure information of the lesion. The attention interaction layer and the spatial feature have a one-to-one correspondence. The first attention interaction layer is used to receive the preset memory matrix, and the last attention interaction layer is used to output the discriminative feature vector. The discriminative feature vector is input into the classification layer of the memory path network, and the classification layer is used to output the lesion subtype result.
5. The method according to claim 1, characterized in that, The training steps for the target model include: Obtain a first training dataset and use the first training dataset to train a first preset neural network model to obtain a first-stage sub-model; Obtain a second training dataset and use the second training dataset to train a second preset neural network model to obtain a second-stage sub-model; The target model is determined based on the first stage sub-model and the second stage sub-model.
6. The method according to claim 5, characterized in that, The first training samples in the first training dataset include a union mask of healthy tissue and potential lesions of the target organ. The first preset neural network model is trained using the first training dataset to obtain a first-stage sub-model, which includes: The first training sample is input into the first preset neural network model; The parameters of the first preset neural network model are updated by minimizing the difference between the predicted contour of the output data of the first preset neural network model and the true contour corresponding to the union mask. The first stage sub-model is determined based on the updated first preset neural network model.
7. The method according to claim 5, characterized in that, The second training samples in the second training dataset include lesion segmentation annotation information, lesion subtype annotation information, and surrounding structure segmentation annotation information corresponding to the target organ. The second training dataset is used to train the second preset neural network model to obtain the second-stage sub-model, which includes: The second training sample is input into the second preset neural network model, and multiple tasks are trained simultaneously, including lesion segmentation task, lesion subtype classification task and surrounding structure segmentation task. The sum of errors between the prediction results of the multiple tasks and the corresponding annotation information is determined, and the parameters of the second preset neural network model are updated based on the sum of errors. The second stage sub-model is determined based on the updated second preset neural network model.
8. The method according to claim 5, characterized in that, Determining the target model based on the first-stage sub-model and the second-stage sub-model includes: Target false positive samples are extracted from the target log, wherein the target log is a log of medical images of the target organ being detected according to a preset model, and the target false positive sample is a sample in which the target organ is normal but the surrounding structure of the target organ is abnormal. Based on the target false positive samples, incremental learning is performed on the first stage sub-model and the second stage sub-model respectively to obtain the target model.
9. A computer-aided diagnostic method based on medical images, characterized in that, include: Receive medical images of a target organ, wherein the medical images are used to detect whether there are any abnormalities in the target organ; The medical image is input into the target model, and the target model outputs multiple detection results. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output the multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
10. A computer-aided diagnostic method for pancreatic cancer based on medical imaging, characterized in that, include: Receive medical images of the pancreas, wherein the medical images are used to detect whether there are any abnormalities in the pancreas; The medical image is input into the target model, and the target model outputs multiple detection results. The first stage of the target model is used to locate the target region where the pancreas is located based on the medical image. The second stage of the target model is used to output the multiple detection results based on the pancreatic region image. The pancreatic region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the pancreas.
11. A computer-aided diagnostic system based on medical images, characterized in that, Including client and server sides, among which: The client is used to upload medical images of the target organ to the server, wherein the medical images are used to detect whether there are any abnormalities in the target organ; The server is used to input the medical image into a target model and output multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output the multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
12. An image processing apparatus, characterized in that, include: A data receiving unit is used to receive medical images of the target organ; The result detection unit is used to input the medical image into the target model and output multiple detection results using the target model. The first stage of the target model is used to locate the target region where the target organ is located based on the medical image. The second stage of the target model is used to output the multiple detection results based on the target region image. The target region image is the image of the target region in the medical image. The multiple detection results include lesion segmentation results, lesion subtype results, and surrounding structure segmentation results. The surrounding structure segmentation results are the segmentation results of tissue structures within a preset range from the target organ.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 10.
14. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 10.
15. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method described in any one of claims 1 to 10.