Target object identification method, object identification model training method, method for detecting visible lymph node in CT image, computer aided diagnosis method, electronic device, storage medium, and program product
Through the Transformer-based object recognition model, combined with multi-scale 2.5D feature fusion and IoU-guided query selection, the inaccuracy problem of lymph node detection in the neural network model in the images is solved, and higher accuracy lymph node recognition is achieved.
Patent Information
- Application Number
- PCT/CN2025/071891
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-01-10
- Publication Date
- 2025-09-04
AI Technical Summary
Existing neural network models cannot accurately identify the location of specific objects in the image, resulting in inaccurate recognition.
Using a Transformer-based object recognition model, the accuracy of lymph node detection is improved by multi-scale 2.5D feature fusion and pre-trained 2D weight combined with 3D context, using IoU prediction head and IoU-guided query selection.
Accurate identification of target objects such as lymph nodes in the image is achieved, avoiding inaccurate detection results and improving the accuracy of detection.
Smart Images

Figure CN2025071891_04092025_PF_FP_ABST
Abstract
Description
Target object recognition method, object recognition model training method, visible lymph node detection method in CT image, computer-aided diagnosis method, electronic device, storage medium and program product
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on February 27, 2024, with application number 202410217003.4 and application name “Target object recognition method, object recognition model training method, target object processing method and information processing method”, the entire content of which is incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to the field of computer technology, and more particularly to a target object recognition method, a computing device, and a storage medium. Background Art
[0003] With the continuous development of computer technology, neural network models are widely used in various service scenarios. In image processing scenarios, neural network models can be used to identify specific objects in images, thereby providing services for image processing scenarios.
[0004] However, in the existing technology, the neural network model cannot accurately identify the object position of a specific object in an image, which leads to inaccurate identification of the specific object. Therefore, how to accurately identify the specific object in the image has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide a method for identifying a target object. One or more embodiments of the present disclosure also relate to a method for training an object recognition model, a method for identifying lesions in liver CT images, a method for detecting visible lymph nodes in CT images, a computer-aided diagnosis method, a method for processing a target object, an information processing method, a target object recognition device, an object recognition model training device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0006] According to a first aspect of an embodiment of the present disclosure, a target object recognition method is provided, including:
[0007] Determine an image to be identified;
[0008] The image to be identified is input into an object recognition model for object recognition to obtain a target object in the image to be identified, wherein the object recognition model is obtained by training sample target objects identified from sample images and sample labels of the sample images, the sample target objects are determined from the multiple sample candidate objects through object type recognition results and object position detection results of multiple sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects, and the multiple sample candidate objects are obtained by performing object recognition on the sample image.
[0009] According to a second aspect of an embodiment of the present disclosure, there is provided a target object recognition device, comprising:
[0010] An image determination module is configured to determine an image to be identified;
[0011] An object recognition module is configured to input the image to be recognized into an object recognition model for object recognition to obtain a target object in the image to be recognized, wherein the object recognition model is obtained by training sample target objects identified from sample images and sample labels of the sample images, the sample target objects are determined from the multiple sample candidate objects through object type recognition results and object position detection results of multiple sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects, and the multiple sample candidate objects are obtained by performing object recognition on the sample images.
[0012] According to a third aspect of an embodiment of the present disclosure, a method for training an object recognition model is provided, comprising:
[0013] Determine a sample image of the object recognition model to be trained, and a sample label corresponding to the sample image;
[0014] Inputting the sample image into the object recognition model to be trained, performing object recognition on the sample image using the object recognition model to be trained, and obtaining a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object;
[0015] Obtaining an object position detection result for each sample candidate object by performing position detection on the object positions of the plurality of sample candidate objects;
[0016] determining a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result;
[0017] Based on the sample target objects and the sample labels, the object recognition model to be trained is trained to obtain an object recognition model.
[0018] According to a fourth aspect of an embodiment of the present disclosure, there is provided an object recognition model training device, comprising:
[0019] A sample determination module is configured to determine a sample image of an object recognition model to be trained, and a sample label corresponding to the sample image;
[0020] a sample candidate object recognition module configured to input the sample image into the to-be-trained object recognition model, perform object recognition on the sample image using the to-be-trained object recognition model, and obtain a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object;
[0021] a position detection module configured to obtain an object position detection result of each sample candidate object by performing position detection on the object positions of the multiple sample candidate objects;
[0022] a sample target object recognition module configured to determine a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result;
[0023] The model training module is configured to train the object recognition model to be trained based on the sample target object and the sample label to obtain an object recognition model.
[0024] According to a fifth aspect of an embodiment of the present disclosure, a method for identifying lesions in a liver CT image is provided, comprising:
[0025] Identify the liver CT images containing the lesion;
[0026] The liver CT image is input into a lesion recognition model for lesion recognition to obtain a target lesion in the liver CT image containing the lesion, wherein the lesion recognition model is obtained by training sample target lesions identified from a liver CT sample image containing the lesion and sample labels of the liver CT sample image containing the lesion, the sample target lesions are determined from the multiple sample candidate lesions through lesion type recognition results and lesion position detection results of the multiple sample candidate lesions, the lesion position detection results are obtained by performing position detection on the lesion positions of the multiple sample candidate lesions, and the multiple sample candidate lesions are obtained by performing lesion recognition on the liver CT sample image containing the lesion.
[0027] According to a sixth aspect of an embodiment of the present disclosure, a computer-aided diagnosis method is provided, which is applied to a client of a medical system, comprising:
[0028] In response to a user's clicking operation on a display interface of the client, determining a medical image to be identified;
[0029] Sending the medical image to be identified to the server of the medical system, and receiving a target object returned by the server, wherein the target object is a lymph node identification result output after performing lymph node identification processing on the medical image to be identified through a lymph node identification model, the lymph node identification model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from the multiple sample candidate lymph nodes through object type identification results and object position detection results of multiple sample candidate lymph nodes, the object position detection result is obtained by performing position detection on the object positions of the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing lymph node identification on the sample medical image;
[0030] The target object is presented to the user through the presentation interface.
[0031] According to a seventh aspect of an embodiment of the present disclosure, a method for detecting visible lymph nodes in a CT image is provided, comprising:
[0032] determining a CT image to be identified;
[0033] The CT image is input into a lymph node recognition model for lymph node recognition to obtain recognition results of visible lymph nodes in the CT image, wherein the lymph node recognition model is obtained by training sample target lymph nodes identified from CT sample images and sample labels of the CT sample images, the sample target lymph nodes are determined from the multiple sample candidate lymph nodes through lymph node type recognition results and lymph node position detection results of multiple sample candidate lymph nodes, the lymph node position detection results are obtained by performing position detection on the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing lymph node recognition on the CT sample image.
[0034] According to an eighth aspect of the embodiments of the present disclosure, a computer-aided diagnosis method is provided, comprising:
[0035] determining a medical image to be identified;
[0036] Inputting the medical image to be identified into a lymph node recognition model to perform lymph node recognition, and obtaining a lymph node recognition result in the medical image to be identified, wherein the lymph node recognition model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from the multiple sample candidate lymph nodes through object type recognition results and object position detection results of multiple sample candidate lymph nodes, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing object recognition on the sample medical image;
[0037] The diagnosis result of the lymph node lesion is determined according to the lymph node recognition result in the medical image to be recognized.
[0038] According to a ninth aspect of an embodiment of the present disclosure, there is provided a computing device, including:
[0039] memory and processor;
[0040] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, lesion recognition method in lymph node CT images, or computer-aided diagnosis method are implemented.
[0041] According to the tenth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, lesion recognition method in lymph node CT images, or computer-aided diagnosis method are implemented.
[0042] According to the eleventh aspect of an embodiment of the present disclosure, a computer program product is provided, wherein, when the computer program product is executed in a computer, the computer is caused to execute the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, lesion recognition method in lymph node CT images, or computer-aided diagnosis method.
[0043] One or more embodiments of the present disclosure provide a target object recognition method, comprising: determining an image to be recognized; inputting the image to be recognized into an object recognition model for object recognition, and obtaining a target object in the image to be recognized, wherein the object recognition model is obtained by training sample target objects identified from sample images and sample labels of the sample images, the sample target objects are determined from the multiple sample candidate objects through object type recognition results and object position detection results of multiple sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects, and the multiple sample candidate objects are obtained by performing object recognition on the sample images.
[0044] Specifically, in the target object recognition method provided by the present disclosure, the object recognition model needs to determine multiple sample candidate objects in the sample image, as well as the object type recognition results and object position detection results of the multiple sample candidate objects during the training process, wherein the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. Then, model training is performed based on the sample labels and the sample target objects determined by the object type recognition results and the object position detection results, thereby obtaining an object recognition model that accurately identifies the position of the target object in the image to be recognized. Based on this, when the image to be recognized is input into the object recognition model for object recognition, the target object can be identified with accurate position, avoiding the problem of inaccurate target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] FIG1 is a schematic diagram of a lymph node detection method according to an embodiment of the present disclosure;
[0046] FIG2 is a schematic diagram of an application of a target object recognition method provided by an embodiment of the present disclosure;
[0047] FIG3 is a flow chart of a target object recognition method provided by one embodiment of the present disclosure;
[0048] FIG4 is a flowchart of an object recognition model training method provided by one embodiment of the present disclosure;
[0049] FIG5 is a flowchart of a processing process of an object recognition model training method provided by one embodiment of the present disclosure;
[0050] FIG6 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0051] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0052] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0053] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0054] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0055] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model (Foundation Model), which is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization capabilities, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0056] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0057] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0058] Lymph Node: Abbreviated as LN, or LNs, refers to lymph nodes.
[0059] Transformer: A neural network architecture.
[0060] Query: A component of the Transformer computing process, which can be translated as query.
[0061] Contrastive representation: Contrastive learning representation.
[0062] CT: (Computed Tomography), also known as electronic computer tomography, uses precisely collimated X-ray beams, gamma rays, ultrasound, etc., together with highly sensitive detectors to perform cross-sectional scans one after another around a certain part of the human body.
[0063] CNN: Convolutional Neural Networks (CNN for short).
[0064] FCN; Fully Convolutional Networks (FCN for short).
[0065] Mask R-CNN: (Mask Region-based Convolutional Neural Network) is a deep learning model for object detection and instance segmentation.
[0066] DINO: A vision model.
[0067] R-CNN: A neural network model for object detection.
[0068] Flatten: One-dimensional processing.
[0069] ResNet: (Deep Residual Network).
[0070] With the continuous advancement of computer technology, neural network models are widely used in various service scenarios. In image processing scenarios, neural network models can be used to identify specific objects in images, thereby providing services for image processing. However, neural network models cannot accurately identify the location of specific objects in images. For example, in the lymph node detection scenario, neural network models are needed to detect lymph nodes from CT images. Specifically, computer-aided detection (CADe) is an active research area in medical imaging, rapidly developing with the development of deep learning technology. Within CADe tasks, lymph node (LN) identification is a critical yet understudied problem, playing an important role in daily clinical practice in both radiology and oncology. As an important component of the human immune system, lymph nodes are widely distributed throughout the body and serve as a major pathway for tumor spread. LN assessment is typically based on 3D computed tomography (CT) scans. Therefore, accurately identifying clinically significant LNs on CT is crucial for cancer diagnosis, staging, treatment planning, and prognosis assessment.
[0071] Lymph node (LN) assessment is a critical, indispensable, yet challenging task in daily clinical practice in radiology and oncology. Accurate LN analysis is crucial for cancer diagnosis, staging, and treatment planning. Even for experienced physicians, identifying scattered, low-contrast, clinically relevant LNs in 3D CT can be difficult, and interobserver variability is substantial. Because many adjacent anatomical structures have similar intensity, shape, or texture (blood vessels, muscle, esophagus, etc.), previous work on automated LN detection often produces high false positives (FPs). LN detection in CT is a challenging task for physicians for the following reasons. First, given the difficulty in distinguishing LNs' intensity relative to adjacent soft tissues, the relative contrast between LNs and adjacent anatomical structures is very low. Second, in addition to intensity, LNs also exhibit similar size and shape (spherical or ellipsoidal) to nearby soft tissues. These similarities make LNs easily confused with blood vessels, muscle, esophagus, pericardial recess, and other structures, as shown in Figure 1, which is a schematic diagram of a lymph node detection method provided by one embodiment of the present disclosure. Therefore, even experienced physicians may miss or misidentify LNs. Furthermore, LNs are dispersed throughout various regions of the body, such as the neck, umbilicus, chest, and abdomen. Therefore, manual inspection of hundreds of 2D CT slices from each patient's CT scan can easily lead to misidentification of clinically significant LNs, especially under time constraints.
[0072] To address the above issues, the present disclosure provides several types of solutions. The first type of solution is a statistical learning solution that uses hand-crafted features or CNN-based solutions to study automatic LN detection. Specifically, the statistical learning solution uses hand-crafted image features such as shape, spatial priors, and volume directional difference filters to capture the appearance of LNs and locate them. By applying FCN or Mask R-CNN to directly segment or detect LNs, the CNN-based solution achieves the function of lymph node detection. However, this first type of solution has many defects. First, these works only detect lymph node enlargement (short axis ≥ 10mm), although studies have shown that simple lymph node enlargement is not a reliable predictor of lymph node malignancy in cancer patients, with a sensitivity of only 60%-80%. Second, there are CNN-based works that attempt to use lymph node station priors (usually unavailable in clinical practice) to detect enlarged and smaller lymph nodes. However, their performance is low, and the average recall rate of this solution is <60%. Third, another limitation of these schemes is that they usually focus on a single body region, such as the chest or abdomen, and lack a universal LN detection model covering major body parts.
[0073] The second category of solutions is based on visual transformers, which formulate object detection as a set prediction task and assign labels via bipartite graph matching. Compared to the first category of CNN-based detectors, this second category achieves lymph node detection performance by improving the denoising training process and utilizing hybrid query selection for anchor initialization. They further extend their performance by adding a mask prediction branch to support segmentation tasks. However, this second category of solutions also has drawbacks, such as being unable to accurately detect the exact lymph nodes in images, and failing to address the high false positive rate problem mentioned above for automatic LN detection.
[0074] The third category of solutions is universal lesion detection. These solutions employ CNN-based detectors, such as R-CNN, for pulmonary nodule detection. Because 3D contextual information in adjacent axial slices is important for distinguishing lesions from other anatomically similar structures, some work has adopted 2.5D approaches by using 2D network architectures with multi-slice inputs. Compared to direct 3D detectors, 2.5D methods with pre-trained 2D model weights have demonstrated superior accuracy and faster runtime. One solution in this third category, the Multi-Task Universal Lesion Analysis Network (MULAN), based on the Mask R-CNN architecture, achieved superior performance on a large lesion dataset (i.e., DeepDisease). Another solution in this third category, the Lesion Sequence (LENS), further improves on MULAN with a novel anchor-free proposal network and a multi-dataset learning strategy. Furthermore, a slice attention transformer module was designed and inserted into the CNN-based detector to incorporate 3D information. However, this second-category solution also suffers from the inability to accurately detect the correct lymph nodes from images, failing to address the high false positive rate problem mentioned above for automatic LN detection. Moreover, the third type of research is a type of research on lesion detection without a Transformer-based backbone.
[0075] The fourth category of approaches utilizes model-based or statistical learning-based approaches. While automatic LN detection has been under development for a long time, it has primarily focused on extracting effective LN features, incorporating organ priors, or leveraging advanced learning models. Therefore, model-based or statistical learning-based approaches are often employed. Based on this, the fourth category of approaches utilizes model-based or statistical learning-based approaches. This approach improves on Mask R-CNN by proposing a global-local attention module and a multi-task uncertainty loss to detect LNs in abdominal MR images. Furthermore, CNN-based LN detection and segmentation models have also been explored. However, these segmentation methods often require additional labels (such as organ or LN delineation) or imaging modalities, which limits their applicability in clinical applications. Furthermore, directly segmenting LNs can yield poor performance because LNs are small and dispersed objects, making instance-based objective learning difficult to supervise using voxel-wise segmentation losses. Previous studies have typically focused on a single body region or disease, or have only detected enlarged lymph nodes (≥10 mm), neglecting smaller, clinically important metastatic lymph nodes.
[0076] Based on this, the present disclosure provides a target object recognition method. The present disclosure also relates to an object recognition model training method, a lesion recognition method in liver CT images, a target object recognition device, an object recognition model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0077] Referring to FIG2 , FIG2 shows an application diagram of a target object recognition method provided according to an embodiment of the present disclosure. Based on FIG2 , it can be seen that a user can send a patient CT image to a server 104 via a terminal 102 . The target object recognition method is applied to the server 104 , which can be understood as a server or cloud. The specific lymph node recognition process is as follows: First, the server 104 inputs the patient CT image into a lymph node recognition model (i.e., an object recognition model) for lymph node recognition. The image processing module in the lymph node recognition model processes the patient CT image to obtain an image feature vector, which is then input into the Transformer encoder of the lymph node recognition model. Second, the Transformer encoder processes the image feature vector to obtain a lymph node recognition result corresponding to the patient CT image. The lymph node recognition result includes a lymph node location box, a classification score, a lymph node image, and an IoU (Intersection over Union) score. The IoU score and the classification score are used to select a more accurate lymph node recognition result from multiple lymph node recognition results and input it into the Transformer decoder. Finally, the Transformer decoder of the lymph node recognition model decodes the relatively accurate lymph node recognition results and outputs multiple lymph node recognition results with higher accuracy, thereby identifying the lymph nodes in the correct position and avoiding the problem of inaccurate lymph node recognition.
[0078] The target object recognition method provided by the present disclosure can be applied to detect target objects in various medical images, such as the detection of visible lymph nodes in CT images, the detection of lung nodules in CT images, and the detection of prostates in MRI images.
[0079] Referring to FIG3 , FIG3 shows a flow chart of a target object recognition method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0080] Step 302: Determine an image to be recognized.
[0081] Among them, the image to be identified can be understood as an image that requires an object recognition model to perform object recognition processing, and the image to be identified contains a target object. For example, the image to be identified can be a CT image, a magnetic resonance imaging (MRI) image, or an image obtained using other medical imaging technologies. For example, a 3D CT image of a person, or a 3D CT image of an animal. The image to be identified can specifically be a medical image of different body parts of a person or an animal (such as the neck, chest, and upper abdomen, etc.). In one or more embodiments provided in the present disclosure, the image to be identified can be a 3D CT image of a patient obtained by scanning.
[0082] In one or more embodiments provided herein, a new LN DE Detection Transformer algorithm model, referred to as LN-DETR (based on the state-of-the-art Transformer detection / segmentation framework Mask DINO), is proposed to achieve more accurate recognition performance. Furthermore, this model incorporates 3D context by enhancing the 2D backbone using multi-scale 2.5D feature fusion. In other words, in the embodiments disclosed herein, an end-to-end LN detection Transformer, referred to as the LN-DETR model, is proposed to address the challenging general LN detection problem. It should be noted that LN-DETR includes key components such as improved denoising training and hybrid query selection for anchor initialization. The LN-DETR model is further improved based on the following key observation: CT is volumetric data and has important 3D context for LN recognition. However, pure 3D DETR is computationally expensive and cannot leverage pre-trained weights, which is crucial for achieving high performance with Transformer models. Therefore, the object recognition method provided in this disclosure enhances LN-DETR by utilizing an effective multi-scale 2.5D fusion scheme and simultaneously utilizes pre-trained 2D weights to incorporate 3D context. The specific method is as follows.
[0083] The step of determining the image to be identified includes:
[0084] Determining image data to be identified, and performing image segmentation on the image data to be identified to obtain a plurality of segmented images;
[0085] determining a target segmented image and other segmented images except the target segmented image from the plurality of segmented images;
[0086] Constructing a target set of images to be identified based on the target segmented image, and constructing other sets of images to be identified based on the other segmented images;
[0087] The target set of images to be recognized and the other sets of images to be recognized are used as images to be recognized.
[0088] The image data to be identified can be a 3D CT image of the patient, and the segmented image can be understood as a two-dimensional slice of the 3D CT image. The target segmented image can be understood as a target CT slice, which is an image slice (a cross section) closest to the location of the lymph nodes (LNs) in the 3D CT scan image.
[0089] Other segmented images can be understood as other slices in the 3D CT image except the target segmented image. For example, with the target CT slice as the center, four upper slices and four lower slices are extracted from the original CT scan image as the three-dimensional background of the central target CT slice. The three-dimensional background can be other segmented images.
[0090] The following example illustrates the application of the target object recognition method provided by this disclosure in a whole-body lymph node detection scenario based on a two-stage Transformer framework. The image data to be identified is a 3D CT image. To detect lymph nodes (LNs) in 3D CT images, multiple slices must be extracted from the original CT scan. The specific steps are as follows.
[0091] First, the target CT slice is identified from the original CT scan. The original CT scan is a 3D CT scan of a patient's body. The target CT slice is an image slice (a cross section) within the 3D CT scan image that approximates the location of lymph nodes (LNs). Next, four upper slices and four lower slices are extracted from the original CT scan, centered on the target CT slice, as a 3D background for the central target CT slice. To detect LNs in the target CT slice, these nine consecutive CT slices are divided into three sets of three-channel images, each independently processed by a shared CNN backbone. In other words, the nine consecutive CT slices are divided into three sets of three-channel images, or three slice sets. For example, slice set 1 contains four upper slices, slice set 2 contains four lower slices, and slice set 3 contains one target CT slice. This facilitates the subsequent accurate identification of the target object (i.e., lymph node) in the image to be identified.
[0092] Step 304: Input the image to be identified into an object recognition model to perform object recognition, and obtain a target object in the image to be identified.
[0093] In which, the object recognition model is obtained by training the sample target objects identified from the sample images and the sample labels of the sample images. The sample target objects are determined from the multiple sample candidate objects through the object type recognition results and object position detection results of multiple sample candidate objects. The object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. The multiple sample candidate objects are obtained by performing object recognition on the sample images.
[0094] The object recognition model can be understood as a model capable of identifying a target object contained in an image. The object recognition model can be a whole-body lymph node detection model based on a two-stage Transformer framework. For example, the object recognition model can be the LN-DETR model described in the above embodiment. In one or more embodiments provided herein, the object recognition model can be a Transformer model.
[0095] The target object can be understood as the object contained in the image to be identified and output by the object recognition model. For example, the target object can be an object such as a lymph node, a lung nodule, etc., and can be set according to the actual application scenario. In one or more embodiments provided in the present disclosure, the target object can be composed of an object position detection result, an object position detection result, an object position of the target object, and an object image of the target object segmented from the image to be identified (which can be represented by a feature vector). Through the object position detection result, the object position detection result, the object position of the target object, and the object image of the target object, a target object can be accurately determined from the image to be identified. For example, the target object can be a lymph node query result output by the object recognition model, and the lymph node query result is composed of a lymph node classification score, a lymph node position box, a lymph node image (a lymph node image segmented from a CT image, which can be represented by a feature vector) and an IoU score of the lymph node.
[0096] The sample image can be understood as an image containing a sample target object as a sample, for example, a 3D CT image of a patient as a sample, and the sample target object can be a lymph node. In one or more embodiments provided in the present disclosure, the object recognition model in the target object recognition method provided in the present disclosure can be a Transformer-based detector. Large-scale data can be used in the training process of the object recognition model. For example, 3D CT scan images of multiple patients, involving different body parts and diseases, can be used as training sample images; and the lymph node marks marked on the 3D CT scan image can be used as sample labels, for example, more than 10,000 lymph nodes with magnified and reduced marks are marked on the 3D CT scan image. Thus, the object recognition model is trained and tested in combination with seven lymph node datasets of different body parts (neck, chest and abdomen) and pathologies to solve the challenging but clinically important LN detection task.
[0097] Sample candidate objects can be understood as objects detected by the object recognition model during model training. Subsequently, the more credible sample candidate objects need to be selected from these sample candidate objects as sample target objects. These sample target objects are the more credible sample candidate objects among multiple sample candidate objects. For example, the sample candidate objects can be multiple lymph nodes detected by the object recognition model; the sample target object can be the more credible lymph node among these multiple lymph nodes.
[0098] The object location detection result can be understood as a detection result that indicates whether the object location of the sample candidate object is accurately identified, such as an Intersection over Union (IoU) score. For example, in one or more embodiments provided herein, unlike objects in natural images, which typically have distinct edges, in CT scans, LN boundaries often differ slightly from adjacent anatomical structures, which also exhibit similar intensity, shape, or texture. Consequently, the detector can also generate a large number of false positives (FPs) or duplicate predictions near true LNs. To address this issue, the target object recognition method provided herein introduces two key advances in LN-DETR, designed to improve query embedding quality. This improves localization accuracy and better distinguishes true LNs from other similar anatomical structures (which appear as FPs or duplicate predictions). Specifically, to improve the representation quality of LN queries, given that LN boundaries are often unclear, an Intersection over Union (IoU) prediction head and IoU-guided query selection are introduced to select LN queries with higher localization accuracy as decoder query initialization. The IoU prediction head can detect LN boundaries and provide corresponding IoU scores (which can be understood as localization confidence). We estimate the localization confidence of query results by adding an IoU prediction task. We also propose an IoU-guided query selection to select LN queries (i.e., sample target objects) with higher localization confidence from multiple LN queries (i.e., sample candidate objects). These selected LN queries serve as the initial content queries and initial anchors for the Transformer decoder, ensuring more accurate initialization.
[0099] The object type recognition result can be understood as a detection result that characterizes whether the sample candidate object is the sample target object. For example, the object type recognition result can be a classification score corresponding to the sample candidate object. The larger the classification score, the more likely the sample candidate object is the sample target object. In the case where the sample candidate object is a plurality of candidate lymph nodes to be identified, the larger the classification score of the candidate lymph node, the more likely the candidate lymph node is a true lymph node (i.e., the sample target object). The object position can be understood as the position of the sample candidate object in the image to be identified.
[0100] In one or more embodiments provided in the present disclosure, inputting the image to be identified into the object recognition model for object recognition to obtain the target object in the image to be identified includes steps 1 and 2:
[0101] Step 1: Input the image to be identified into an object recognition model, use the object recognition model to determine multiple candidate objects in the image to be identified, and confirm the object type recognition result and object position detection result of each candidate object.
[0102] Specifically, in one or more embodiments provided by the present disclosure, determining multiple candidate objects in the image to be recognized using the object recognition model and confirming the object type recognition result and object position detection result of each candidate object includes:
[0103] Determining the image features to be recognized corresponding to the image to be recognized using the object recognition model;
[0104] Performing object recognition on the features of the image to be recognized, determining a plurality of candidate objects in the image to be recognized, and object type recognition results of the plurality of candidate objects;
[0105] Perform position detection on the object positions of the multiple candidate objects to obtain an object position detection result corresponding to each candidate object.
[0106] The image features to be identified may be understood as vector features corresponding to the image to be identified.
[0107] Candidate objects can be understood as objects detected by the object recognition model. Subsequently, the more reliable candidate objects need to be selected as target objects. For example, the candidate objects can be multiple lymph nodes detected by the object recognition model; the target object can be a lymph node with a relatively accurate location and high reliability among the multiple lymph nodes.
[0108] The object type recognition result can be understood as a detection result indicating whether the candidate object is the target object. For example, the object type recognition result can be a classification score corresponding to the candidate object. The higher the classification score, the more likely the candidate object is the target object. In the case where the candidate object is multiple candidate lymph nodes, the higher the classification score of the candidate lymph node, the more likely the candidate lymph node is a true lymph node (i.e., the target object).
[0109] The object position detection result can be understood as a detection result that represents whether the object position of the candidate object is accurately identified. For example, the object position detection result is an IoU score.
[0110] Specifically, the target object recognition method provided by the present disclosure can input the image to be recognized into the object recognition model, and the object recognition model will first perform feature extraction on the image to be recognized, thereby determining the image features to be recognized corresponding to the image to be recognized; secondly, it will perform object recognition on the image features to determine multiple candidate objects in the image to be recognized, and perform object classification recognition on the multiple candidate objects, thereby determining the object type recognition results of the multiple candidate objects; finally, it will determine the object positions of the multiple candidate objects, and perform position detection on the multiple candidate objects based on the object positions, thereby obtaining the object position detection results corresponding to each candidate object. This facilitates the subsequent selection of the target object with accurate position from the multiple candidate objects based on the object type recognition results and the object position detection results.
[0111] Continuing with the above example, the candidate objects are candidate lymph nodes, the object type recognition result is the classification score, and the object location detection result is the Intersection over Union (IoU) score. Based on this, after obtaining the slice set, the image features corresponding to the slice set are input into the encoding layer contained in the Transformer encoder. The feature vector output by the Transformer encoding layer is input into the prediction head contained in the Transformer encoder. This prediction head is responsible for predicting multiple lymph node recognition results (i.e., lymph node query results) in the CT slice image. The lymph node query result includes the lymph node (LNs) location box, IoU score, classification score, and lymph node image.
[0112] In one or more embodiments provided in the present disclosure, determining the image features to be recognized corresponding to the image to be recognized using the object recognition model includes:
[0113] Processing the image to be recognized using the image processing module in the object recognition model to obtain candidate image features;
[0114] The encoding module in the object recognition model is used to encode the candidate image features to obtain the image features to be recognized corresponding to the image to be recognized.
[0115] The image processing module can be understood as the module used for feature extraction in the object recognition model. The image processing module can be a CNN backbone (or CNN backbone network). In one or more embodiments provided herein, the LN-DETR framework consists of a CNN backbone network with multi-scale 2.5D feature fusion, as well as a Transformer encoder and decoder. The CNN backbone network is used to perform multi-scale 2.5D feature fusion. Although the 3D environment is crucial for LN detection, the performance of 3D detectors is generally inferior to 2D models initialized with weights pre-trained using large-scale data. To bridge this gap, the target object recognition method provided herein applies a 2.5D feature-level fusion scheme to LN-DETR. Unlike using only the output of the FPN for prediction, 2.5D fused features from multiple levels are fed into the Transformer encoder to avoid information loss and enable cross-scale token interaction in the encoder layer. Specifically, to detect LNs in the target CT slice, after the nine consecutive CT slices in the above embodiment are divided into three groups of three-channel images, each group is independently processed by a shared CNN backbone to obtain fused image features (i.e., candidate image features).
[0116] The encoding module can be understood as a module in the object recognition model used to encode features, for example, the encoding layer of the object recognition model.
[0117] Continuing with the above example, to detect lymph nodes (LNs) in a target CT slice, multiple slices are extracted from the original CT scan. These consecutive CT slices are then divided into three groups of three-channel images, or three slice sets. These slice sets are then input as training sample images into the LN-DETR model. The CNN backbone in the LN-DETR model performs feature fusion on these slice sets. Each slice set is independently extracted using the shared CNN backbone to obtain candidate image features. After obtaining the candidate image features, the image features of the slice set are converted to one dimension using a flatten layer. The feature vector of the slice image after this one-dimensional processing is added to the position embedding corresponding to the slice feature vector to obtain 2.5D image tokens (which can be understood as a feature vector). These 2.5D image tokens are then input into the Transformer encoder layer, which outputs the feature vector (i.e., the image feature to be identified). This facilitates the subsequent Transformer encoder to accurately identify the target object.
[0118] In one or more embodiments provided in the present disclosure, the processing of the image to be recognized by the image processing module in the object recognition model to obtain candidate image features includes:
[0119] Using the image processing module in the object recognition model, feature extraction is performed on multiple sets of images to be recognized to obtain image features corresponding to each set of images to be recognized;
[0120] Performing feature concatenation processing and feature conversion processing on the image features corresponding to the plurality of sets of images to be identified to obtain target image features;
[0121] Determining a target image set to be identified from the multiple image sets to be identified, and replacing image features corresponding to the target image set to be identified with the target image features;
[0122] The candidate image features are obtained based on the image features corresponding to the multiple sets of images to be identified.
[0123] Continuing with the above example, the CNN backbone network in the LN-DETR performs feature fusion processing on the slice set. Specifically, first, the three slice sets are subjected to feature extraction through concatenation and 1×1 transformation to obtain the image features (i.e., feature maps) corresponding to each slice set. Secondly, the image features corresponding to each set are fused to obtain the fused image features. Finally, the image features of the original central target slice in the slice set are replaced with the fused image features, while the image features of the other two sets remain unchanged. By determining the image features corresponding to each slice set as candidate image features, the purpose of quickly obtaining candidate image features is achieved. It should be noted that this fusion operation is applied to all four ResNet blocks in the CNN backbone network.
[0124] In one or more embodiments provided herein, the target object recognition method provided herein uses a CNN backbone network to process CT images. The CNN backbone network in this disclosure proposes a DETR (i.e., DINO) with an improved denoising anchor box, which can be optimized end-to-end. Its enhanced version of denoising training is an extension of DN-DETR. Mask DINO further extends DINO by adding a mask prediction branch that supports different segmentation tasks. Therefore, it contributes to improving the performance of challenging LN detection.
[0125] In one or more embodiments provided in the present disclosure, performing object recognition on the features of the image to be recognized, determining multiple candidate objects in the image to be recognized, and object type recognition results of the multiple candidate objects include:
[0126] Using the object position recognition module in the object recognition model to perform object position recognition on the image features to be recognized, and determine the object positions of multiple candidate objects in the image to be recognized;
[0127] Inputting the object positions of the multiple candidate objects and the features of the image to be identified into the object recognition module in the object recognition model to perform object recognition, thereby obtaining multiple candidate objects in the image to be identified;
[0128] The object classification module in the object recognition model is used to perform object type recognition on each candidate object, and an object type recognition score corresponding to each candidate object is determined.
[0129] It should be noted that the Transformer encoder includes a prediction module (prediction heads), which consists of an IoU prediction head (consisting of a single MLP layer), a classification head, a bounding box regression head, and a mask generation head. These heads operate in parallel. The object location recognition module can be understood as a bounding box regression head: it locates the location of lymph nodes using bounding boxes, thereby determining the location box of the lymph nodes. The object classification module can be understood as a classification head: it classifies objects in the image and identifies lymph nodes in the image. The object recognition module can be understood as a mask generation head: it performs image segmentation to obtain object images of candidate objects, such as lymph node images. The object location detection module can be understood as an IoU prediction head: it uses IoU scores to estimate the query localization quality and uses IoU confidence to guide query ranking and selection in the final encoder layer.
[0130] Continuing with the above example, after obtaining the feature vector output by the encoding layer, the feature vector needs to be input into the prediction head contained in the Transformer encoder, and the bounding box regression head is used to predict the lymph node location box based on the feature vector, thereby locating the position of the lymph node through the location box, thereby determining the location box of the lymph node. The mask generation head is used to segment the lymph nodes based on the location box and feature vector of the lymph node, thereby extracting the lymph nodes from the CT image. The classification head classifies the lymph nodes based on the feature vector, thereby classifying and evaluating the lymph nodes in the image, and obtaining the classification score of each lymph node in the CT image. The specific process can be:
[0131] For the bounding box regression head: The feature vector corresponding to the CT slice image is input into the bounding box regression head. Based on the feature vector, the bounding box regression head can predict the lymph node location box (also called lymph node bounding box) and identify multiple lymph node location boxes in the CT slice image. Based on these location boxes, the location of the lymph node can be determined.
[0132] For the mask generation head, the feature vector corresponding to the CT slice image and the lymph node location box are input into the mask generation head. The mask generation head performs image segmentation on the lymph nodes framed by the lymph node location box to obtain the feature vector representation corresponding to the lymph node image.
[0133] For the classification head: The feature vector corresponding to the CT slice image and the lymph node location box are input into the classification head. The classification head evaluates the lymph node framed by the lymph node location box and determines the classification score for the lymph node. The higher the classification score, the higher the probability that the object framed by the lymph node location box is a lymph node.
[0134] By obtaining each lymph node location frame, lymph node image, and lymph node classification score (object type recognition score) in the CT image, it is convenient to accurately determine the target object from multiple candidate objects based on the object type recognition score.
[0135] In one or more embodiments provided by the present disclosure, performing position detection on the plurality of candidate objects to obtain an object position detection result corresponding to each candidate object includes:
[0136] The object positions of the multiple candidate objects and the image features to be identified are input into the object position detection module in the object recognition model for position detection to obtain an object position detection score corresponding to each candidate object.
[0137] Continuing with the above example, after obtaining the feature vector output by the encoding layer, the feature vector needs to be input into the prediction head included in the Transformer encoder, and the IoU prediction head is used to evaluate the accuracy of the lymph node position based on the feature vector, thereby obtaining the IoU prediction (positioning confidence). The specific execution process of the IoU prediction head is as follows: the feature vector corresponding to the CT slice image and the lymph node position box are input into the IoU prediction head, and the IoU prediction head is used to perform positioning analysis on the lymph nodes framed by the lymph node position box to obtain the IoU score (i.e., IoU confidence). The IoU score can be used to estimate the query positioning quality, and the IoU confidence can be used to guide the query sorting and selection of the last layer encoder. This facilitates the subsequent accurate determination of the target object from multiple candidate objects based on the object position detection score.
[0138] Step 2: Based on the object type recognition result and the object position detection result, determine the target object from the multiple candidate objects.
[0139] Specifically, in one or more embodiments provided in the present disclosure, the object position detection result is an object position detection score, and the object type recognition result is an object type recognition score;
[0140] The determining the target object from the plurality of candidate objects based on the object type recognition result and the object position detection result includes:
[0141] multiplying the object position detection score and the object type recognition score to obtain object recognition scores for the multiple candidate objects;
[0142] The target object is selected from the plurality of candidate objects based on the object recognition score.
[0143] Continuing with the previous example, we use the calculated IoU prediction (position confidence) to multiply the query's IoU score by its classification score to create a new query ranking score. Based on the new query ranking score (more accurate LNs bounding box) at the encoder output, we select the top K queries to identify the target object with accurate location, avoiding the problem of inaccurate target objects.
[0144] In one or more embodiments provided by the present disclosure, the overall framework of the LN-DETR proposed in the target object recognition method consists of a CNN backbone with multi-scale 2.5D feature fusion and a transformer encoder and decoder. During the model training process, it includes (1) an IoU-guided query ranking and selection module in the last layer of the encoder (based on an additional IoU prediction head); and (2) a query contrast learning module (naturally utilizing the denoised anchor boxes in the mask DINO) to improve the query representation ability for distinguishing true LN queries from nearby FP or repeated queries. Specifically, the overall architecture of the LN-DETR model adopts a multi-scale deformable attention module to aggregate multi-scale feature maps. In order to accelerate training convergence, additional denoised queries and denoising losses are adopted. In addition, it calculates the highest-scoring queries (used as region proposals) from the output of the transformer encoder to initialize content queries and reference boxes, which are fed back to the decoder for subsequent refinement. When selecting the top K queries, only classification confidence is considered as the ranking criterion. However, classification confidence does not necessarily correlate with the quality of the predicted bounding box; predicted LN bounding boxes with higher classification confidence actually overlap less with the corresponding ground-truth boxes. Therefore, these inaccurate initial reference boxes can lead to suboptimal predictions in the decoder. For example, an IoU prediction branch estimates localization confidence for subsequent box refinement. Unlike these works, the object recognition method proposed in this disclosure adds an IoU prediction head to both the Transformer encoder and decoder, uses IoU scores to estimate query localization quality, and leverages IoU confidence to guide query ranking and selection in the final encoder layer. Another observed issue is that the LN predictions of the original Mask DINO often contain many duplicates or false positives due to the fuzzy LN boundaries and similar adjacent anatomical structures. This is likely due to the one-to-one matching step of the Hungarian algorithm, which only assigns each ground-truth to its most closely matching query and forces all unmatched queries to predict the same background label, regardless of their relative ranking. This results in insufficient supervision to distinguish locally similar queries. To alleviate this problem, a simple yet effective query contrastive learning module (naturally leveraging the denoised anchor boxes in Mask DINO) is introduced to improve query representation capabilities to distinguish true LN queries from nearby FP or duplicate queries. Based on this, the training process for the object recognition model is as follows.
[0145] In one or more embodiments provided by the present disclosure, before inputting the image to be identified into the object recognition model for object recognition to obtain the target object in the image to be identified, the method further includes:
[0146] Determine a sample image of the object recognition model to be trained, and a sample label corresponding to the sample image;
[0147] Inputting the sample image into the object recognition model to be trained, performing object recognition on the sample image using the object recognition model to be trained, and obtaining a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object;
[0148] Obtaining an object position detection result for each sample candidate object by performing position detection on the object positions of the plurality of sample candidate objects;
[0149] determining a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result;
[0150] Based on the sample target object and the sample label, the object recognition model to be trained is trained to obtain the object recognition model.
[0151] For the above training process of the object recognition model to be trained, please refer to the steps of the object recognition model training method described below, which will not be described in detail here.
[0152] In the target object recognition method provided by the present disclosure, the object recognition model, during the training process, needs to determine multiple sample candidate objects in a sample image, as well as object type recognition results and object position detection results for the multiple sample candidate objects, wherein the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. Model training is then performed based on the sample target objects and sample labels determined by the object type recognition results and the object position detection results, thereby obtaining an object recognition model that accurately determines the position of the target object in the image to be recognized. Based on this, when the image to be recognized is input into the object recognition model for object recognition, the target object can be identified with accurate position, avoiding the problem of inaccurate target objects.
[0153] Referring to FIG4 , FIG4 shows a flowchart of an object recognition model training method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0154] Step 402: Determine sample images for the object recognition model to be trained and sample labels corresponding to the sample images.
[0155] Step 404: Input the sample image into the object recognition model to be trained, and use the object recognition model to be trained to perform object recognition on the sample image to obtain multiple sample candidate objects and object type recognition results corresponding to each sample candidate object.
[0156] In one or more embodiments provided in the present disclosure, performing object recognition on the sample image using the object recognition model to be trained to obtain multiple sample candidate objects and object type recognition results corresponding to each sample candidate object includes:
[0157] Determining sample image features corresponding to the sample image using the object recognition model to be trained;
[0158] Performing object position recognition on the sample image features using the object position recognition module in the object recognition model to be trained, and determining the object positions of a plurality of sample candidate objects in the sample image;
[0159] Inputting the object positions of the multiple sample candidate objects and the sample image features into the object recognition module in the object recognition model to be trained to perform object recognition, thereby obtaining multiple sample candidate objects in the sample image;
[0160] The object classification module in the object recognition model to be trained is used to perform object type recognition on each sample candidate object, and an object type recognition score corresponding to each sample candidate object is determined.
[0161] The following describes the application of the object recognition model training method provided by the present disclosure in the training of a LN-DETR model. The LN-DETR framework consists of a CNN backbone network with multi-scale 2.5D feature fusion and a Transformer encoder and decoder. The CNN backbone network is capable of multi-scale 2.5D feature fusion. Although the 3D environment is crucial for LN detection, the performance of 3D detectors is generally inferior to 2D models initialized with weights pre-trained with large-scale data. To bridge this gap, the 2.5D feature-level fusion scheme in
[15] is applied to LN-DETR. By feeding 2.5D fused features from multiple levels into the Transformer encoder, information loss is avoided and cross-scale token interaction in the encoder layer is enabled. Specifically, to detect LNs (i.e., sample target objects) in a sample CT slice (i.e., sample image), four upper and lower slices are extracted from the original CT scan as the 3D background of the central target slice. These nine consecutive CT slices are then divided into three groups of 3-channel images. Each group is independently processed by a shared CNN backbone, and the three groups are then fused through concatenation and 1×1 transformation. Afterwards, the feature map of the original central target patch is replaced by the fused feature map, and the feature maps of the upper and lower sets remain unchanged.
[0162] After obtaining the candidate image features, the image features of the slice set are converted to one dimension through the flatten layer; the feature vector of the slice image after one-dimensional processing and the position vector (Position Embedding) corresponding to the slice feature vector are added to obtain 2.5D image tokens (2.5D Image tokens, which can be understood as a feature vector); the 2.5D image tokens are input into the encoding layer of the Transformer encoder to obtain the feature vector output by the encoding layer (i.e., the sample image features).
[0163] The Transformer encoder includes a prediction module (prediction heads), which consists of an IoU prediction head (the IoU prediction head consists of a single MLP layer), a classification head, a bounding box regression head, and a mask generation head. After obtaining the feature vector output by the encoding layer, it is necessary to input the feature vector into the prediction head included in the Transformer encoder. The bounding box regression head predicts the lymph node location box based on the feature vector, thereby locating the lymph node using the location box and determining the lymph node location box. The mask generation head performs lymph node segmentation based on the lymph node location box and feature vector, thereby extracting lymph nodes from the sample CT image. The classification head classifies the lymph nodes based on the feature vector, thereby classifying and evaluating the lymph nodes in the image and obtaining a classification score for each lymph node in the CT image. Specifically, the bounding box regression head is fed with the feature vector corresponding to the sample CT slice image (i.e., the sample image features). Based on this feature vector, the bounding box regression head predicts the lymph node location box (also known as the lymph node bounding box), identifying multiple lymph node location boxes (i.e., the object locations of the sample candidate objects) in the sample CT slice image. Based on these location boxes, the lymph nodes can be located.
[0164] For the mask generation head, the feature vector corresponding to the sample CT slice image and the lymph node location box are input into the mask generation head. The mask generation head performs image segmentation on the lymph nodes framed by the lymph node location box to obtain the feature vector representation corresponding to the lymph node image (i.e., the sample candidate object).
[0165] For the classification head: The feature vector corresponding to the sample CT slice image and the lymph node location box are input into the classification head. The classification head can evaluate the lymph node framed by the lymph node location box and determine the classification score (i.e., object type recognition score) for the lymph node. The higher the classification score, the higher the probability that the object framed by the lymph node location box is a lymph node.
[0166] By obtaining each lymph node location frame, lymph node image, and lymph node classification score in the sample CT image, it is convenient to accurately determine the sample target object from multiple sample candidate objects based on the object type recognition score.
[0167] Step 406: Obtain an object position detection result for each of the sample candidate objects by performing position detection on the object positions of the multiple sample candidate objects.
[0168] In one or more embodiments provided by the present disclosure, performing position detection on the object positions of the multiple sample candidate objects to obtain the object position detection result of each sample candidate object includes:
[0169] The object positions of the multiple sample candidate objects and the sample image features are input into the object position detection module in the object recognition model to be trained for position detection, and the object position detection score corresponding to each sample candidate object is obtained.
[0170] Continuing with the previous example, to select queries with high classification and localization accuracy, the IoU-guided query selection module needs to be trained. First, an additional IoU prediction head is trained. This prediction head is introduced at the final layer of the Transformer encoder and all decoder layers with shared parameters to predict the query's localization accuracy (IoU). The IoU prediction head consists of a single MLP layer (similar to [ 1] ), running in parallel with the classification head, bounding box regression head, and mask generation head. Specifically, the IoU prediction head is implemented as follows: the feature vector corresponding to the sample CT slice image and the lymph node location box are input into the IoU prediction head. The IoU prediction head then performs a localization analysis on the lymph nodes framed by the lymph node location box to obtain an IoU score (i.e., IoU confidence). The IoU score can then be used to estimate the query localization quality, and the IoU confidence is used to guide query ranking and selection in the final encoder layer. This IoU prediction head is also applied to queries in each decoder layer to improve the accuracy of IoU prediction.
[0171] In one or more embodiments provided by the present disclosure, it is necessary to train the IoU prediction head. To train the IoU prediction head, the true IoU values between GTs and matching box predictions are used to supervise the IoU prediction. Given a CT slice (sample image) with M GT boxes, the M query features matched are represented as {q1, q2, ..., qM}, and the IoU prediction for each query is Assuming that Hungarian Matching assigns the i-th GT to the j-th query, the IoU score of the query box prediction to the matching GT box can be calculated, which is recorded as Based on this, the IoU score predicted by the IoU prediction head is obtained. The subsequent IoU loss of the i-th GT is defined as the following formula 1:
[0172] Step 408: Determine a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result.
[0173] Continuing with the previous example, using the calculated IoU prediction (localization confidence), a new query ranking score is used by multiplying the query's IoU score with its classification score. The top K queries (i.e., sample target objects) are selected based on the new query ranking scores (more accurate LNs bounding boxes) at the encoder output.
[0174] Step 410: Based on the sample target object and the sample label, the object recognition model to be trained is trained to obtain an object recognition model.
[0175] In one or more embodiments provided by the present disclosure, training the object recognition model to be trained based on the sample target object and the sample label to obtain the object recognition model includes:
[0176] Determine the sample candidate object label, object type recognition score label, object position label, object position detection score label, and sample target object label included in the sample label;
[0177] Determine a first loss value based on the object type recognition score and the object type recognition score label, determine a second loss value based on the object position and the object position label, determine a third loss value based on the sample candidate object label and the plurality of sample candidate objects, and determine a fourth loss value based on the object position detection score and the object position detection score label;
[0178] Determining a fifth loss value based on the sample target object and the sample target object label;
[0179] Based on the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value, the object recognition model to be trained is trained until a model training stop condition is reached to obtain an object recognition model.
[0180] Continuing with the above example, during the training phase, the total loss is the original loss, i.e., the loss value of the classification head, the loss value of the bounding box regression head, and the loss value of the mask head, as well as the combination of the IoU prediction loss and the query comparison loss proposed in one or more embodiments of the present disclosure. For details, please refer to the following formula 2.
[0181] in Refers to the loss value of the classification head, that is, the first loss value, Refers to the loss value of the bounding box regression head, that is, the second loss value, Refers to the loss value of the covered pier, that is, the third loss value, Refers to the IoU prediction loss, that is, the fourth loss value, Refers to the query contrast loss, or the fifth loss value. Loss values λ1 to λ5 represent weights for each loss component. These values remain the same as those in the original Mask DINO. Experimentally, we set λ4 to 10 and λ5 to 1. During inference, the query contrast branch is removed, and IoU-guided query selection is applied to the encoder and decoder outputs. The LN-DETR model is trained using this total loss until the model training stopping condition is met. This results in an object recognition model that accurately identifies the target object's corresponding location, avoiding inaccurate target object recognition.
[0182] In one or more embodiments provided by the present disclosure, determining the fifth loss value based on the sample target object and the sample target object label includes:
[0183] Performing noise processing on the sample target object label to obtain a plurality of noise object labels corresponding to the sample target object label, wherein the number of the plurality of noise object labels is consistent with the number of the plurality of sample target objects;
[0184] Determining corresponding noise object labels for a plurality of sample target objects, and calculating associations between each sample target object and the noise object label corresponding to each sample target object to obtain a plurality of object association groups;
[0185] calculating an association score between a sample target object and a noise object label in each object association group, and dividing the plurality of object association groups into a first object association group and a second object association group based on the association score;
[0186] A fifth loss value is determined based on a similarity between the first object association group and the second object association group.
[0187] The sample target object label can be understood as a real object corresponding to the sample target object, serving as a sample label. For example, the sample target object label can be a real lymph node image in the sample image.
[0188] The noise object label can be understood as the sample target object label after noise processing, wherein the noise processing can be understood as adding noise and removing noise.
[0189] The association score can be understood as a score that characterizes the degree of association between the sample target object and the noise object label. The higher the score, the more associated the sample target object and the noise object label.
[0190] The similarity can be understood as a value representing the degree of similarity between the first object association group and the second object association group. The higher the value, the more similar the first object association group and the second object association group are.
[0191] It should be noted that to improve the representation quality of LN queries, an IoU prediction head and IoU-guided query selection are introduced to select LN queries with higher localization accuracy as decoder query initialization. In addition, to reduce FP, query contrastive learning is proposed, with the goal of strengthening lymph node queries that match the true target query (obtained from the denoised anchor box) relative to unmatched queries. By introducing a query contrastive learning module at the decoder output, this module explicitly strengthens positive queries to their most closely matched ground truth (GT) query (derived from denoising training in Mask DINO) rather than mismatched negative query predictions. As can be seen from the study, these two components effectively improve LN detection performance.
[0192] Continuing with the previous example, query contrastive learning is introduced during model training. To better distinguish LNs from similar adjacent anatomical structures or repeated predictions at the feature level, a query contrastive learning module is introduced at the decoder output. The construction and calculation process of positive and negative query pairs are shown below.
[0193] Specifically, the K query results are used as initial contents, and the predicted location boxes are used as initial proposals, which are input into the decoding layer of the Transformer decoder. At the same time, the GT vectors (GT embeddings) and the corresponding GT boxes + noise (GT boxes noise) are input into the decoding layer. The noisy GT vectors are denoised to obtain multiple denoised GT queries (multiple noisy object labels). The GT query consists of the GT vector and the GT ground truth box (i.e., label).
[0194] In this scheme, each GT and its most matching output query form a positive query pair, while all other non-matching queries for that GT are treated as negative queries. In this setting, the GT needs to have its query representation, for example, it can be derived by processing the GT box and embedding the labels through the Transformer decoder. It should be noted that an intuitive way to construct positive and negative query pairs is to use Hungarian matching results: each GT and its most matching output query form a positive query pair, while all other non-matching queries for that GT are treated as negative queries. In this setting, the GT needs to have its query representation, for example, it can be derived by processing the GT box and embedding the labels through the Transformer decoder. GT queries can be easily obtained through the denoising training process in Mask DINO. Therefore, using multiple sets of denoising training, multiple GT queries corresponding to the same GT can be obtained to form multiple positive pairs, which encourages more divergence and robustness, further facilitating contrastive learning.
[0195] Specifically, assume that there are N denoising groups, each group contains M denoised GT queries, namely anchor queries, so the total anchor queries q DN It is shown in the following formula 3:
[0196] Where M is the number of GT LN boxes in the CT slice. Match the K query results and the corresponding predicted position boxes with the GT query in the denoising group. Assuming that the matching cost (i.e., association score) between the K-th query result and the i-th GT in the CT slice is the smallest, then q k For query with N anchor points The matching pair is the positive query pair, and the other unmatched K-1 query results are Negative query pairs.
[0197] Afterwards, a comparative calculation can be performed on the positive query pairs and the negative query pairs. Through the comparative calculation, the similarity between all positive anchor pairs (positive query pairs) and negative anchor pairs (negative query pairs) is calculated, and the loss value is determined based on the similarity. Specifically, this scheme does not directly measure the similarity between different query embeddings. Instead, a simple shared MLP layer φ is first used to project the query embedding into the latent space. By projecting the query embedding, the similarity between all positive anchor and negative anchor pairs is calculated. The InfoNCE loss (i.e., query contrast loss) is then used to bring the matched positive queries closer to their designated anchor queries, while moving away from all other unmatched negative queries. The total query contrast loss formula for a given CT slice is shown in the following formula 4.
[0198] where <·,·> is the cosine similarity and τ is the temperature coefficient, which is set to 0.05. This contrast loss is applied to the output query of the last decoder layer.
[0199] In one or more embodiments provided herein, after model training is completed, the trained object recognition model is evaluated. Specifically, LN-DETR is first evaluated on the LN detection task. Seven LN datasets from different body parts and diseases were collected, including multiple patients and more than 10,000 labeled LN instances. Five of these datasets were used for model development and internal testing, and the remaining two datasets were used for independent external testing. CNN-based and Transformer-based detection and segmentation methods were compared. To further demonstrate the effectiveness of LN-DETR, the DeepLesion dataset was used for training and evaluation results. For details on the datasets and evaluation metrics for this object recognition model, please refer to the description below. Regarding the LN dataset: Seven LN datasets were collected and organized, including multiple patients, with more than 10,000 annotated LNs, covering different body parts (neck, chest, and upper abdomen), and different diseases (head and neck cancer, cancer, cancer, COVID, and other diseases). See Table 1 for details.
[0200] Table 1
[0201] The dataset column in Table 1 refers to the data source that provides the seven LN detection datasets. "NIH-LN" in this dataset column is a public dataset; Center 1, Center 2, Center 3, Center 4, and Center 5 refer to five different clinical centers, indicating that the datasets may be internal datasets collected from five different clinical centers. HN, Eso, and Mul represent head and neck cancer, cancer, and multiple diseases, respectively. Therefore, "Center 1 HN" and "Center 5 HN" in the dataset column can refer to head and neck cancer data provided by Clinical Center 1 and Clinical Center 5; "Center 2 Eso" and "Center 3 Eso" can refer to cancer data provided by Clinical Center 2 and Clinical Center 3; "Center 2 Lung" refers to lung disease (e.g., lung cancer) data provided by Clinical Center 2; and "Center 4 Mul" refers to multiple disease data provided by Clinical Center 4.
[0202] It should be noted that the eighth row (ie, the last row) in Table 1 is used to count the total number of values. For example, "all" in the eighth row is used to represent the total number of values of all data sources.
[0203] The #Patient column in Table 1 indicates the number of patients. In other words, it shows the number of patients from which each data source provided data. For example, the value "89" in the #Patient column in the first row of Table 1 indicates that the data from the "NIH-LN" public dataset came from 89 patients. Similarly, the other values in the #Patient column indicate the number of patients from which the data provided by each clinical center came.
[0204] It should be noted that the “1067” in the eighth row of Table 1 indicates that the data provided by all data sources come from 1067 patients.
[0205] The #LNs column in Table 1 indicates the number of lymph nodes (LNs) in each dataset. Specifically, the #LNs column displays the number of LNs included in the data provided by each data source. For example, the value "1956" in the #LNs column in the first row of Table 1 indicates that the number of LNs in the "NIH-LN" public dataset is 1956. Similarly, the other values in the #LNs column indicate the number of LNs included in the data provided by each clinical center.
[0206] It should be noted that "10435" in the eighth row of Table 1 indicates that the number of lymph nodes included in the data provided by all data sources is 10435.
[0207] The Average Resolution (mm) column in Table 1 refers to the average resolution of the data in each dataset. The data in the dataset can be CT images, so the Average Resolution refers to the average resolution of CT images. For example, the value "(0.82, 0.82, 2.0)" in the Average Resolution (mm) column in the first row of Table 1 indicates that the average resolution of CT images in the "NIH-LN" public dataset is (0.82, 0.82, 2.0). Similarly, the other values in the Average Resolution (mm) column represent the average resolution of CT images provided by each clinical center.
[0208] Among them, the background column in Table 1 refers to the purpose of each dataset. Table 1 counts 7 LN detection datasets, 5 of which are used as internal data for model development and internal testing, and the remaining 2 are used as independent external test sets.
[0209] The above datasets have enlarged (short axis > 1 cm) and smaller LNs. Table 1 shows the detailed number of patients, number of LNs, imaging scheme, and evaluation settings. Although NIH-LN is a public LN dataset, the other six datasets are from five clinical centers. The object recognition model training method provided in this disclosure uses NIH-LN and four datasets from centers 1-3 (covering head and neck cancer, esophageal cancer, and lung cancer) as internal datasets to develop and internally test model performance. The datasets from centers 4 and 5 are used as independent external test sets, among which center 4Mul contains three different types of patients, namely lung cancer, cancer, and infectious lung disease.
[0210] For the five internal datasets, each dataset was randomly split into 70% training, 10% validation, and 20% patient-level testing. The training and validation data from the five datasets were used together to develop and select LN detection models, and the remaining test patients from the five datasets were retained to report internal testing results. The Generic Lesion Dataset: DeepDisease consists of 32,735 lesions from 4,427 patients. It contains a variety of lesions, including lung nodules, liver / kidney / bone lesions, enlarged lymph nodes, and more. The official training / validation / test split was used, and results were reported on the test set.
[0211] Among them, for the evaluation indicators: For LN detection, the object recognition model training method provided by the present disclosure uses the free response receiver operating characteristic (FROC) curve as the evaluation indicator, and reports the sensitivity / recall rate of 0.5, 1, 2, and 4 FPs for each patient / CT volume. The 2D detection boxes of all methods are merged into 3D detection boxes. When the detected 3D box is compared with the GT 3D box, if their 3D intersection on the detected bounding box ratio (IoBB) is greater than 0.3, the predicted box is considered a true positive. For lymph node detection, since metastatic lymph nodes may be as small as 5 mm, the post-processing size threshold can be set to 5 mm so that enlarged and smaller LNs (short axis ≥ 5 mm) can be detected during the inference process. If a GT LN smaller than 5 mm is detected, it is not counted in TP (True Positive) and FP (False Positive). During training, the object recognition model training method provided in this disclosure uses LN annotations of various sizes. For DeepLesion detection, its official evaluation metrics are used: sensitivity / recall at 0.5, 1, 2, and 4 FPs per image / CT slice. The recall reported by DeepLesion is based on 2D images / CT slices; therefore, 3D box merging is not performed here.
[0212] In one or more embodiments provided in this illustration, the object recognition model training method provided by the present disclosure performs extensive comparative evaluation on LN detection, including a general target detection method and two instance segmentation methods.
[0213] It should be noted that in one or more embodiments provided in the present disclosure, ResNet is used as the CNN backbone for feature extraction. The model is initialized using weights pre-trained on the COCO dataset. During training, a batch size of 8 is used with an initial learning rate of 2e-4. A cosine learning rate scheduler is used to reduce the learning rate to 1e-5 with a warm-up step of 500. A weight decay of 1e-4 is set to avoid overfitting. In addition, the training data is enhanced by normalizing the 3D CT volume to a resolution of 0.8×0.8×2 mm and randomly applying horizontal flipping, cropping, scaling, and random noise.
[0214] All blocks in the ResNet employ 2.5D fusion, and the outputs of the last three blocks are sent to the Transformer encoder. Position and level embeddings are added to the flattened tokens in the encoder, providing spatial and horizontal position priors. To generate initial content queries and anchor boxes after the encoder, the top 300 queries are selected based on the proposed IOU-guided ranking criterion. In the decoder, the number of denoised queries is set to 100. During inference, the top 20 query predictions are selected as the final LN detection output.
[0215] In one or more embodiments provided herein, seven LN datasets were collected and collated, considering that disease type datasets require detailed information. NIH-LN is a public LN dataset, and the remaining six datasets come from five clinical centers. Specifically, NIH-LN includes 89 cancer patients. Clinical Center 1-HN includes 256 patients with head and neck cancer. Clinical Center 2 provides 91 patients with esophageal cancer (denoted as Clinical Center 2-Eso) and 97 patients with cancer (denoted as Clinical Center 2-lung). Clinical Center 3-Eso consists of another 300 patients with esophageal cancer. Clinical Centers 1-3 plus NIH-LN serve as internal data to develop and internally test LN detection performance. Datasets from Clinical Centers 4 and 5 were used as independent external test data, with Clinical Center-Mul. including 184 patients with different types of diseases (lung cancer, esophageal cancer, and infectious lung disease), and Clinical Center-HN including 50 patients with head and neck cancer.
[0216] In one or more embodiments provided herein, ResNet is used as the backbone. During training, the total number of training steps is increased to 50 to achieve further convergence, with all other settings remaining the same as for LN-DETR. During inference, the LN bounding boxes of these instance segmentation methods are predicted from their masks.
[0217] In one or more embodiments provided in the present disclosure, the object detection model obtained by training verifies the effectiveness of the proposed multi-scale 2.5D fusion, IoU-guided query selection (IQS) and contrastive learning (CL), and the results are summarized in Table 2.
[0218] Table 2
[0219] Among them, "2.5D" in Table 2 refers to multi-scale 2.5D fusion, "CL" refers to contrastive learning, and "IQS" refers to IoU-guided query selection. Based on this, the first column in Table 2 means selecting different schemes from "2.5D", "IQS" and "IQS" for lymph node recognition. For example, "√" in the first column of the second row of Table 2 (located below "2.5D") means selecting the multi-scale 2.5D fusion scheme for lymph node recognition. Similarly, "√√" in the first column of the third row of Table 2 (located below "2.5D" and "CL") means selecting the multi-scale 2.5D fusion and contrastive learning combination scheme for lymph node recognition. In this way, it means selecting different schemes from "2.5D", "IQS" and "IQS" for lymph node recognition.
[0220] Among them, the second column in Table 2 shows the effectiveness of each lymph node identification scheme in the process of selecting different schemes from "2.5D", "IQS" and "IQS" for lymph node identification. The effectiveness is expressed by the FPs recall rate (i.e., recall rate (%)@FPs in Table 2).
[0221] The "@0.5" in the second column refers to sample data using 0.5 FPs (FPs indicates a high false positive rate); the "@1" in the second column refers to sample data using 1 FPs; the "@2" in the second column refers to sample data using 2 FPs; and the "@4" in the second column refers to sample data using 4 FPs. It should be noted that after determining @0.5, @1, @2, and @4, the recall rates of multiple lymph node identification schemes were tested on @0.5, @1, @2, and @4. The "Average" in the second column refers to the average recall rate of @0.5, @1, @2, and @4.
[0222] Among them, the multiple values in the second column of Table 2 represent the recall rates and average recall rates of various lymph node recognition schemes at @0.5, @1, @2 and @4. For example, "39.34" in the second column of the second row of Table 2 means that the recall rate of the multi-scale 2.5D fusion scheme on the sample data of 0.5FPs is 39.34. "47.09" in the second column of the second row of Table 2 means that the recall rate of the multi-scale 2.5D fusion scheme on the sample data of 1FPs is 47.09. "51.32" in the second column of the second row of Table 2 refers to the average recall rate of the multi-scale 2.5D fusion scheme at @0.5, @1, @2 and @4 (i.e., the average recall rate). Similarly, the other values in the second column of Table 2 are also used to represent the recall rates and average recall rates of various lymph node recognition schemes at @0.5, @1, @2 and @4.
[0223] It should be noted that the first row of Table 2 refers to the recall rate and average recall rate of lymph node identification on @0.5, @1, @2 and @4 using other lymph node identification schemes.
[0224] Specifically, the effectiveness of the three components of LN-DETR, namely 2.5D feature fusion, IoU-guided query selection, and query contrastive learning, is demonstrated in Table 2. First, 2.5D feature fusion improves the average recall by 1.08% in the range of 0.5 to 4 FPs / patient (i.e., +1.08% in Table 2, indicating an increase in the average recall from 51.32% to 52.40%). Second, based on 2.5D fusion (average recall of 52.40%), the average recall of LN detection is improved by 1.04% (from 52.40% to 53.44%) and 2.11% (from 52.40% to 54.51%) using only query contrastive learning and IoU-guided query selection, respectively. More importantly, combining these two modules, LN-DETR improves the average recall by nearly 4%.
[0225] It should be noted that “+2.12%”, “+3.19%” and “+4.95%” in Table 2 indicate improvements of 2.12%, 3.19% and 4.95% compared to the average recall rate (51.32%) of other lymph node identification schemes.
[0226] Based on the above embodiments, it can be seen that the object recognition model training method provided by the present disclosure trains a model that uses position-enhanced query selection and contrast query representation in Transformer to effectively detect lymph nodes in CT scans. That is, a LN detection Transformer LN-DETR, which combines IoU-guided query selection and contrast query learning to enhance the representation ability of LN queries, which is crucial for improving detection sensitivity and reducing FP or repeated predictions. LN-DETR is also enhanced by combining 3D environments with an effective multi-scale 2.5D fusion scheme. By training and evaluating CT scans of multiple patients, the object recognition model training method provided by the present disclosure significantly improved the average recall rate by at least 4-5% in internal (5 datasets) and external (2 datasets) tests. The DeepLesion dataset was used to further evaluate the effect of LN-DETR on general lesion detection tasks, and its effectiveness was demonstrated by achieving good performance. When further evaluated on the general lesion detection task using DeepDisease, our method achieves state-of-the-art performance of 88.46% with an average recall ranging from 0.5 to 4 FPs per image.
[0227] The object recognition model training method provided by the present disclosure solves the critical but challenging task of lymph node detection. Based on the latest Transformer detection / segmentation framework Mask DINO, a new LN detection Transformer LN-DETR is proposed to achieve more accurate performance. In addition, an io-guided query selection with high positioning accuracy is introduced as the initialization of the decoder query, and a query comparison learning module is introduced to improve the quality of learning queries and reduce false positives and duplicate predictions. The object recognition model training method provided by the present disclosure was trained and tested on 3D CT scans of multiple patients with more than 10,000 labeled LNs (the largest LN dataset to date), significantly improving the average recall rate by at least 4-5% in internal and external tests. When further evaluated on the general lesion detection task using DeepLesion, LN-DETR achieved a recall rate of 88.47%.
[0228] The object recognition model training method provided by the present disclosure utilizes the object recognition model to be trained to perform object recognition on a sample image to obtain multiple sample candidate objects and object type recognition results corresponding to each sample candidate object; and obtains object position detection results by performing position detection on the object positions of the multiple sample candidate objects; and then trains the object recognition model to be trained by using the sample target objects and sample labels determined by the object type recognition results and the object position detection results to obtain an object recognition model that can accurately identify the target object positions corresponding to the target objects, thereby avoiding the problem of inaccurate target object recognition.
[0229] The following, in conjunction with Figure 5, takes the application of the object recognition model training method provided by the present disclosure in the lymph node recognition scenario as an example to further illustrate the object recognition model training method. Figure 5 shows a processing flow chart of an object recognition model training method provided by an embodiment of the present disclosure. It should be noted that Figure 5 is provided in two parts, the first part and the second part, to show the processing flow of the object recognition model training method, which specifically includes the following steps.
[0230] Step 502: Obtain an original CT scan as a training sample. In order to detect lymph nodes (LNs) in a target CT slice, extract multiple slices from the original CT scan.
[0231] Specifically, a large-scale LN CT scan is obtained as sample data. The CT scan contains more than 10,000 manually annotated lymph node (LN) outlines from multiple patients and involves different body parts and diseases.
[0232] The sample data is divided into a training sample set and a test sample set, which are used for model training and model testing respectively.
[0233] The steps of extracting multiple slices from the original CT scan are:
[0234] First, a target CT slice is identified in the original CT scan. The original CT scan is a 3D CT scan image of a patient's body. The target CT slice is the image slice (a cross section) in the 3D CT scan image that is closest to the location of the lymph nodes (LNs).
[0235] Secondly, with the target CT slice as the center, four upper slices and four lower slices are extracted from the original CT scan as the three-dimensional background of the central target CT slice.
[0236] Step 504: Perform concatenation and 1×1 conversion processing on multiple slices to obtain fusion features.
[0237] First, the nine consecutive CT slices are divided into three groups of three-channel images, that is, three slice sets. For example, slice set 1 contains four upper slices, slice set 2 contains four lower slices, and slice set 2 contains one target CT slice.
[0238] Second, the slice sets are input into the model as training samples, and each set is processed independently by a shared CNN backbone.
[0239] Again, feature extraction is performed on the three slice sets to obtain corresponding image features (i.e., feature mapping), and the image features corresponding to the three slice sets are fused through concatenation and 1×1 conversion processing to obtain fused image features.
[0240] Finally, the image features of the original central target slice in the slice set are replaced with the fused image features, and the image features of the other two sets remain unchanged.
[0241] Step 506: The slice set is processed into one dimension through a flatten layer to obtain a one-dimensional image feature vector.
[0242] Step 508: Add the one-dimensionalized image feature vector and the position vector (Position Embedding) corresponding to the image feature vector to obtain a 2.5D image token (i.e., 2.5D Image tokens, which can be understood as a feature vector); and input the 2.5D image token into the encoding layer contained in the Transformer encoder.
[0243] Step 510: The feature vector output by the Transformer encoding layer is input to the prediction heads contained in the Transformer encoder. The prediction heads are responsible for predicting the lymph node (LNs) location box, IOU score, classification score and lymph node image in the CT slice image.
[0244] Specifically, the prediction head includes: IoU prediction head (IoU prediction head consists of a single MLP layer), classification head, box regression heads and mask generation heads.
[0245] The bounding box regression head uses the feature vector corresponding to the CT slice image as input. Based on this feature vector, the bounding box regression head predicts the lymph node location boxes (also known as lymph node bounding boxes) and identifies multiple lymph node location boxes in the CT slice image. Based on these location boxes, the lymph nodes can be located.
[0246] For the classification head: The feature vector corresponding to the CT slice image and the lymph node location box are input into the classification head. The classification head evaluates the lymph node framed by the lymph node location box and determines the classification score for the lymph node. The higher the classification score, the higher the probability that the object framed by the lymph node location box is a lymph node.
[0247] For the mask generation head, the feature vector corresponding to the CT slice image and the lymph node location box are input into the mask generation head. The mask generation head performs image segmentation on the lymph nodes framed by the lymph node location box to obtain the feature vector representation corresponding to the lymph node image.
[0248] For the IoU prediction head: The feature vector corresponding to the CT slice image and the lymph node location box are input into the IoU prediction head. The IoU prediction head is used to perform positioning analysis on the lymph nodes framed by the lymph node location box to obtain an IoU score (i.e., IoU confidence). The IoU score can be used to estimate the query positioning quality, and the IoU confidence is used to guide the query sorting and selection in the final encoder layer.
[0249] It should be noted that during the model training process, the IoU prediction head, classification head, bounding box regression head, and mask generation head are trained using sample images and sample labels. The lymph node location box, classification score, lymph node image, and IoU score obtained based on the sample CT image in step 510 can be used to calculate the loss function later.
[0250] Taking the training process of the IoU prediction head as an example, the specific training process of the IoU prediction head is:
[0251] First, the classification scores of the CT slices and the true IoU between the lymph node boxes are used as labels; the CT slice images are used as sample images. The labels can be: the true IoU value between the ground truth (GTs) and the matching box (lymph node location box) prediction, which is used to supervise the IoU prediction.
[0252] Secondly, the feature vector corresponding to the sample image is input into the bounding box regression head to determine the location box of the lymph node in the sample image.
[0253] Again, the feature vector of the sample image and the lymph node location box are input into the IoU prediction head to obtain the IoU score corresponding to the lymph node location box.
[0254] Finally, in the loss calculation stage, assuming that Hungarian Matching assigns the i-th GT to the j-th query, the IoU score between the location box of the j-th query and the matching GT box can be calculated. The IoU loss of the i-th GT is defined as:
[0255] Step 512: For the multiple query results obtained in step 514, multiply the IoU scores of the query results by the classification scores to obtain new query result ranking scores, and select the K query results with the highest scores and the corresponding predicted location boxes.
[0256] The query result includes: the classification score corresponding to the lymph node, the lymph node image, and the IoU score.
[0257] The predicted location box refers to: the location box of the lymph node.
[0258] Step 514: Input the K query results as initial contents and the predicted location boxes as initial proposals to the decoding layer of the Transformer decoder. At the same time, input the GT embeddings and the corresponding GT boxes + noise to the decoding layer.
[0259] It should be noted that GT refers to: the manually annotated box (Ground-Truth) in an image, that is, the label; and the position box estimated by the algorithm is predictions, which can be called Bboxes, that is, bounding boxes.
[0260] Specifically, in this scheme, each GT and its most matching output query need to form a positive query pair, while all other non-matching queries of the GT are regarded as negative queries. In this setting, the GT needs to have its query representation, for example, it can be derived by processing the GT box and embedding the label through the Transformer decoder.
[0261] Step 516: Denoise the feature vector output by the decoding layer and perform matching.
[0262] The specific method is:
[0263] First, the GT vector with added noise is denoised to obtain multiple denoised GT queries, which are composed of GT vectors and GT ground truth boxes (i.e., labels). Assume that there are N denoised groups after denoising, and each group contains M denoised GT queries, i.e., anchor queries. Therefore, the total number of anchor queries q DN for:
[0264] where M is the number of GT LN boxes in the CT slice.
[0265] Step 518: Matching the K query results and the corresponding predicted location boxes with the GT queries in the denoising group.
[0266] Assuming that the matching cost between the K-th query result and the i-th GT in the CT slice is the smallest, then q k For query with N anchor points The matching pair is the positive query pair, and the other unmatched K-1 query results are Negative query pairs.
[0267] Step 520: Calculate the similarity between all positive anchor pairs (positive query pairs) and negative anchor pairs (negative query pairs) through comparison calculation, and determine the loss value based on the similarity.
[0268] For example, a shared MLP layer φ can be used to project the embedding of the query pair into the latent space. By projecting the query embedding, the similarity between all positive and negative anchor pairs is calculated.
[0269] Afterwards, the InfoNCE loss is used to pull the matched positive queries closer to their designated anchor queries, while moving away from all other unmatched negative queries. Based on this, the total query contrast loss formula for a given CT slice is shown in Equation 4 above.
[0270] Step 522: Calculate the total loss value.
[0271] During the training phase, the total loss is the original loss, i.e., the loss of the classification head, the loss of the bounding box regression head, and the loss of the mask head, as well as the combination of the IoU prediction loss and the query contrast loss proposed in the above steps. For details, see the above formula 2.
[0272] Among them, the loss value of the classification head is: the classification score obtained by processing the sample image through the classification head, and the sample label corresponding to the classification head (that is, the actual lymph node classification score) is calculated.
[0273] Loss value of the bounding box regression head: The lymph node location box obtained by processing the sample image through the bounding box regression head is calculated based on the sample label corresponding to the bounding box regression head (i.e. the above-mentioned GT box, the real lymph node location box).
[0274] The loss value of the mask terminal is: the lymph node image obtained by performing image segmentation on the sample image through the mask terminal, and the sample label corresponding to the mask terminal (that is, the above-mentioned real lymph node image) is calculated.
[0275] The LN-DETR model is trained using this total loss value until the model training stopping condition is reached.
[0276] It should be noted that when the model training is completed and the LN-DETR model is applied, the patient's CT image can be input into the LN-DETR model for lymph node identification. The specific steps are as follows:
[0277] First: The CT image is processed through the CNN backbone network in the LN-DETR model to obtain the image feature vector, which is then input into the Transformer encoder.
[0278] Secondly, the encoding layer contained in the Transformer encoder encodes the image feature vector and inputs the output feature vector into the prediction head contained in the Transformer encoder. It is processed using the IoU prediction head, classification head, bounding box regression head and mask generation head to obtain the lymph node location box, classification score, lymph node image and IoU score corresponding to the CT image.
[0279] Again, the IoU score of the query result is multiplied by the classification score to obtain a new query result ranking score, and the K query results with the highest scores and the corresponding predicted location boxes are selected and input into the Transformer decoder.
[0280] Finally, the Transformer decoder decodes the K query results and corresponding predicted location boxes contained in the encoding layer, and outputs multiple lymph node query results with high accuracy.
[0281] Based on the above steps, it can be seen that the object recognition model training method provided by the present disclosure is a solution that can simultaneously detect all visible lymph nodes larger than 5mm in multiple body parts (covering the head and neck, chest, and upper abdomen). It is also a lymph node detection method based on the Transformer framework. By adding 2.5D feature fusion technology features to the model, and proposing location-enhanced query selection and contrastive query representation of queries, the detection accuracy is improved. In addition, in addition to achieving good performance in lymph node detection tasks, this solution also achieves good performance in general lesion detection tasks, which will provide basic support for subsequent lymph node-related tasks.
[0282] In addition, the solution provided by the present disclosure has also achieved good performance in tasks such as detecting lung nodules in CT images and detecting prostate cancer in MRI images, and can be widely used to detect various target objects in medical images.
[0283] The present disclosure also provides a method for identifying lesions in liver CT images, including:
[0284] Identify the liver CT images containing the lesion;
[0285] The liver CT image is input into a lesion recognition model for lesion recognition to obtain a target lesion in the liver CT image containing the lesion, wherein the lesion recognition model is obtained by training sample target lesions identified from a liver CT sample image containing the lesion and sample labels of the liver CT sample image containing the lesion, the sample target lesions are determined from the multiple sample candidate lesions through lesion type recognition results and lesion position detection results of the multiple sample candidate lesions, the lesion position detection results are obtained by performing position detection on the lesion positions of the multiple sample candidate lesions, and the multiple sample candidate lesions are obtained by performing lesion recognition on the liver CT sample image containing the lesion.
[0286] The lesion recognition model is obtained by training the object recognition model training method in the above embodiment. The lesion recognition model can be understood as the object recognition model in the above target object recognition method or the object recognition model training method, and the target lesion can be understood as the target object in the above target object recognition method.
[0287] The liver CT image containing the lesion can be understood as an image to be identified in a target object identification method in the above embodiment.
[0288] The present disclosure provides a method for identifying lesions in liver CT images. During the training process, the lesion recognition model needs to determine multiple sample candidate lesions in a sample liver CT image containing lesions, as well as lesion type recognition results and lesion location detection results for the multiple sample candidate lesions, wherein the lesion location detection results are obtained by performing position detection on the lesion locations of the multiple sample candidate lesions. Model training is then performed based on the sample target lesions and sample labels determined by the lesion type recognition results and lesion location detection results, thereby obtaining an object recognition model that accurately determines the location of target lesions in liver CT images containing lesions. Based on this, when the liver CT image containing the lesion is input into the lesion recognition model for lesion recognition, the target lesion can be identified with accurate location, avoiding the problem of inaccurate target lesions.
[0289] The above is a schematic diagram of a method for identifying lesions in liver CT images according to this embodiment. It should be noted that the technical solution of this method for identifying lesions in liver CT images is based on the same concept as the technical solutions of the target object identification method and the object recognition model training method described above. For details not described in detail in the technical solution of the method for identifying lesions in liver CT images, please refer to the description of the technical solutions of the target object identification method and the object recognition model training method described above.
[0290] The present disclosure also provides a method for identifying lesions in a lymph node CT image, including:
[0291] Identify the CT images of the lymph nodes containing the lesion;
[0292] The lymph node CT image is input into a lesion recognition model for lesion recognition to obtain a target lesion in the lymph node CT image containing the lesion, wherein the lesion recognition model is obtained by training with sample target lesions identified from a lymph node CT sample image containing the lesion and sample labels of the lymph node CT sample image containing the lesion, the sample target lesions are determined from the multiple sample candidate lesions through lesion type recognition results and lesion location detection results of the multiple sample candidate lesions, the lesion location detection results are obtained by performing position detection on the lesion locations of the multiple sample candidate lesions, and the multiple sample candidate lesions are obtained by performing lesion recognition on the lymph node CT sample image containing the lesion.
[0293] The lesion recognition model is obtained by training the object recognition model training method in the above embodiment. The lesion recognition model can be understood as the object recognition model in the above target object recognition method or the object recognition model training method, and the target lesion can be understood as the target object in the above target object recognition method.
[0294] The lymph node CT image containing the lesion can be understood as an image to be identified in a target object identification method in the above embodiment.
[0295] The present disclosure provides a method for identifying lesions in lymph node CT images. During the training process, the lesion recognition model needs to determine multiple sample candidate lesions in a sample lymph node CT image containing lesions, as well as lesion type recognition results and lesion location detection results for the multiple sample candidate lesions, wherein the lesion location detection results are obtained by performing position detection on the lesion locations of the multiple sample candidate lesions. Model training is then performed based on the sample target lesions and sample labels determined by the lesion type recognition results and lesion location detection results, thereby obtaining an object recognition model that accurately determines the location of target lesions in lymph node CT images containing lesions. Based on this, when the lymph node CT image containing the lesion is input into the lesion recognition model for lesion recognition, the target lesion can be identified with accurate location, avoiding the problem of inaccurate target lesions.
[0296] The above is a schematic diagram of a method for identifying lesions in lymph node CT images according to this embodiment. It should be noted that the technical solution of this method for identifying lesions in lymph node CT images is based on the same concept as the technical solutions of the target object identification method and the object recognition model training method described above. For details not described in detail in the technical solution of the method for identifying lesions in lymph node CT images, please refer to the description of the technical solutions of the target object identification method and the object recognition model training method described above.
[0297] The present disclosure also provides a target object processing method, which is a computer-aided diagnosis method and is applied to a client of a medical system. The method includes:
[0298] In response to a user's clicking operation on a display interface of the client, determining a medical image to be identified;
[0299] Sending the medical image to be identified to a server of the medical system, and receiving a target object returned by the server, wherein the target object is an object output after object recognition processing is performed on the medical image to be identified by an object recognition model, the object recognition model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from a plurality of sample candidate objects through object type recognition results and object position detection results of a plurality of sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the plurality of sample candidate objects, and the plurality of sample candidate objects are obtained by performing object recognition on the sample medical images;
[0300] The target object is presented to the user through the presentation interface.
[0301] The target object processing method provided by one or more embodiments of the present disclosure is a computer-aided diagnosis method. After a user sends a medical image to be identified to a server via a client, the server's object recognition model performs object recognition processing on the medical image to be identified and outputs a target object. During the training process, the object recognition model needs to determine multiple sample candidate objects in the sample medical image, as well as the object type recognition results and object position detection results of the multiple sample candidate objects, wherein the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. Then, based on the sample labels and the sample target objects determined by the object type recognition results and the object position detection results, a model is trained to obtain an object recognition model that accurately identifies the position of the target object in the medical image to be identified. Based on this, when the medical image to be identified is input into the object recognition model for object recognition, the target object can be identified with an accurate position, avoiding the problem of inaccurate target objects. In addition, the target object is displayed to the user via the client's display interface, avoiding the problem of the user being informed of an inaccurate target object.
[0302] The above is a schematic scheme of a target object processing method and a computer-aided diagnosis method of this embodiment. It should be noted that the technical scheme of the target object processing method and the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, and information processing method are based on the same concept. For details not described in detail in the technical scheme of the target object processing method, please refer to the description of the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, and information processing method.
[0303] The present disclosure also provides an information processing method, which is a computer-aided diagnosis method. The method includes:
[0304] determining a medical image to be identified;
[0305] Inputting the medical image to be identified into an object recognition model for object recognition to obtain a target object in the medical image to be identified, wherein the object recognition model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from a plurality of sample candidate objects through object type recognition results and object position detection results of the plurality of sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the plurality of sample candidate objects, and the plurality of sample candidate objects are obtained by performing object recognition on the sample medical images;
[0306] An information processing result is determined according to the target object in the medical image to be identified.
[0307] The information processing result can be understood as the diagnosis result of a specific disease, such as the diagnosis result of lymphoma, the diagnosis result of lung cancer, etc.
[0308] One or more embodiments of the present disclosure provide an information processing method and a computer-aided diagnosis method. During the training process, the object recognition model needs to determine multiple sample candidate objects in the sample medical image, as well as the object type recognition results and object position detection results of the multiple sample candidate objects, wherein the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. Model training is then performed based on sample labels and sample target objects determined by the object type recognition results and the object position detection results, thereby obtaining an object recognition model that accurately identifies the position of the target object in the medical image to be identified. Based on this, when the medical image to be identified is input into the object recognition model for object recognition, the target object with accurate position can be identified, avoiding the problem of inaccurate target object. Furthermore, in the process of determining the information processing result based on the accurate target object, the accuracy of the information processing result can be guaranteed, avoiding obtaining an erroneous information processing result.
[0309] The above is a schematic scheme of an information processing method and computer-aided diagnosis method of this embodiment. It should be noted that the technical scheme of the information processing method and the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, and target object processing method are based on the same concept. For details not described in detail in the technical scheme of the information processing method, please refer to the description of the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, and target object processing method.
[0310] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a target object recognition device, the device comprising:
[0311] An image determination module is configured to determine an image to be identified;
[0312] An object recognition module is configured to input the image to be recognized into an object recognition model for object recognition to obtain a target object in the image to be recognized, wherein the object recognition model is obtained by training sample target objects identified from sample images and sample labels of the sample images, the sample target objects are determined from the multiple sample candidate objects through object type recognition results and object position detection results of multiple sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects, and the multiple sample candidate objects are obtained by performing object recognition on the sample images.
[0313] Optionally, the object recognition module is further configured to:
[0314] Inputting the image to be identified into an object recognition model, determining a plurality of candidate objects in the image to be identified using the object recognition model, and confirming an object type recognition result and an object position detection result of each candidate object;
[0315] The target object is determined from the plurality of candidate objects based on the object type recognition result and the object position detection result.
[0316] Optionally, the object recognition module is further configured to:
[0317] Determining the image features to be recognized corresponding to the image to be recognized using the object recognition model;
[0318] Performing object recognition on the features of the image to be recognized, determining a plurality of candidate objects in the image to be recognized, and object type recognition results of the plurality of candidate objects;
[0319] Perform position detection on the object positions of the multiple candidate objects to obtain an object position detection result corresponding to each candidate object.
[0320] Optionally, the object recognition module is further configured to:
[0321] Processing the image to be recognized using the image processing module in the object recognition model to obtain candidate image features;
[0322] The encoding module in the object recognition model is used to encode the candidate image features to obtain the image features to be recognized corresponding to the image to be recognized.
[0323] Optionally, the object recognition module is further configured to:
[0324] Using the image processing module in the object recognition model, feature extraction is performed on multiple sets of images to be recognized to obtain image features corresponding to each set of images to be recognized;
[0325] Performing feature concatenation processing and feature conversion processing on the image features corresponding to the plurality of sets of images to be identified to obtain target image features;
[0326] Determining a target image set to be identified from the multiple image sets to be identified, and replacing image features corresponding to the target image set to be identified with the target image features;
[0327] The candidate image features are obtained based on the image features corresponding to the multiple sets of images to be identified.
[0328] Optionally, the object recognition module is further configured to:
[0329] Using the object position recognition module in the object recognition model to perform object position recognition on the image features to be recognized, and determine the object positions of multiple candidate objects in the image to be recognized;
[0330] Inputting the object positions of the multiple candidate objects and the features of the image to be identified into the object recognition module in the object recognition model to perform object recognition, thereby obtaining multiple candidate objects in the image to be identified;
[0331] The object classification module in the object recognition model is used to perform object type recognition on each candidate object, and an object type recognition score corresponding to each candidate object is determined.
[0332] Optionally, the object recognition module is further configured to:
[0333] The object positions of the multiple candidate objects and the image features to be identified are input into the object position detection module in the object recognition model for position detection to obtain an object position detection score corresponding to each candidate object.
[0334] Optionally, the object position detection result is an object position detection score, and the object type recognition result is an object type recognition score;
[0335] The object recognition module is further configured to:
[0336] multiplying the object position detection score and the object type recognition score to obtain object recognition scores for the multiple candidate objects;
[0337] The target object is selected from the plurality of candidate objects based on the object recognition score.
[0338] Optionally, the image determination module is further configured to:
[0339] Determining image data to be identified, and performing image segmentation on the image data to be identified to obtain a plurality of segmented images;
[0340] determining a target segmented image and other segmented images except the target segmented image from the plurality of segmented images;
[0341] Constructing a target set of images to be identified based on the target segmented image, and constructing other sets of images to be identified based on the other segmented images;
[0342] The target set of images to be recognized and the other sets of images to be recognized are used as images to be recognized.
[0343] Optionally, the target object recognition device further includes a model training module configured to:
[0344] Determine a sample image of the object recognition model to be trained, and a sample label corresponding to the sample image;
[0345] Inputting the sample image into the object recognition model to be trained, performing object recognition on the sample image using the object recognition model to be trained, and obtaining a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object;
[0346] Obtaining an object position detection result for each sample candidate object by performing position detection on the object positions of the plurality of sample candidate objects;
[0347] determining a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result;
[0348] Based on the sample target object and the sample label, the object recognition model to be trained is trained to obtain the object recognition model.
[0349] The present disclosure provides a target object recognition device. During the training process, the object recognition model needs to determine multiple sample candidate objects in a sample image, as well as object type recognition results and object position detection results for these multiple sample candidate objects. The object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects. Model training is then performed based on the sample target objects and sample labels determined by the object type recognition results and the object position detection results, thereby obtaining an object recognition model that accurately determines the position of the target object in the image to be recognized. Based on this, when the image to be recognized is input into the object recognition model for object recognition, the target object can be accurately identified, avoiding the problem of inaccurate target objects.
[0350] The above is a schematic diagram of a target object recognition device according to this embodiment. It should be noted that the technical solution of the target object recognition device and the technical solution of the target object recognition method described above are based on the same concept. For details not described in detail in the technical solution of the target object recognition device, please refer to the description of the technical solution of the target object recognition method described above.
[0351] Corresponding to the above method embodiment, the present disclosure also provides an object recognition model training device embodiment, the device comprising:
[0352] A sample determination module is configured to determine a sample image of an object recognition model to be trained, and a sample label corresponding to the sample image;
[0353] a sample candidate object recognition module configured to input the sample image into the to-be-trained object recognition model, perform object recognition on the sample image using the to-be-trained object recognition model, and obtain a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object;
[0354] a position detection module configured to obtain an object position detection result of each sample candidate object by performing position detection on the object positions of the multiple sample candidate objects;
[0355] a sample target object recognition module configured to determine a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result;
[0356] The model training module is configured to train the object recognition model to be trained based on the sample target object and the sample label to obtain an object recognition model.
[0357] Optionally, the sample candidate object identification module is further configured to:
[0358] Determining sample image features corresponding to the sample image using the object recognition model to be trained;
[0359] Performing object position recognition on the sample image features using the object position recognition module in the object recognition model to be trained, and determining the object positions of a plurality of sample candidate objects in the sample image;
[0360] Inputting the object positions of the multiple sample candidate objects and the sample image features into the object recognition module in the object recognition model to be trained to perform object recognition, thereby obtaining multiple sample candidate objects in the sample image;
[0361] The object classification module in the object recognition model to be trained is used to perform object type recognition on each sample candidate object, and an object type recognition score corresponding to each sample candidate object is determined.
[0362] Optionally, the position detection module is further configured to:
[0363] The object positions of the multiple sample candidate objects and the sample image features are input into the object position detection module in the object recognition model to be trained for position detection, and the object position detection score corresponding to each sample candidate object is obtained.
[0364] Optionally, the model training module is further configured to:
[0365] Determine the sample candidate object label, object type recognition score label, object position label, object position detection score label, and sample target object label included in the sample label;
[0366] Determine a first loss value based on the object type recognition score and the object type recognition score label, determine a second loss value based on the object position and the object position label, determine a third loss value based on the sample candidate object label and the plurality of sample candidate objects, and determine a fourth loss value based on the object position detection score and the object position detection score label;
[0367] Determining a fifth loss value based on the sample target object and the sample target object label;
[0368] Based on the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value, the object recognition model to be trained is trained until a model training stop condition is reached to obtain an object recognition model.
[0369] Optionally, the model training module is further configured to:
[0370] Performing noise processing on the sample target object label to obtain a plurality of noise object labels corresponding to the sample target object label, wherein the number of the plurality of noise object labels is consistent with the number of the plurality of sample target objects;
[0371] Determining corresponding noise object labels for a plurality of sample target objects, and calculating associations between each sample target object and the noise object label corresponding to each sample target object to obtain a plurality of object association groups;
[0372] calculating an association score between a sample target object and a noise object label in each object association group, and dividing the plurality of object association groups into a first object association group and a second object association group based on the association score;
[0373] A fifth loss value is determined based on a similarity between the first object association group and the second object association group.
[0374] The object recognition model training device provided by the present disclosure uses the object recognition model to be trained to perform object recognition on a sample image to obtain multiple sample candidate objects and the object type recognition results corresponding to each sample candidate object; and obtains object position detection results by performing position detection on the object positions of the multiple sample candidate objects; and then trains the object recognition model to be trained by using the sample target objects and sample labels determined by the object type recognition results and the object position detection results, to obtain an object recognition model that can accurately identify the target object position corresponding to the target object, thereby avoiding the problem of inaccurate target object recognition.
[0375] The above is a schematic diagram of an object recognition model training device according to this embodiment. It should be noted that the technical solution of the object recognition model training device and the technical solution of the object recognition model training method described above are based on the same concept. For details not described in detail in the technical solution of the object recognition model training device, please refer to the description of the technical solution of the object recognition model training method described above.
[0376] Figure 6 shows a block diagram of a computing device 600 according to one embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0377] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0378] In one embodiment of the present disclosure, the aforementioned components of computing device 600 and other components not shown in FIG6 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG6 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0379] Computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 may also be a mobile or stationary server.
[0380] Among them, the processor 620 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, target object processing method or information processing method.
[0381] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in CT images, target object processing method, or information processing method described above are of the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in CT images, target object processing method, or information processing method described above.
[0382] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, target object processing method, or information processing method.
[0383] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, target object processing method, or information processing method described above are of the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, target object processing method, or information processing method described above.
[0384] One embodiment of the present disclosure also provides a computer program product, wherein, when the computer program product is executed in a computer, the computer is caused to execute the steps of the above-mentioned target object recognition method, object recognition model training method, lesion recognition method in liver CT images, target object processing method or information processing method.
[0385] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of this computer program product and the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, target object processing method, or information processing method described above are based on the same concept. For details not described in detail in the technical scheme of the computer program product, please refer to the description of the technical schemes of the target object recognition method, object recognition model training method, liver lesion recognition method in liver CT images, target object processing method, or information processing method described above.
[0386] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0387] The computer instructions include computer program product codes, which may be in source code form, object code form, executable files, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program product code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0388] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.
[0389] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0390] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. A target object recognition method, comprising: Determine an image to be identified; The image to be identified is input into an object recognition model for object recognition to obtain a target object in the image to be identified, wherein the object recognition model is obtained by training sample target objects identified from sample images and sample labels of the sample images, the sample target objects are determined from the multiple sample candidate objects through object type recognition results and object position detection results of multiple sample candidate objects, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate objects, and the multiple sample candidate objects are obtained by performing object recognition on the sample image.
2. The target object recognition method according to claim 1, wherein inputting the image to be recognized into the object recognition model for object recognition to obtain the target object in the image to be recognized comprises: Inputting the image to be identified into an object recognition model, determining a plurality of candidate objects in the image to be identified using the object recognition model, and confirming an object type recognition result and an object position detection result of each candidate object; The target object is determined from the plurality of candidate objects based on the object type recognition result and the object position detection result.
3. The target object recognition method according to claim 2, wherein the determining of a plurality of candidate objects in the image to be recognized using the object recognition model and confirming an object type recognition result and an object position detection result for each candidate object comprises: Determining the image features to be recognized corresponding to the image to be recognized using the object recognition model; Performing object recognition on the features of the image to be recognized, determining a plurality of candidate objects in the image to be recognized, and object type recognition results of the plurality of candidate objects; Perform position detection on the object positions of the multiple candidate objects to obtain an object position detection result corresponding to each candidate object.
4. The target object recognition method according to claim 3, wherein determining the image features corresponding to the image to be recognized by using the object recognition model comprises: Processing the image to be recognized using the image processing module in the object recognition model to obtain candidate image features; The encoding module in the object recognition model is used to encode the candidate image features to obtain the image features to be recognized corresponding to the image to be recognized.
5. The target object recognition method according to claim 4, wherein the step of processing the image to be recognized using the image processing module in the object recognition model to obtain candidate image features comprises: Using the image processing module in the object recognition model, feature extraction is performed on multiple sets of images to be recognized to obtain image features corresponding to each set of images to be recognized; Performing feature concatenation processing and feature conversion processing on the image features corresponding to the plurality of sets of images to be identified to obtain target image features; Determining a target image set to be identified from the multiple image sets to be identified, and replacing image features corresponding to the target image set to be identified with the target image features; The candidate image features are obtained based on the image features corresponding to the multiple sets of images to be identified.
6. The target object recognition method according to claim 3, wherein the performing object recognition on the features of the image to be recognized, determining multiple candidate objects in the image to be recognized, and object type recognition results of the multiple candidate objects, comprises: Using the object position recognition module in the object recognition model to perform object position recognition on the image features to be recognized, and determine the object positions of multiple candidate objects in the image to be recognized; Inputting the object positions of the multiple candidate objects and the features of the image to be identified into the object recognition module in the object recognition model to perform object recognition, thereby obtaining multiple candidate objects in the image to be identified; The object classification module in the object recognition model is used to perform object type recognition on each candidate object, and an object type recognition score corresponding to each candidate object is determined.
7. The target object recognition method according to claim 6, wherein the performing position detection on the plurality of candidate objects to obtain an object position detection result corresponding to each candidate object comprises: The object positions of the multiple candidate objects and the image features to be identified are input into the object position detection module in the object recognition model for position detection to obtain an object position detection score corresponding to each candidate object.
8. The target object recognition method according to claim 2, wherein the object position detection result is an object position detection score, and the object type recognition result is an object type recognition score; The determining the target object from the plurality of candidate objects based on the object type recognition result and the object position detection result includes: multiplying the object position detection score and the object type recognition score to obtain object recognition scores for the multiple candidate objects; The target object is selected from the plurality of candidate objects based on the object recognition score.
9. The target object recognition method according to any one of claims 1 to 8, wherein determining the image to be recognized comprises: Determining image data to be identified, and performing image segmentation on the image data to be identified to obtain a plurality of segmented images; determining a target segmented image and other segmented images except the target segmented image from the plurality of segmented images; Constructing a target set of images to be identified based on the target segmented image, and constructing other sets of images to be identified based on the other segmented images; The target set of images to be recognized and the other sets of images to be recognized are used as images to be recognized.
10. The target object recognition method according to any one of claims 1 to 9, before inputting the image to be recognized into the object recognition model for object recognition and obtaining the target object in the image to be recognized, further comprising: Determine a sample image of the object recognition model to be trained, and a sample label corresponding to the sample image; Inputting the sample image into the object recognition model to be trained, performing object recognition on the sample image using the object recognition model to be trained, and obtaining a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object; Obtaining an object position detection result for each sample candidate object by performing position detection on the object positions of the plurality of sample candidate objects; determining a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result; Based on the sample target object and the sample label, the object recognition model to be trained is trained to obtain the object recognition model.
11. A method for training an object recognition model, comprising: Determine a sample image of the object recognition model to be trained, and a sample label corresponding to the sample image; Inputting the sample image into the object recognition model to be trained, performing object recognition on the sample image using the object recognition model to be trained, and obtaining a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object; Obtaining an object position detection result for each sample candidate object by performing position detection on the object positions of the plurality of sample candidate objects; determining a sample target object from the plurality of sample candidate objects based on the object type recognition result and the object position detection result; Based on the sample target objects and the sample labels, the object recognition model to be trained is trained to obtain an object recognition model.
12. The object recognition model training method according to claim 11, wherein the object recognition model to be trained is used to perform object recognition on the sample image to obtain a plurality of sample candidate objects and an object type recognition result corresponding to each sample candidate object, comprising: Determining sample image features corresponding to the sample image using the object recognition model to be trained; Performing object position recognition on the sample image features using the object position recognition module in the object recognition model to be trained, and determining the object positions of a plurality of sample candidate objects in the sample image; Inputting the object positions of the multiple sample candidate objects and the sample image features into the object recognition module in the object recognition model to be trained to perform object recognition, thereby obtaining multiple sample candidate objects in the sample image; The object classification module in the object recognition model to be trained is used to perform object type recognition on each sample candidate object, and an object type recognition score corresponding to each sample candidate object is determined.
13. The object recognition model training method according to claim 12, wherein the step of performing position detection on the plurality of sample candidate objects to obtain the object position detection result of each sample candidate object comprises: The object positions of the multiple sample candidate objects and the sample image features are input into the object position detection module in the object recognition model to be trained for position detection, and the object position detection score corresponding to each sample candidate object is obtained.
14. The object recognition model training method according to claim 13, wherein the training of the object recognition model to be trained based on the sample target object and the sample label to obtain the object recognition model comprises: Determine the sample candidate object label, object type recognition score label, object position label, object position detection score label, and sample target object label included in the sample label; Determine a first loss value based on the object type recognition score and the object type recognition score label, determine a second loss value based on the object position and the object position label, determine a third loss value based on the sample candidate object label and the plurality of sample candidate objects, and determine a fourth loss value based on the object position detection score and the object position detection score label; Determining a fifth loss value based on the sample target object and the sample target object label; Based on the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value, the object recognition model to be trained is trained until a model training stop condition is reached to obtain an object recognition model.
15. The object recognition model training method according to claim 13, wherein determining the fifth loss value based on the sample target object and the sample target object label comprises: Performing noise processing on the sample target object label to obtain a plurality of noise object labels corresponding to the sample target object label, wherein the number of the plurality of noise object labels is consistent with the number of the plurality of sample target objects; Determining corresponding noise object labels for a plurality of sample target objects, and calculating associations between each sample target object and the noise object label corresponding to each sample target object to obtain a plurality of object association groups; calculating an association score between a sample target object and a noise object label in each object association group, and dividing the plurality of object association groups into a first object association group and a second object association group based on the association score; A fifth loss value is determined based on a similarity between the first object association group and the second object association group.
16. A computer-aided diagnosis method, applied to a client of a medical system, comprising: In response to a user's clicking operation on a display interface of the client, determining a medical image to be identified; Sending the medical image to be identified to the server of the medical system, and receiving a target object returned by the server, wherein the target object is a lymph node identification result output after performing lymph node identification processing on the medical image to be identified through a lymph node identification model, the lymph node identification model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from the multiple sample candidate lymph nodes through object type identification results and object position detection results of multiple sample candidate lymph nodes, the object position detection result is obtained by performing position detection on the object positions of the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing lymph node identification on the sample medical image; The target object is presented to the user through the presentation interface.
17. A method for detecting visible lymph nodes in a CT image, comprising: determining a CT image to be identified; The CT image is input into a lymph node recognition model for lymph node recognition to obtain recognition results of visible lymph nodes in the CT image, wherein the lymph node recognition model is obtained by training sample target lymph nodes identified from CT sample images and sample labels of the CT sample images, the sample target lymph nodes are determined from the multiple sample candidate lymph nodes through lymph node type recognition results and lymph node position detection results of multiple sample candidate lymph nodes, the lymph node position detection results are obtained by performing position detection on the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing lymph node recognition on the CT sample image.
18. A computer-aided diagnosis method comprising: determining a medical image to be identified; Inputting the medical image to be identified into a lymph node recognition model to perform lymph node recognition, and obtaining a lymph node recognition result in the medical image to be identified, wherein the lymph node recognition model is obtained by training sample target objects identified from sample medical images and sample labels of the sample medical images, the sample target objects are determined from the multiple sample candidate lymph nodes through object type recognition results and object position detection results of multiple sample candidate lymph nodes, the object position detection results are obtained by performing position detection on the object positions of the multiple sample candidate lymph nodes, and the multiple sample candidate lymph nodes are obtained by performing object recognition on the sample medical image; The diagnosis result of the lymph node lesion is determined according to the lymph node recognition result in the medical image to be recognized.
19. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the target object recognition method according to any one of claims 1 to 10, the object recognition model training method according to any one of claims 11 to 15, the computer-aided diagnosis method according to any one of claims 16-17, and the lesion recognition method in lymph node CT images according to claim 18 are implemented.
20. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the target object recognition method described in any one of claims 1 to 10, the object recognition model training method described in any one of claims 11 to 15, the computer-aided diagnosis method described in any one of claims 16-17, and the lesion recognition method in lymph node CT images described in claim 18.
21. A computer program product, wherein When the computer program product is executed in a computer, the computer is caused to execute the steps of the target object recognition method described in any one of claims 1 to 10, the object recognition model training method described in any one of claims 11 to 15, the computer-aided diagnosis method described in any one of claims 16-17, and the lesion recognition method in lymph node CT images described in claim 18.
Citation Information
Patent Citations
Medical image detection
CN110622168A
Image processing method and device
CN116797554A
Cell detection and classification method and system based on grouping prompt learning
CN116844161A
Target object recognition method, object recognition model training method, target object processing method and information processing method
CN117809121A
Device and method for universal lesion detection in medical images
US20210224603A1
Cited By
Lymph node detection method, system and equipment based on anatomical environment perception and medium
CN121458728A