Method, equipment and medium for rib detection in medical imaging

The neural network model is used to encode and decode the features of the rib area in the CT image, and the rib query item is used to determine the rib detection result, which solves the problem of poor rib segmentation accuracy and achieves accurate rib instance segmentation.

CN115082389BActive Publication Date: 2025-10-03ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210648147.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-10-03
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

Existing rib segmentation methods have poor accuracy in CT images and are affected by conditions such as rib adhesion, structural damage, or image field of view.

Method used

A neural network model is used to encode and decode the features of the rib area in medical images. The rib query item is used to determine the feature vector to achieve rib detection and segmentation, including rib prediction category and posture parameter prediction value.

Benefits of technology

It achieves accurate segmentation of rib instances in three-dimensional space, reduces the interference of rib adhesion or structural damage on segmentation, and improves the intuitiveness and accuracy of the segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082389B_ABST
    Figure CN115082389B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device and medium for rib detection in medical images. Using a neural network model, feature encoding is performed on different rib regions in the medical image to obtain multiple feature vectors. Based on the correspondence between the learned rib query items and the different rib regions in the medical image, the feature vector of at least one rib query item can be determined. The feature vectors of the at least one rib query item are decoded in parallel to obtain the rib detection results corresponding to the at least one rib query item. The rib query item is endowed with semantic information in anatomy, so that it can focus on the features of different rib regions, thereby realizing controllable, instance-level rib detection. The above-mentioned image detection method enables any rib detection result to include the rib prediction category and the predicted value of the posture parameters of the rib detection box in three-dimensional space, so that the position of the rib instance can be accurately determined in three-dimensional space to accurately realize rib instance segmentation in medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, and medium for detecting ribs in medical images. Background Art

[0002] Computer Aided Diagnosis (CAD) technology can help doctors detect lesions based on the powerful analytical and computing capabilities of computers, combined with imaging, medical image processing technology and other possible physiological and biochemical methods.

[0003] Automatically segmenting instance-level ribs from computed tomography (CT) images is a prerequisite for many rib-related applications. However, existing rib segmentation methods suffer from poor accuracy due to issues such as rib adhesion, structural damage, and image field of view (FOV). Therefore, a new solution is needed. Summary of the Invention

[0004] Various aspects of the present application provide a method, device, and medium for detecting ribs in medical images, so as to accurately segment instance-level ribs from medical images.

[0005] An embodiment of the present application provides a method for rib detection in medical images, comprising: obtaining a three-dimensional medical image containing ribs; utilizing a neural network model to perform feature encoding on different rib regions in the medical image to obtain a plurality of feature vectors; determining, based on the correspondence between learned rib query items and different rib regions in the medical image, a feature vector corresponding to each of the at least one rib query items; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; performing parallel decoding on the feature vectors corresponding to each of the at least one rib query items to obtain a rib detection result for each of the at least one rib query items; the rib detection result corresponding to any rib query item includes: a predicted rib category corresponding to the rib query item and a predicted value of a pose parameter of a rib detection box in three-dimensional space.

[0006] An embodiment of the present application also provides a rib detection method for medical images, which is applied to AR devices, including: obtaining a three-dimensional medical image containing ribs; using a neural network model to feature encode different rib regions in the medical image to obtain multiple feature vectors; determining the feature vector corresponding to at least one rib query item based on the correspondence between the learned rib query items and the different rib regions in the medical image; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; parallel decoding of the feature vectors corresponding to the at least one rib query item to obtain a rib detection result corresponding to the at least one rib query item; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space; and superimposing and displaying the rib detection result corresponding to the at least one rib query item in the captured real image.

[0007] An embodiment of the present application also provides a method for rib detection in medical images, comprising: responding to a call request from a client to a first interface, obtaining a three-dimensional medical image containing ribs from the interface parameters of the first interface; utilizing a neural network model to perform feature encoding on different rib regions in the medical image to obtain a plurality of feature vectors; determining, based on the correspondence between the learned rib query items and the different rib regions in the medical image, a feature vector corresponding to each of the at least one rib query items; wherein the at least one rib query item corresponds one-to-one to at least one rib instance to be detected; performing parallel decoding on the feature vectors corresponding to each of the at least one rib query items to obtain a rib detection result corresponding to each of the at least one rib query items; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space.

[0008] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the method provided in the embodiment of the present application.

[0009] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps in the method provided in the embodiment of the present application.

[0010] In the rib detection method for medical images provided in an embodiment of the present application, a neural network model can be used to perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors. Based on the learned correspondence between the rib query items and the different rib regions in the medical image, the feature vector corresponding to at least one rib query item can be determined. The feature vector corresponding to each of the at least one rib query items is decoded in parallel to obtain the rib detection result corresponding to each of the at least one rib query items. The rib query items are endowed with semantic information in anatomy, so that they can focus on the features of different rib regions, thereby achieving controllable, instance-level rib detection. At the same time, any rib detection result, including the rib prediction category and the predicted value of the pose parameters of the rib detection box in three-dimensional space, can accurately determine the location of the rib instance in three-dimensional space based on the predicted value of the pose parameters, thereby accurately achieving rib instance segmentation in three-dimensional medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0012] Figure 1 A flowchart of a rib detection method for medical images provided by an exemplary embodiment of the present application;

[0013] Figure 2 A schematic diagram of the structure of a neural network model provided for an exemplary embodiment of the present application;

[0014] Figure 3 A schematic diagram of rib segmentation results when a rib query item is provided;

[0015] Figure 4 Schematic diagram of the binary matching process between rib query items and rib category truth values;

[0016] Figure 5 Schematic diagram of the weighted adjacency matrix used to determine the penalty term for bipartite matching;

[0017] Figure 6 The figure shows the segmentation and labeling results of 24 rib instances in medical images.

[0018] Figure 7 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0019] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] In response to the technical problems existing in the prior art, a solution is provided in some embodiments of the present application. The technical solutions provided in each embodiment of the present application are described in detail below with reference to the accompanying drawings.

[0021] Figure 1 A flowchart of a rib detection method for medical imaging provided by an exemplary embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0022] Step 101: Acquire a three-dimensional medical image containing ribs.

[0023] Step 102: Using a neural network model, perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors.

[0024] Step 103: Determine a feature vector corresponding to each of the at least one rib query items based on the learned correspondence between the rib query items and different rib regions in the medical image; wherein the at least one rib query item corresponds one-to-one to the at least one rib instance to be detected.

[0025] Step 104: Parallel decoding is performed on the feature vectors corresponding to the at least one rib query item to obtain a rib detection result for each of the at least one rib query item. The rib detection result corresponding to any rib query item includes: a rib prediction category and a predicted value of a pose parameter of a rib detection box in three-dimensional space.

[0026] The rib segmentation method provided in this embodiment can be used to segment instance-level ribs from medical images. Medical images are images obtained by performing axial tomographic scanning of the chest and abdomen, including the ribs, using CT technology. Under normal circumstances, a CT image of the human chest and abdomen shows 24 ribs, evenly distributed on the left and right sides of the chest cavity.

[0027] In this embodiment, since the number of ribs in medical images is relatively fixed, the rib detection problem of medical images can be converted into a set prediction problem of detection frames to improve detection efficiency. The neural network model can be implemented based on a transformer architecture. In the transformer-based neural network model, the encoder can learn to focus on the features of different rib regions on the input image through training, and can learn multiple target query items (object query) through the decoder, and can decode multiple targets in parallel. The number of target query items can be customized, and the binding relationship between the target query items and the query objects can be learned during the training process. Based on the learned binding relationship between the target query items and the query objects, the neural network model can achieve controllable object detection and segmentation.

[0028] In this embodiment, after the medical image is input into the neural network model, an encoder can be used in the neural network model to perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors.

[0029] When used for rib instance segmentation, the decoder can learn target queries about rib instances, described below as rib queries. In this embodiment, the ribs can be numbered sequentially from top to bottom and from left to right, starting with the first rib in the upper left corner, according to their distribution pattern. Rib numbers 1 to 24 are obtained and used as rib queries. During the training phase, rib numbers 1 to 24 can all be used as rib queries to learn the intrinsic relationship between each rib query and the location of its corresponding rib instance in the medical image.

[0030] The decoder input includes at least one rib query item and multiple feature vectors output by the encoder. Each rib query item corresponds to a specific rib to be detected. The ribs to be detected can be user-specified or, by default, a subset or all of the ribs in the medical image.

[0031] In some embodiments, during the prediction phase, the at least one rib query item may include a rib query item corresponding to a target rib input by the user. For example, if a user requests detection of the third rib on the left side, the number 3 corresponding to the third rib on the left side may be used as a rib query item to participate in the rib instance detection process.

[0032] In other implementation examples, during the prediction phase, the at least one rib query item may include: a plurality of preset rib query items corresponding one-to-one to a plurality of ribs in the medical image. That is, the default rib query items involved in the rib instance detection process include rib numbers 1 to 24.

[0033] During the decoder training phase, a bipartite matching method is used to bind rib query terms to the ground truth of the rib category. This allows the decoder to learn the correspondence between rib query terms and different rib regions in the medical image. In other words, each rib query term can be trained to focus on a different rib region in the medical image. Based on this correspondence, the decoder can determine the feature vector corresponding to each of the at least one rib query terms input. In other words, each rib query term acquires the feature vector of the region of interest.

[0034] The decoder can perform parallel decoding on the feature vectors corresponding to at least one rib query item to obtain the rib detection results of at least one rib query item. The rib detection results corresponding to any rib query item may include: the rib prediction category corresponding to the rib query item and the predicted value of the posture parameter of the rib detection frame in the three-dimensional space. The rib prediction category includes: 24 different rib categories and non-rib categories (i.e., background categories). The rib category can be marked with a rib number. The posture parameter is used to describe the position and posture of the detection frame of the rib instance. The position can be the position of the center point of the detection frame of the rib instance in the three-dimensional space, and the posture can be the rotation angle of the detection frame in the three-dimensional space.

[0035] In some embodiments, the pose parameters of the rib detection frame in three-dimensional space may include at least nine degrees of freedom (DOF) pose parameters, namely, the coordinates (x, y, z) of the center point of the rib detection frame in the XYZ coordinate system corresponding to the medical image, the scale information (w, h, d) on the X-axis, Y-axis, and Z-axis, and the rotation angles (α, β, γ) on the X-axis, Y-axis, and Z-axis. That is, the predicted pose parameters output by the neural network may include at least the predicted values ​​of the parameters for the aforementioned nine degrees of freedom. The X direction corresponds to the rightward direction of the human body, the Y direction corresponds to the front (chest) and back (back) directions of the human body, and the Z direction corresponds to the top (head) and bottom (feet) directions of the human body.

[0036] In this embodiment, a neural network model can be used to perform feature encoding on different rib regions in a medical image to obtain multiple feature vectors. Based on the learned correspondence between rib query items and different rib regions in the medical image, a feature vector corresponding to at least one rib query item can be determined. The feature vectors corresponding to each of the at least one rib query items are decoded in parallel to obtain a rib detection result corresponding to each of the at least one rib query items. The rib query items are endowed with anatomical semantic information, allowing them to focus on the features of different rib regions, thereby achieving controllable, instance-level rib detection. Furthermore, any rib detection result, including the predicted rib category and the predicted pose parameter value of the rib detection box in three-dimensional space, can accurately determine the location of the rib instance in three-dimensional space based on the predicted pose parameter value, thereby accurately achieving rib instance segmentation in three-dimensional medical images.

[0037] Based on the above embodiments, after obtaining the rib detection results corresponding to at least one rib query item, the neural network model can also segment rib instances from the medical image according to the rib detection results of at least one rib query item to obtain an image segmentation result.

[0038] This implementation employs a "detect first, segment later" paradigm, with each rib query corresponding to a rib detection result, and each rib detection result segmented into a rib instance. This enables instance-level rib segmentation, minimizing interference from rib adhesion or structural damage, and improving the intuitiveness and accuracy of rib segmentation results.

[0039] In some exemplary embodiments, the neural network model used in this application can be based on Figure 2 The illustrated Transformer network structure or its variant network implementation is not limited in this embodiment.

[0040] like Figure 2 As shown in FIG, when performing the segmentation operation of the rib instance, the neural network model mainly includes: an input layer, a RoI (region of interest) extractor, a rib query module (Rib-Query), a detection box detection module, and a segmentation network connected in sequence.

[0041] The RoI extractor and segmentation network can be implemented based on independent segmentation modules to perform different segmentation tasks at different spatial resolutions. The segmentation module can be implemented as a U-Net, V-Net, or FCN (Fully Convolutional Network) module, including but not limited to these. The rib query module is implemented based on a Transformer network, primarily comprising a feature extractor, an encoder, and a decoder.

[0042] The following will be combined Figure 2 The detection and segmentation process of the rib instance is further illustrated.

[0043] After the medical image is input into the neural network model, a RoI extractor can be used to extract the three-dimensional region of interest (ROI) containing the ribs from the medical image. Medical images are typically anisotropic, with inconsistent pixel spacing in the X, Y, and Z scanning directions of the local coordinate system. The pixel spacing in the X and Y directions is smaller and has a higher resolution, typically 0.5 mm. The pixel spacing in the Z direction is larger, typically 1 to 3 mm. To facilitate processing by the segmentation network, the pixel spacing in the medical image can be adjusted to obtain an isotropic image. This is then input into the RoI extractor, ensuring consistent pixel spacing in the X, Y, and Z directions. The RoI extractor can then extract the three-dimensional region of interest (ROI) containing the ribs from the isotropic image. For example, the pixel spacing in the X, Y, and Z directions of the three-dimensional medical image can be adjusted to 3 mm. The RoI extractor can then be trained and inferred at a resolution of 3 mm to more accurately extract the ROI containing the ribs.

[0044] In the rib query module, the feature extractor can extract features from the region of interest to obtain spatial features. Figure 2 As shown in the figure, the flattened spatial features output by the feature extractor are input to the encoder. The encoder primarily consists of a self-attention network and a feed-forward neural network (FFN). In the encoder, the self-attention network performs self-attention encoding on the input spatial features based on the self-attention mechanism, generating feature vectors corresponding to different rib regions.

[0045] During the training phase, a feature extractor may be used to extract features from the region of interest. Before obtaining spatial features, data augmentation may be performed on the region of interest to improve the generalization performance and training efficiency of the model. Optionally, data augmentation may be performed on the region of interest by performing at least one of random displacement, random scaling, random rotation, random cropping, and random erasing.

[0046] Because ribs vary in position, orientation, and scale, local ambiguity may exist between different CT scans. By performing operations such as random displacement, random scaling, and random rotation, the input data can be made more diverse to accurately identify each rib instance. Random cropping can be performed along the Z axis in three-dimensional space. This can truncate some rib instances to mimic various data from clinical situations. Random erasing can remove ribs at the bottom of the image from the training sample with a certain probability to mitigate potential over-prediction.

[0047] The feature vectors corresponding to different rib regions output by the encoder are input to the decoder. The decoder input also includes at least one rib query item, such as Figure 2 q1, q2…q23, q24 are shown. The decoder is primarily composed of a self-attention network, an attention network, and a feedforward neural network. After the at least one rib query item is input into the decoder, self-attention calculations can be performed in the self-attention network, thereby enabling information exchange between the at least one rib query item and achieving an overall arrangement of rib instances. After self-attention calculations, each rib query item is input into the attention network. For any rib query item, the attention network can perform attention calculations on the input rib query item and the feature vector corresponding to the rib query item based on the attention mechanism to obtain a decoding vector.

[0048] Each decoded vector output by the decoder is fed into the classification layer and the bounding box prediction layer. The classification layer performs classification calculations based on the decoded vector of the rib query item to obtain the predicted rib category corresponding to the rib query item. The bounding box prediction layer performs bounding box prediction based on the decoded vector of the rib query item to obtain the predicted pose parameters of the rib detection box corresponding to the rib query item in 3D space.

[0049] The predicted values ​​of the posture parameters include at least one of the predicted value of the center position of the detection frame, the predicted value of the scale of the detection frame, and the predicted value of the rotation angle of the detection frame.

[0050] In the local coordinate system, the 9-DOF detection frame can be represented by a parameter group of (x, y, z, w, h, d, α, β, γ). Among them, (x, y, z) are the coordinates of the center point of the detection frame on the X-axis, Y-axis, and Z-axis; w is the scale of the detection frame on the X-axis, h is the scale of the detection frame on the Y-axis, and d is the scale of the detection frame on the Z-axis; α is the rotation angle of the detection frame on the X-axis, β is the rotation angle of the detection frame on the Y-axis, and γ is the rotation angle of the detection frame on the Z-axis. Figure 2 As shown in the figure, the rib query module outputs the rib detection results corresponding to the rib query items, which can be input into the detection frame detection module (i.e. Figure 29-DOF detection frame detection module is shown in the figure). The detection frame detection module can draw a 9-DOF rib detection frame corresponding to each rib query item in the medical image based on the predicted values ​​of the detection frame's pose parameters. After determining the rib detection frame, a segmentation network can be used to perform rib segmentation. In this embodiment, to obtain more accurate rib segmentation results, a finer spatial resolution can be used to independently segment each rib within a locally cropped field of view (FOV). The following will be used to illustrate the rib detection results of any rib query.

[0051] For any rib detection result, the segmentation network can segment a sub-volume block from the medical image based on the predicted values ​​of the pose parameters of the detection box in the rib detection result in three-dimensional space. Optionally, the sub-volume block can be obtained from the original input medical image. The original input medical image has an original resolution, in which the pixel spacing is smaller (for example, 2mm, 1mm or even 0.5mm), which is conducive to obtaining more accurate segmentation results. After obtaining the sub-volume block, binary segmentation can be performed on the sub-volume block to obtain the segmentation results of the rib instance and non-rib area in the sub-volume block.

[0052] In this embodiment, the segmentation network may include multiple parallel rib segmentation heads. When the rib query module outputs rib detection results for multiple rib query items, multiple sub-volume blocks may be segmented from the original resolution medical image based on the rib detection results for the multiple rib query items. Binary segmentation may then be performed on the multiple sub-volume blocks in parallel using the multiple rib segmentation heads.

[0053] During binary segmentation, the rib region of interest is segmented as the foreground, while other tissue (including adjacent ribs) is segmented as the background. This segmentation result can be represented using a binary mask. In the mask, pixels with a value of 1 represent pixels in the rib instance, while pixels with a value of 0 represent pixels in the background. After obtaining the binary masks for multiple ribs, they are merged with the corresponding classification prediction labels and spatial positions to form the final instance segmentation result.

[0054] In this implementation, rib detection and segmentation based on the rib query term yields manageable rib detection and segmentation results. Furthermore, the use of a 9-degree-of-freedom (DOF) detection bounding box to estimate the location of rib instances allows for more accurate 3D rib segmentation and effectively addresses anomalies such as rib adhesion and structural damage.

[0055] In the above and following embodiments of the present application, based on the rib query item, a controllable rib segmentation result query function can be provided to the user.

[0056] In some optional embodiments, based on the rib detection results for each of the at least one rib query items, a rib instance is segmented from the medical image. After obtaining the image segmentation result, a user-input segmentation request for any rib instance may be received, a target rib query item corresponding to the rib instance may be determined, and a rib segmentation result corresponding to the target rib query item may be determined from the image segmentation result as the segmentation result for the rib instance. After determining the segmentation result for the rib instance, the segmentation result for the rib instance may be prominently displayed in the medical image.

[0057] For example, Figure 3 As shown, a user can enter "Query for the fifth rib on the right side." Based on the mapping relationship between the left and right rib order and the rib query item, the rib query item for the fifth rib on the right side can be determined to be "rib number 5." If the rib segmentation result corresponding to the at least one rib query item includes a rib segmentation result corresponding to rib number 5, the rib segmentation result corresponding to rib number 5 can be returned as the segmentation result corresponding to the fifth rib on the right side. In medical images, the segmentation result for the fifth rib on the right side can be highlighted or specially marked to make the query result more intuitive.

[0058] Based on this implementation, directional segmentation of ribs is achieved, and the specified rib segmentation results can be flexibly returned to the user according to user needs.

[0059] The above embodiment illustrates the rib detection and segmentation logic implemented by a neural network during the inference phase. In addition to the aforementioned rib detection method, the present embodiment also provides a method for training a neural network model. During training, the forward propagation process of the neural network model is as described in the aforementioned embodiment. Based on the rib detection method provided in the aforementioned embodiment, the neural network model can detect instance-level ribs from the input medical image to complete the forward propagation process, which will not be further described.

[0060] During the training phase, the neural network model can be optimized based on forward propagation, with the goal of reducing inference error. Specifically, after obtaining rib detection results for at least one rib query item, the neural network model can be further trained in a supervised manner using a preset supervisory signal.

[0061] Optionally, the rib category truth value and the rib detection frame pose parameter truth value corresponding to each of the at least one rib query items may be obtained. The rib category truth value may be pre-set. The rib detection frame pose parameter truth value may be calculated from a chest and abdomen scan image. An exemplary explanation is provided below.

[0062] After obtaining the three-view images obtained by scanning the chest and abdomen on the X, Y, and Z axes, three-dimensional space reconstruction can be performed using the three-view images to reconstruct the posture of the ribs in three-dimensional space and further visualize the reconstructed rib posture. Among them, rib instances can be annotated on the three-view images, and the annotated labels (such as color) of the same rib instances are the same. After the rib instances are reconstructed in three dimensions, the rib instances in three-dimensional space can be annotated using the same annotation method for rib instances on the three-view images, thereby visually displaying and distinguishing different rib instances.

[0063] Based on the 3D reconstructed and visualized rib instances, principal component analysis (PCA) is used to calculate a parametric representation of the rib instance, and the resulting parametric representation is used as the true pose parameter value of the rib detection box. PCA calculates the eigenvalues ​​and eigenvectors of the rib voxel coordinate covariance matrix, and sorts the eigenvalues ​​to identify the axes of the local coordinate system with corresponding eigenvectors. In the local coordinate system, a parameter set (x, y, z, w, h, d, α, β, γ) can be used to represent the 9-degree-of-freedom detection box.

[0064] During the training process, a bipartite matching method can be used to find a permutation of the rib query item to map the prediction result of the rib query item to the true value of the rib category, such as Figure 4 As shown. Figure 4 In the diagram, the matching relationship between the rib query item q and the rib category ground truth x can be learned during training to achieve a better match. To find a better match, a bipartite matching loss can be constructed to bind the rib query item to the rib category ground truth, thereby training the neural network model to find the region of interest for each rib query item.

[0065] Optionally, the matching loss of bipartite matching may include at least three parts: classification loss, bounding box prediction loss, and a penalty term. The classification loss may be determined based on the error between the true rib category value and the predicted rib category of each of the at least one rib instance. The bounding box prediction loss may be determined based on the error between the true pose parameter value and the predicted pose parameter value of the rib detection box of each of the at least one rib instance. When parameters with multiple degrees of freedom are used to describe the bounding box, the bounding box loss may include a multi-degree-of-freedom regression loss. The penalty term refers to the penalty term between the rib query item corresponding to each of the at least one rib instance and the true rib category value. During the training process, the penalty term can be configured to guide the rib query module to learn the matching relationship between the rib query item and the true rib category value. When the rib query item matches or is relatively close to the true rib category value specified by the user, the penalty term is small; otherwise, the penalty term is large.

[0066] In some optional embodiments, to facilitate the rapid acquisition of penalty items, indexes may be set for the rib query items and the rib category truth values, and a weighted adjacency matrix may be set based on the expected matching relationship between the rib query items and the rib category truth values. The penalty items may be obtained using the weighted adjacency matrix. If the rib query items with the same expected index are the best match to the rib category truth values, then one implementation of the weighted adjacency matrix may be as follows: Figure 5 As shown. In the weighted adjacency matrix, the indexes of the multiple rib query items are stored in rows, and the indexes of the multiple rib category true values ​​are stored in columns. When the index of the rib query item is the same as the index of the rib category true value, the penalty value is 0. The greater the difference between the index of the rib query item and the index of the rib category true value, the greater the penalty value corresponding to the intersection position of the index values. For example, when the index of the rib query item is 1 and the index of the rib category true value is 9, the penalty value is 8. When the index of the rib query item is 1 and the index of the rib category true value is 2, the penalty value is 1.

[0067] For any rib query item, the penalty value between the rib query item and the rib category truth value corresponding to the rib query item is queried from the preset weighted adjacency matrix. The rib category truth value corresponding to the rib query item refers to the inferred category truth value of the rib instance corresponding to the rib query item.

[0068] In each round of training, after determining the matching loss based on the above implementation, the assignment relationship between the at least one rib query item and the rib category true value can be updated according to the matching loss to learn the correspondence between the rib query item and different rib regions in the medical image.

[0069] The following examples will illustrate this.

[0070] Assume that a chest and abdomen CT scan contains N rib instances to be detected, where 1≤N≤C, where C is the maximum number of ribs in a normal chest and abdomen scan. Typically, C is set to 24. The rib set can be expressed as: Among them, x i represents the i-th rib instance, c i represents the true value of the rib category of the i-th rib instance, represents the true value of the center position of the detection box of the i-th rib instance, represents the scale truth value of the detection box of the i-th rib instance, Represents the true rotation angle of the detection box of the i-th rib instance.

[0071] The set of rib query items can be described as q = {q i,0≤i≤Q}, the total number of query items is set to C+1, Q=C+1. In the rib query module, the decoder has C+1 output channels, where channel 0 is used to output the background category, and channels 1 to C are used to output the classification results of C rib instances. Among them, the weighted adjacency matrix can be marked as M, M∈R (Q +1)×(Q+1) .

[0072] The index of the rib query item corresponding to the i-th rib instance is described as σ(i), and the detection result corresponding to the rib query item with index σ(i) is described as The rib query module is The matching loss on can be described as:

[0073]

[0074] Among them, c i is the true value of the rib category of the i-th rib instance, is the ith rib instance in c i Probabilistic predictions on classes, is the predicted value of the center position of the detection box corresponding to the rib query item σ(i), is the true value of the center position of the detection box corresponding to the i-th rib instance. is the predicted value of the scale of the detection box corresponding to the rib query item σ(i), is the true value of the scale of the detection box corresponding to the i-th rib instance. is the predicted value of the rotation angle of the detection box corresponding to the rib query item σ(i), is the true value of the rotation angle of the detection box corresponding to the i-th rib instance. i ] indicates that the rib query item σ(i) and the rib category truth value c of the i-th rib instance i The penalty term between λ and λ. C ,λ p ,λ s ,λ α ,λ m They are rib category, detection frame center position, detection frame scale, detection frame rotation angle and weighted coefficient of weighted adjacency matrix.

[0075] In the above embodiment, by adding a penalty term to the matching loss and using a weighted adjacency matrix to configure the size of the penalty value between the rib query term and the rib category true value, each rib query term can be manipulated to learn to predict rib instances of a specified category.

[0076] In each training round, the Hungarian algorithm can be used to find the best matching relationship between the rib query item and the true value of the rib category based on the above matching cost. Therefore, each rib query item can establish a correspondence with different rib regions on the medical image during the training process, and continuously learn the prediction task of the corresponding category. In the inference process after training, each rib query item can be inferred and predicted based on the learned correspondence and the prediction task of the corresponding category. During the training phase, the overall model loss function of the neural network model on N rib instances can be described as:

[0077]

[0078] Based on the above model loss function, the gradient descent method can be used to optimize the neural network model. When the above model loss function converges to a specified range, the training can be stopped and the trained neural network model can be output.

[0079] During the training process, matching loss was used to constrain the neural network model, allowing it to learn the intrinsic connection between rib query items and feature vectors from different regions of the medical image. This allowed the model to accurately identify rib instances from different regions of the medical image based on their characteristics.

[0080] In some scenarios, the rib detection methods for medical images provided in the aforementioned embodiments can be encapsulated as software tools that can be used by third parties, such as SaaS (Software-as-a-Service) tools. The SaaS tool can be implemented as a plug-in or an application. The plug-in or application can be deployed on a server and can open a specified interface to third-party users such as clients. For ease of description, in this embodiment, the specified interface is described as a first interface. Furthermore, third-party users such as clients can conveniently access and use the above-mentioned method provided by the server device by calling the first interface. The server can be a conventional server or a cloud server, and this embodiment does not impose any restrictions.

[0081] Taking the SaaS tool corresponding to the rib detection method of medical images as an example, the server can respond to the client's call request to the first interface, and obtain a three-dimensional medical image containing ribs from the interface parameters of the first interface; use the neural network model to feature encode different rib areas in the medical image to obtain multiple feature vectors; based on the correspondence between the learned rib query items and the different rib areas in the medical image, determine the feature vector corresponding to at least one rib query item; wherein, the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; the feature vectors corresponding to each of the at least one rib query items are decoded in parallel to obtain the rib detection results corresponding to each of the at least one rib query items; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space.

[0082] When providing rib segmentation service, the server may, after obtaining the rib detection results corresponding to each of the at least one rib query items, segment at least one rib instance from the medical image based on the detection results corresponding to each of the at least one rib query items, obtain an image segmentation result, and return the image segmentation result to the client for viewing.

[0083] In this embodiment, the server can provide the client with rib segmentation services in medical images based on the SaaS tool running thereon, thereby reducing the client's computing pressure and computing cost.

[0084] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0085] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The sequence numbers of the operations, such as 101, 102, etc., are merely used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0086] It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types.

[0087] In addition to the rib segmentation scenario, the segmentation method provided in the above and following embodiments of this application can also be extended to other scenarios to segment objects with a certain distribution pattern in space. For example, it can be applied to commodity segmentation scenarios in shelf images, etc., which is not limited in this embodiment. Figure 6 , a typical application scenario of the rib detection method of medical images provided in an embodiment of the present application is exemplified.

[0088] In a typical application scenario, the rib detection method for medical images provided in the embodiment of the present application can be applied to the rib detection and segmentation process in chest and abdominal radiography images. After obtaining the CT radiography images of the patient's chest and abdomen, the abdominal radiography images can be input into the electronic device. The electronic device can use a neural network model to perform feature encoding on the different rib regions in the medical image to obtain multiple feature vectors; based on the correspondence between the learned rib query items and the different rib regions in the medical image, determine the feature vector corresponding to each of the at least one rib query items; wherein, the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; the feature vectors corresponding to each of the at least one rib query items are decoded in parallel to obtain the rib detection results corresponding to each of the at least one rib query items; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space.

[0089] After obtaining the rib detection results corresponding to the at least one rib query item, the electronic device may segment at least one rib instance from the medical image based on the detection results corresponding to the at least one rib query item to obtain an image segmentation result. The at least one rib instance may be marked and distinguished using different colors.

[0090] In some embodiments, when a user provides rib search information, the electronic device can segment the rib instances specified by the user through a neural network model and display them, such as Figure 3 In other embodiments, when the user does not provide rib search information, the electronic device may segment all rib instances in the medical image by a neural network model by default and display them, such as Figure 6 The segmentation results of rib instances numbered 1 to 24 are shown.

[0091] In this embodiment, controllable rib instance segmentation is achieved based on the rib query item, and the rib segmentation accuracy in the three-dimensional space is improved based on the posture prediction value of the detection box in the three-dimensional space.

[0092] It is worth noting that the rib detection method of medical imaging provided in the aforementioned embodiments of this application can also be implemented by an AR (Augmented Reality) device. The AR device can be implemented as an AR head-mounted display device, or can be implemented as a smart terminal device equipped with an AR application, which is not limited in this embodiment. In remote diagnosis and consultation scenarios, virtual images and real images can be obtained through AR devices, and the virtual images and real images can be superimposed and displayed. In this embodiment, the virtual image can be obtained based on the three-dimensional medical image.

[0093] The AR device can obtain a three-dimensional medical image containing ribs, which can be input by the user or obtained by the AR device from the server, and this embodiment does not impose any restrictions. A neural network model for rib detection runs on the AR device. The neural network model can perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors; and determine the feature vector corresponding to at least one rib query item based on the correspondence between the learned rib query items and the different rib regions in the medical image; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected. The neural network model decodes the feature vectors corresponding to the at least one rib query item in parallel to obtain the rib detection results corresponding to the at least one rib query item; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space.

[0094] After obtaining the detection results output by the neural network model, the AR device can superimpose and display the rib detection results corresponding to the at least one rib query item in the captured real image, thereby realizing the superimposed display of the real image and the virtual rib detection results.

[0095] When displaying the rib detection results corresponding to the at least one rib query item, the AR device may reconstruct a three-dimensional detection box in virtual three-dimensional space based on the predicted pose parameters of the detection box corresponding to any rib query item in three-dimensional space, and display the reconstructed three-dimensional detection box. The corresponding rib category truth value may be marked near the three-dimensional detection box to facilitate viewing and distinguishing the detection results of different rib instances displayed in three-dimensional space.

[0096] In some embodiments, after the neural network model obtains the rib detection results corresponding to the at least one rib query item, it can perform a rib instance segmentation operation on the medical image to obtain at least one segmented rib instance. In this embodiment, because the rib detection results include the predicted values ​​of the pose parameters of the rib detection frame in three-dimensional space, a three-dimensional rib instance can be segmented from the three-dimensional medical image. When displaying the segmentation results, the AR device can display the segmentation results of the three-dimensional rib instance in a virtual three-dimensional space, thereby more clearly displaying the anatomical structure of the ribs.

[0097] In the above embodiment, based on the AR device, the instance-level rib detection or segmentation results can be displayed in three-dimensional space, and the instance-level rib detection or segmentation results can be linked with the real information captured in the real scene (such as the lesion site) to be displayed, which can supplement the real scene with richer and more accurate rib anatomical information.

[0098] Figure 7 is a structural diagram of an electronic device provided by an exemplary embodiment of the present application, such as Figure 7 As shown, the electronic device includes: a memory 701 and a processor 702.

[0099] The memory 701 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device.

[0100] The memory 701 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0101] Processor 702 is coupled to memory 701 and is configured to execute a computer program in memory 701, configured to: acquire a three-dimensional medical image containing ribs; utilize a neural network model to perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors; determine, based on the learned correspondence between rib query items and different rib regions in the medical image, a feature vector corresponding to each of at least one rib query items; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; and perform parallel decoding on the feature vectors corresponding to each of the at least one rib query items to obtain a rib detection result corresponding to each of the at least one rib query items; wherein each rib detection result includes: a predicted rib category and a predicted value of a pose parameter of a rib detection box in three-dimensional space.

[0102] Optionally, the at least one rib query item includes: a rib query item corresponding to a target rib input by a user; or a plurality of preset rib query items corresponding one-to-one to a plurality of ribs in the medical image.

[0103] Optionally, after the processor 702 performs parallel decoding on the feature vectors corresponding to the at least one rib query item to obtain the rib detection results corresponding to the at least one rib query item, it is further used to: segment at least one rib instance from the medical image according to the detection results corresponding to the at least one rib query item to obtain an image segmentation result.

[0104] Optionally, when the processor 702 segments at least one rib instance from the medical image based on the rib detection results corresponding to the at least one rib query item to obtain an image segmentation result, the processor 702 is specifically used to: for any rib detection result, segment a sub-volume block from the medical image based on the predicted value of the posture parameters of the detection box in the rib detection result in three-dimensional space; perform binary segmentation on the sub-volume block to obtain a segmentation result of the rib instance and the non-rib area in the sub-volume block.

[0105] Optionally, after the processor 702 segments at least one rib instance from the medical image based on the rib detection results corresponding to the at least one rib query item and obtains the image segmentation result, it is further used to: obtain a segmentation request for any rib instance input by the user, and determine the target rib query item corresponding to the rib instance; determine the rib segmentation result corresponding to the target rib query item from the image segmentation result as the segmentation result of the rib instance; and highlight the segmentation result of the rib instance in the medical image.

[0106] Optionally, when the processor 702 performs feature encoding on different rib regions in the medical image to obtain multiple feature vectors, it is specifically used to: adjust the pixel spacing in the medical image to obtain an isotropic medical image; use a segmentation network to extract a three-dimensional region of interest containing ribs from the isotropic medical image; perform feature extraction on the region of interest to obtain spatial features; and based on the self-attention mechanism, perform self-attention encoding on the spatial features to obtain feature vectors corresponding to the different rib regions.

[0107] Optionally, before extracting features from the region of interest and obtaining spatial features, the processor 702 is further configured to: perform at least one of random displacement, random scale transformation, random rotation, random cropping, and random erasing on the region of interest to perform data enhancement on the region of interest.

[0108] Optionally, when the processor 702 performs parallel decoding on the feature vectors corresponding to the at least one rib query item to obtain the rib detection results of the at least one rib query item, it is specifically used to: for any rib query item, decode the feature vector corresponding to the rib query item based on the attention mechanism to obtain a decoding vector; perform classification calculation according to the decoding vector to obtain a rib prediction category corresponding to the rib query item; and perform bounding box prediction according to the decoding vector to obtain a pose parameter prediction value of the rib detection box corresponding to the rib query item in three-dimensional space; wherein the pose parameter prediction value includes: at least one of a prediction value of the center position of the detection box, a prediction value of the scale of the detection box, and a prediction value of the rotation angle of the detection box.

[0109] Optionally, after performing parallel decoding on the feature vectors corresponding to the at least one rib query item and obtaining the rib detection results for the at least one rib query item, the processor 702 is further configured to: obtain the true value of the rib category and the true value of the pose parameters of the rib detection box corresponding to the at least one rib query item; determine a classification loss based on the error between the true value of the rib category and the rib prediction category of the at least one rib query item; determine a bounding box prediction loss based on the error between the true value of the pose parameters and the predicted value of the pose parameters of the rib detection box of the at least one rib query item; and determine a penalty term between the rib query item corresponding to each of the at least one rib instances and the true value of the rib category; determine a matching loss based on the classification loss, the bounding box prediction loss, and the penalty term; and update the assignment relationship between the at least one rib query item and the true value of the rib category based on the matching loss to learn the correspondence between the rib query item and different rib regions in the medical image.

[0110] Optionally, when determining the penalty term between the rib query item corresponding to each of the at least one rib instances and the label true value, the processor 702 is specifically used to: for any rib query item, query the penalty value between the rib query item and the rib category true value corresponding to the rib query item from a preset weighted adjacency matrix; wherein, in the weighted adjacency matrix, the indexes of multiple rib query items are stored in rows, and the indexes of multiple rib category true values ​​are stored in columns. The greater the difference between the index value of the rib query item and the index value of the rib category true value, the greater the penalty value corresponding to the intersection position of the index value.

[0111] Further, if Figure 7 As shown, the electronic device also includes: a communication component 703, a display 704, a power supply component 705 and other components. Figure 7 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 7 Components shown.

[0112] Among them, the communication component 703 is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0113] The display 704 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0114] The power supply component 705 provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0115] In this embodiment, a neural network model can be used to perform feature encoding on different rib regions in a medical image to obtain multiple feature vectors. Based on the learned correspondence between rib query items and different rib regions in the medical image, a feature vector corresponding to at least one rib query item can be determined. The feature vectors corresponding to each of the at least one rib query items are decoded in parallel to obtain a rib detection result corresponding to each of the at least one rib query items. The rib query items are endowed with anatomical semantic information, allowing them to focus on the features of different rib regions, thereby achieving controllable, instance-level rib detection. Furthermore, any rib detection result, including the predicted rib category and the predicted pose parameter value of the rib detection box in three-dimensional space, can accurately determine the location of the rib instance in three-dimensional space based on the predicted pose parameter value, thereby accurately achieving rib instance segmentation in three-dimensional medical images.

[0116] It should be noted that Figure 7The illustrated electronic device, in addition to being able to perform data processing operations according to the data processing logic described in the aforementioned embodiments, can also perform the following operations according to the rib detection method for medical images described below: the processor 702 is specifically used to: respond to a call request from a client to a first interface, and obtain a three-dimensional medical image containing ribs from the interface parameters of the first interface; use a neural network model to feature encode different rib regions in the medical image to obtain multiple feature vectors; determine the feature vector corresponding to each of the at least one rib query items based on the correspondence between the learned rib query items and the different rib regions in the medical image; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; parallel decode the feature vectors corresponding to each of the at least one rib query items to obtain the rib detection results corresponding to each of the at least one rib query items; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection box in three-dimensional space.

[0117] The electronic device may return the rib detection results corresponding to each of the at least one rib query items to the client. In some embodiments, after obtaining the rib detection results corresponding to each of the at least one rib query items, the electronic device may further segment at least one rib instance from the medical image based on the detection results corresponding to each of the at least one rib query items, obtain an image segmentation result, and return the image segmentation result to the client. This will not be further described.

[0118] It is also worth noting that, in some embodiments, Figure 7 The electronic device shown in the figure can be realized as an AR device, except Figure 7 In addition to the components shown in the figure, the electronic device may also include an image acquisition device (not shown). The image acquisition device is used to capture real images. The processor 702 is specifically used to: obtain a three-dimensional medical image containing ribs; use a neural network model to perform feature encoding on different rib regions in the medical image to obtain multiple feature vectors; determine the feature vector corresponding to at least one rib query item based on the correspondence between the learned rib query items and the different rib regions in the medical image; wherein the at least one rib query item corresponds one-to-one to at least one rib instance to be detected; perform parallel decoding on the feature vectors corresponding to the at least one rib query item to obtain the rib detection result corresponding to the at least one rib query item; any rib detection result includes: a rib prediction category and a predicted value of the posture parameters of the rib detection frame in three-dimensional space; and, through the display 704, superimpose and display the rib detection result corresponding to the at least one rib query item on the real image captured by the image acquisition device.

[0119] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which can implement the steps in the above method embodiment when executed by a processor.

[0120] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable electronic device to produce a machine, so that the instructions executed by the processor of the computer or other programmable electronic device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable electronic device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable electronic device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0124] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0125] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0126] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0127] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0128] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A rib detection method for medical imaging, applied to AR equipment, characterized in that: include: Acquire three-dimensional medical images including ribs; Using a neural network model, feature encoding is performed on different rib regions in the medical image to obtain multiple feature vectors; Determining a feature vector corresponding to at least one rib query item based on the learned correspondence between the rib query item and different rib regions in the medical image; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; Parallel decoding is performed on the feature vectors corresponding to the at least one rib query item to obtain a rib detection result corresponding to the at least one rib query item; each rib detection result includes: a rib prediction category and a predicted value of a pose parameter of a rib detection box in three-dimensional space; In the captured real image, superimposing and displaying the rib detection results corresponding to each of the at least one rib query items; Among them, the matching loss of the neural network model during the training process includes: classification loss, bounding box prediction loss and penalty items; the method for obtaining the penalty items includes: for any rib query item, querying the penalty value between the rib query item and the rib category true value corresponding to the rib query item from a preset weighted adjacency matrix; wherein, in the weighted adjacency matrix, the indexes of multiple rib query items are stored in rows, and the indexes of multiple rib category true values ​​are stored in columns. The greater the difference between the index value of the rib query item and the index value of the rib category true value, the greater the penalty value corresponding to the intersection position of the index value.

2. A rib detection method for medical images, characterized in that: include: Acquire three-dimensional medical images including ribs; Using a neural network model, feature encoding is performed on different rib regions in the medical image to obtain multiple feature vectors; Determining a feature vector corresponding to at least one rib query item based on the learned correspondence between the rib query item and different rib regions in the medical image; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; Parallel decoding is performed on the feature vectors corresponding to the at least one rib query item to obtain a rib detection result corresponding to the at least one rib query item; each rib detection result includes: a rib prediction category and a predicted value of a pose parameter of a rib detection box in three-dimensional space; Among them, the matching loss of the neural network model during the training process includes: classification loss, bounding box prediction loss and penalty items; the method for obtaining the penalty items includes: for any rib query item, querying the penalty value between the rib query item and the rib category true value corresponding to the rib query item from a preset weighted adjacency matrix; wherein, in the weighted adjacency matrix, the indexes of multiple rib query items are stored in rows, and the indexes of multiple rib category true values ​​are stored in columns. The greater the difference between the index value of the rib query item and the index value of the rib category true value, the greater the penalty value corresponding to the intersection position of the index value.

3. The method according to claim 2, characterized in that The at least one rib query item includes: a rib query item corresponding to a target rib input by a user; or a plurality of preset rib query items corresponding one-to-one to a plurality of ribs in the medical image.

4. The method according to claim 2, characterized in that After decoding the feature vectors corresponding to the at least one rib query item in parallel to obtain the rib detection results corresponding to the at least one rib query item, the method further includes: At least one rib instance is segmented from the medical image according to the detection results corresponding to each of the at least one rib query items to obtain an image segmentation result.

5. The method according to claim 4, characterized in that Segmenting at least one rib instance from the medical image according to the rib detection results corresponding to each of the at least one rib query items to obtain an image segmentation result, including: For any rib detection result, segmenting a sub-volume block from the medical image according to a predicted value of a pose parameter of a detection box in the rib detection result in three-dimensional space; Binary segmentation is performed on the sub-volume block to obtain a segmentation result of the rib instance and the non-rib area in the sub-volume block.

6. The method according to claim 4, characterized in that Segmenting at least one rib instance from the medical image according to the rib detection results corresponding to the at least one rib query item, and obtaining the image segmentation result, further comprising: Obtaining a segmentation request for any rib instance input by a user, and determining a target rib query item corresponding to the rib instance; Determining a rib segmentation result corresponding to the target rib query item from the image segmentation results as a segmentation result of the rib instance; In the medical image, the segmentation result of the rib instance is highlighted.

7. The method according to claim 2, characterized in that Feature encoding is performed on different rib regions in the medical image to obtain multiple feature vectors, including: Adjusting the pixel spacing in the medical image to obtain an isotropic medical image; A segmentation network is used to extract a three-dimensional region of interest including the ribs from the isotropic medical image; Extracting features from the region of interest to obtain spatial features; Based on the self-attention mechanism, the spatial features are self-attention encoded to obtain feature vectors corresponding to the different rib regions.

8. The method according to claim 7, characterized in that Before extracting features from the region of interest to obtain spatial features, the method further includes: At least one of random displacement, random scale transformation, random rotation, random cropping, and random erasure is performed on the region of interest to perform data enhancement on the region of interest.

9. The method according to claim 2, characterized in that Parallel decoding of the feature vectors corresponding to the at least one rib query item to obtain a rib detection result for each of the at least one rib query item includes: For any rib query item, decode the feature vector corresponding to the rib query item based on the attention mechanism to obtain a decoding vector; Perform classification calculation based on the decoding vector to obtain a rib prediction category corresponding to the rib query item; and Bounding box prediction is performed based on the decoded vector to obtain a pose parameter prediction value of the rib detection box corresponding to the rib query item in three-dimensional space; wherein the pose parameter prediction value includes: a prediction value of the center position of the detection box, a prediction value of the scale of the detection box, and a prediction value of the rotation angle of the detection box.

10. The method according to any one of claims 2 to 9, characterized in that: After decoding the feature vectors corresponding to the at least one rib query item in parallel to obtain the rib detection results for the at least one rib query item, the method further includes: Obtaining a true value of a rib category and a true value of a pose parameter of a rib detection frame corresponding to each of the at least one rib query items; determining a classification loss based on an error between a true rib category value of each of the at least one rib query items and a predicted rib category; determining a bounding box prediction loss based on an error between a true value of a pose parameter of a rib detection box and a predicted value of the pose parameter of each of the at least one rib query items; and Determining a penalty term between a rib query item corresponding to each of the at least one rib instance and a true value of the rib category; Determine a matching loss based on the classification loss, the bounding box prediction loss, and the penalty term; According to the matching loss, the assignment relationship between the at least one rib query item and the rib category true value is updated to learn the correspondence between the rib query item and different rib regions in the medical image.

11. A rib detection method for medical imaging, characterized in that: include: In response to a call request to the first interface from the client, obtaining a three-dimensional medical image containing ribs from interface parameters of the first interface; Using a neural network model, feature encoding is performed on different rib regions in the medical image to obtain a plurality of feature vectors; and based on the learned correspondence between the rib query items and the different rib regions in the medical image, a feature vector corresponding to each of at least one rib query items is determined; wherein the at least one rib query item has a one-to-one correspondence with at least one rib instance to be detected; Parallel decoding is performed on the feature vectors corresponding to the at least one rib query item to obtain a rib detection result corresponding to the at least one rib query item; each rib detection result includes: a rib prediction category and a predicted value of a pose parameter of a rib detection box in three-dimensional space; Among them, the matching loss of the neural network model during the training process includes: classification loss, bounding box prediction loss and penalty items; the method for obtaining the penalty items includes: for any rib query item, querying the penalty value between the rib query item and the rib category true value corresponding to the rib query item from a preset weighted adjacency matrix; wherein, in the weighted adjacency matrix, the indexes of multiple rib query items are stored in rows, and the indexes of multiple rib category true values ​​are stored in columns. The greater the difference between the index value of the rib query item and the index value of the rib category true value, the greater the penalty value corresponding to the intersection position of the index value.

12. An electronic device, characterized in that: include: memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute the one or more computer instructions to perform the steps of the method according to any one of claims 1 to 11.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of the method according to any one of claims 1 to 11 can be implemented.

Citation Information

Patent Citations

  • Multi-modal medical image multi-organ positioning method based on one-to-one target query Transform

    CN114359642A