Image Matching Network Model Optimization Method, Patient Body Surface ROI Recognition Method and Device
By optimizing the image matching network model, the optimal image matching network model is generated, which solves the problems of low accuracy and long time-consuming ROI recognition in the prior art, and realizes high-precision and efficient automatic recognition of ROI on the surface of the patient, supporting the accurate execution of radiation therapy.
Patent Information
- Application Number
- CN202311671128.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-12-06
AI Technical Summary
The neural network model for radiation therapy in the prior art is low in accuracy and long-term in identifying the patient's surface ROI, and cannot meet the high accuracy and efficiency needs of radiation therapy.
By acquiring the patient's image pairs, forming an image data set, and performing multiple iterative optimization training on the image matching network model, the feature extraction module and feature matching module, including self-attention unit and cross-attention unit, optimize the representation state and matching degree of feature points, generate the optimal image matching network model, calculate the homography matrix, and realize automatic recognition of the patient's body surface ROI.
It improves the recognition accuracy and recognition speed of the patient's surface ROI, reduces the deviation caused by manual outline, and ensures the accuracy and efficiency of radiation therapy.
Smart Images

Figure CN117830610B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to an image matching network model optimization method, a patient body surface ROI recognition method and a device. Background Art
[0002] Radiotherapy, as one of the more advanced methods of tumor treatment, occupies an important position in the comprehensive treatment of tumors. Tumor radiotherapy mainly generates radiation through linear accelerators, and uses the radiation to irradiate the tumor target area to kill cancer cells and achieve the purpose of treatment. Because respiratory movement during radiotherapy can cause the tumor target area to move, the therapeutic radiation cannot accurately irradiate the target area, but increases the risk of irradiating normal tissues and organs, causing irreversible damage to the patient's body. Therefore, how to achieve precise radiotherapy has become a key research issue in the field of modern radiotherapy.
[0003] There are currently two methods to assist in achieving precise radiotherapy. The first method is to align the patient's scanned cone beam CT (CBCT) image with the planned CT (Computed Tomography) image to obtain the target area posture deviation and correct it. The defects of this method are: it causes radiation harm to the patient and cannot dynamically monitor and correct the target area motion deviation. The second method is to use the structural optical body surface tracking system to obtain the patient's body surface contour in real time through the optical body surface system and monitor the movement of the body surface contour ROI (Region Of Interest) with the strongest correlation with the target area. Once the body surface contour ROI motion monitoring value exceeds the threshold, the accelerator interlock will be automatically triggered to stop the beam, thereby reducing the irradiation of normal tissues and organs. This method has the advantage of no radiation for patients, and the method can indirectly monitor the changes in the target area posture by real-time monitoring of the body surface contour ROI movement. After collecting the body surface contour and selecting the ROI in the CT scanning room, in clinical practice, the radiotherapy process still relies on the therapist to manually draw the ROI that is as consistent as possible with the ROI selected in the CT room based on naked eye observation and experience. This manual delineation will have a large deviation, which brings uncertainty to the motion tracking of the patient's 3D surface ROI. With the development of machine learning technology, there are also solutions based on machine learning technology to identify the patient's body surface ROI to avoid the uncertainty caused by manual ROI delineation.
[0004] A patient motion tracking system configured to automatically generate regions of interest (ROIs), with publication number CN112132860A and publication date December 25, 2020, uses an ROI generation processor configured to output an ROI-marked 3D surface to a display and a motion tracking module using stored ROI descriptive data and the 3D surface to identify the patient's body surface ROI. The 3D surface with ROI markings is used by the motion tracking module to track the patient's motion during positioning and / or treatment of the patient in a treatment room. Among them, the ROI generation processor generates the 3D surface with ROI markings through an ROI model configured as a trained convolutional neural network. However, the convolutional neural network model used in this method has low accuracy in identifying ROI markings and takes a long time, which is not conducive to the high requirements for ROI identification accuracy and speed in scenarios such as radiotherapy. Summary of the Invention
[0005] To overcome the problems in the related art, the present invention provides an image matching network model optimization method, a method and device for identifying a patient's body surface ROI, to solve the problems of low accuracy and long time consumption in identifying ROI markings using the neural network model in the prior art.
[0006] According to a first aspect of the present invention, there is provided an image matching network model optimization method, the method comprising:
[0007] Obtain a plurality of pairs of images of patients to form an image data set; wherein, the pair of images includes a first image of the patient during an examination process and a second image during a treatment process;
[0008] Input the image data set into the image matching network model for multiple rounds of iterative optimization training to obtain an optimal image matching network model; the image matching network model is used to perform feature matching on the first image and the second image in the pair of images to generate a set of matching point pairs for the pair of images.
[0009] Preferably, the image matching network model includes a feature extraction module and a feature matching module;
[0010] The feature extraction module is used to extract a first set of feature points of the first image and a second set of feature points of the second image respectively;
[0011] The feature matching module is used to calculate the similarity and matching degree between each feature point in the first set of feature points and the second set of feature points, and generate a set of matching point pairs.
[0012] Preferably, the feature matching module includes a plurality of operation layers; each operation layer includes an update module and a prediction module;
[0013] The updating module is used to update the representation state of each feature point in the first feature point set and the second feature point set;
[0014] The prediction module is used to calculate the similarity and matching degree between each feature point according to the representation state of each feature point in the first feature point set and the second feature point set, and generate a set of matching point pairs.
[0015] Preferably, the updating module comprises two self-attention units corresponding to the first feature point set and the second feature point set respectively, and a cross-attention unit;
[0016] The self-attention unit and the cross-attention unit are used to aggregate the image information through a multi-layer perceptron to update the representation state of each feature point:
[0017]
[0018] in, represents the representation state of feature point i in image I, Indicates the image aggregation message originating from image S, [.|.] indicates and To overlay;
[0019] Wherein, the image aggregation message It is calculated by the weighted average of the representation state of each feature point j in the image S produced by the attention mechanism:
[0020]
[0021] Where W is the projection matrix, represents the attention score between feature points i and j in images I and S;
[0022] In the self-attention unit, each feature point i in the image I pays attention to all feature points in the same image I. At this time, the image S = I, and the representation state x of each feature point i is transformed into i Decompose into key vector k i and the query vector q i , and calculate the attention score between feature points i and j as:
[0023]
[0024] in, is the rotation encoding of the relative position between feature points, p i and p i are the positions of feature points i and j respectively;
[0025] In the cross-attention unit, each feature point i in the image I attends to all feature points in another image S, and the representation state x of each feature point i is calculated through a linear transformation. i The key vector k of i , and the attention score between the feature points i and j in the images I and S is calculated as:
[0026]
[0027] where only one similarity calculation is performed between the message from image S to image I and the message from image I to image S.
[0028] Preferably, the prediction module is used to:
[0029] Calculate the similarity between each feature point in the two images:
[0030]
[0031] where and are the representation states of the feature points i and j in the images A and B respectively, represents the pairwise similarity score matrix between the feature points i and j, and M and N are the numbers of feature points in the images A and B respectively;
[0032] Calculate the matching degree of each feature point in each image:
[0033] σ i = Sigmoid(Linear(x i )) ∈ [0, 1]
[0034] where σ i represents the matching score of the feature point i; and when the feature point cannot be detected in another image, the feature point is marked as an unmatched point, and at this time σ i → 0;
[0035] According to the similarity and the matching degree, calculate and generate the assignment matrix between each feature point in the two images:
[0036]
[0037] where P ij represents the assignment matrix between the feature points i and j in the images A and B;
[0038] When both feature point i and feature point j are not non - matchable points, and the similarity between feature point i and feature point j is higher than the similarity between feature point i or feature point j and any other feature points in image B or image A, the feature point i and the feature point j are formed into a matching point pair;
[0039] Among all the matching point pairs, select all the assignment matrices P ij greater than the preset assignment threshold τ and greater than any other elements in its row and column, and the matching point pairs form the set of matching point pairs.
[0040] Preferably, each of the operation layers further includes a classifier module; the classifier module is used for:
[0041] Calculate the confidence of each feature point in the first feature point set and the second feature point set through a multi - layer perceptron:
[0042] c i = Sigmoid(MLP(x i )) ∈ [0, 1]
[0043] where x i is the representation state of feature point i, and c i is the confidence of feature point i;
[0044] Judge whether the proportion of the confidence of all feature points in each operation layer that exceeds the specified layer confidence standard is greater than the preset credible proportion threshold, and stop the inference of the set of matching points when the proportion is greater than the credible proportion threshold:
[0045]
[0046] where, is the confidence of feature point i in image A or image B, N and M are the numbers of feature points in image A and image B respectively, α is the credible proportion threshold, λ l is the layer confidence standard for the corresponding operation layer l, and the layer confidence standard λ l decreases layer by layer according to the verification accuracy of each classifier module;
[0047] If the proportion is not greater than the credible proportion threshold, discard all the feature points in the operation layer whose confidence exceeds the layer confidence standard and whose matching degree is non - matchable, and pass the remaining feature points to the next operation layer.
[0048] According to the second aspect of the present invention, there is provided a method for identifying the ROI on the patient's body surface, the method including:
[0049] Obtain the first image and the first three-dimensional body surface contour of the patient in the CT scanning room, as well as the second image and the second three-dimensional body surface contour in the accelerator treatment room;
[0050] Select a first body surface ROI on the first three-dimensional body surface contour and map the first body surface ROI onto the first image to generate a first image ROI;
[0051] Input the first image and the second image into an image matching network model obtained by training according to the image matching network model optimization method of any embodiment of the present invention to generate a set of matching point pairs, and calculate a homography transformation matrix from the first image to the second image through the set of matching point pairs;
[0052] Transform the first image ROI onto the second image through the homography transformation matrix to obtain a second image ROI, and map the second image ROI onto the second three-dimensional body surface contour to generate a second body surface ROI.
[0053] Preferably, the obtaining of the first image and the first three-dimensional body surface contour of the patient during the examination process, as well as the second image and the second three-dimensional body surface contour of the patient during the treatment process, includes:
[0054] Obtain the first image and the first three-dimensional body surface contour of the patient in the CT scanning room, as well as the second image and the second three-dimensional body surface contour of the patient in the accelerator treatment room.
[0055] Preferably, the mapping of the first body surface ROI onto the first image to generate a first image ROI includes:
[0056] Read the camera parameter file to obtain the internal parameter matrix M 3×4 and the external parameter matrix N 4×4 ;
[0057] According to the calculation formula for converting world coordinates to pixel coordinates:
[0058]
[0059]
[0060] where X w 、Y w and Z w are respectively the x, y, and z-axis coordinates of the three-dimensional point in the world coordinate system; X c 、Y c and Z c are respectively the x, y, and z-axis coordinates of the three-dimensional point after being transformed into the camera coordinate system; x and y are the x and y-axis coordinates of the three-dimensional point after being transformed into the pixel coordinate system;
[0061] Convert each three-dimensional point of the first three-dimensional body surface contour in the world coordinate system to the first image through the above formula to generate corresponding two-dimensional pixel points, where the closed area formed by the pixel points projected by all three-dimensional points in the first body surface ROI is the first image ROI.
[0062] Preferably, calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs includes:
[0063] Apply each matching point pair (P, Q) in the set of matching point pairs to the homography transformation formula to obtain the homography transformation matrix from the first image to the second image:
[0064]
[0065]
[0066] where H is the homography transformation matrix.
[0067] Preferably, calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs includes:
[0068] Obtain all matching point pairs corresponding to the feature points of the first image within the range of the first image ROI in the set of matching point pairs to form a set of ROI matching point pairs;
[0069] Calculate the homography transformation matrix from the first image to the second image through the set of ROI matching point pairs.
[0070] According to the third aspect of the present invention, there is provided a patient body surface ROI recognition device, the device includes:
[0071] A data acquisition module for acquiring the first image and the first three-dimensional body surface contour of the patient during the examination process, and the second image and the second three-dimensional body surface contour of the patient during the treatment process;
[0072] An ROI selection module for selecting a first body surface ROI on the first three-dimensional body surface contour and mapping the first body surface ROI to the first image to generate a first image ROI;
[0073] An image matching module for inputting the first image and the second image into an image matching network model trained by the image matching network model optimization method according to any embodiment of the present invention to generate a set of matching point pairs, and calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs;
[0074] The ROI recognition module is used to transform the first image ROI onto the second image through the homography transformation matrix, obtain the second image ROI, and map the second image ROI onto the second three-dimensional body surface contour to generate the second body surface ROI.
[0075] The present invention discloses an image matching network model optimization method, a patient body surface ROI recognition method and device. By obtaining patient image pairs to optimize and train the image matching network model, an optimal image matching network model is obtained. And in sequence, the first body surface ROI selected on the first three-dimensional body surface contour obtained during the examination is mapped onto the first image obtained during the examination to generate the first image ROI; then, through the homography transformation matrix calculated from the set of matching point pairs generated by inputting the first image and the second image obtained during the treatment process into the optimal image matching network model obtained by training, the first image ROI is transformed into the second image to generate the second image ROI; then the second image ROI is mapped into the second three-dimensional body surface contour obtained during the treatment process to generate the second body surface ROI, thereby realizing that when the patient is in the treatment process, the ROI corresponding to the ROI selected during the examination can be automatically recognized on the three-dimensional body surface of the patient, avoiding the possible deviation of manual delineation and reducing the risk brought to the motion tracking of the 3D surface ROI of the patient. Moreover, the present invention uses the optimal image matching network model obtained by optimization training to match the first image and the second image in the patient body surface ROI recognition process, making the final body surface ROI recognition result more accurate and less time-consuming.
[0076] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings
[0077] Figure 1 is a flowchart of an image matching network model optimization method shown according to an embodiment of the present invention.
[0078] Figure 2 is a schematic structural diagram of an image matching network model shown according to an embodiment of the present invention.
[0079] Figure 3 is a flowchart of a patient body surface ROI recognition method shown according to an embodiment of the present invention.
[0080] Figure 4a is a comparison schematic diagram of the automatic recognition results of the image ROI of the patient's left abdomen shown according to an embodiment of the present invention.
[0081] Figure 4bIt is a comparison schematic diagram of the automatic recognition results of the ROI of an image of a patient's right abdomen shown according to an embodiment of the present invention.
[0082] Figure 4c It is a comparison schematic diagram of the automatic recognition results of the ROI of an image of a patient's lower abdomen shown according to an embodiment of the present invention.
[0083] Figure 5a It is a comparison schematic diagram of the automatic recognition results of the surface ROI of a patient's left abdomen shown according to an embodiment of the present invention.
[0084] Figure 5b It is a comparison schematic diagram of the automatic recognition results of the surface ROI of a patient's right abdomen shown according to an embodiment of the present invention.
[0085] Figure 5c It is a comparison schematic diagram of the automatic recognition results of the surface ROI of a patient's lower abdomen shown according to an embodiment of the present invention.
[0086] Figure 6 It is a schematic structural diagram of an image matching network model optimization device shown according to an embodiment of the present invention.
[0087] Figure 7 It is a schematic structural diagram of a patient surface ROI recognition device shown according to an embodiment of the present invention.
[0088] Figure 8 It is a schematic structural diagram of a computer device shown according to an embodiment of the present invention. Detailed implementation manners
[0089] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0090] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0091] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0092] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0093] As Figure 1 shown Figure 1 is a flowchart of an optimization method for an image matching network model according to an embodiment of the present invention, including the following steps:
[0094] Step S101, obtaining image pairs of a number of patients to form an image data set; wherein, the image pairs include a first image during the examination process of the patient and a second image during the treatment process;
[0095] Step S102, inputting the image data set into the image matching network model for multi-round iterative optimization training to obtain an optimal image matching network model; the image matching network model is used to perform feature matching on the first image and the second image in the image pair to generate a set of matching point pairs of the image pair.
[0096] In the present invention, the examination process may include the stage of physical examination of the patient before treatment, the stage of planning the treatment plan for the patient, and any other preparatory stage before treatment, which is not limited in the present invention; while the treatment process includes any implementation stage of treating the patient, such as the process of radiotherapy of a patient with tumor radiotherapy in the accelerator treatment room, which is not limited in the present invention.
[0097] In step S101, the first image of the patient during the examination process may be the first image obtained by the patient in the CT scanning room, and the second image of the patient during the treatment process may be the second image obtained by the patient in the accelerator treatment room. In other embodiments, the patient may also obtain the first image and the second image in other ways and in other scenarios, which is not limited in the present invention.
[0098] In step S102, the image matching network model may be iteratively trained by using an image data set composed of image pairs of multiple different patients to obtain an image matching network model with the best accuracy.
[0099] Specifically, as Figure 2 shown Figure 2It is a schematic structural diagram of an image matching network model shown according to an embodiment of the present invention. The image matching network model used in step S302 may include a feature extraction module and a feature matching module. Among them, the feature extraction module may be used to extract a first feature point set of a first image and a second feature point set of a second image respectively. The feature matching module may be used to calculate the similarity and matching degree between each feature point in the first feature point set and the second feature point set, and generate a set of matching point pairs.
[0100] Specifically, the above-mentioned feature extraction module may use any existing image feature extraction method, such as SIFT (Scale Invariant Feature Transform), SuperPoint or other image feature extraction algorithms, to extract the first feature point set of the first image and the second feature point set of the second image. The present invention places no restrictions on this.
[0101] Specifically, the first feature set and the second feature set include a number of feature points of the first image and the second image. Each feature point is jointly composed of a set of feature point positions p and related image feature descriptors d to form a local feature (p, d). Among them, the feature point position p is composed of an x coordinate, a y coordinate and a detection confidence c, that is, p i =(x, y, c); and the image feature descriptor d is obtained by the feature extraction algorithm.
[0102] Specifically, the above-mentioned feature matching module may be used to predict a partial matching relationship between the local feature sets extracted from the first image and the second image, the first feature set and the second feature set.
[0103] Specifically, the above-mentioned feature matching module may include several operation layers. Each operation layer may include an update module and a prediction module. Among them, the update module may be used to update the representation state of each feature point in the first feature point set and the second feature point set. The prediction module may be used to calculate the similarity and matching degree between each feature point according to the representation state of each feature point in the first feature point set and the second feature point set, and generate a set of matching point pairs.
[0104] Specifically, the number of operation layers in the feature matching module may be 9 layers, or may be set to other layers according to actual needs, so that the computational amount can be relatively small while the feature matching is more sufficient, that is, to achieve a balance between measurement accuracy and matching time consumption.
[0105] Specifically, the representation state of each feature point in the first feature point set and the second feature point set may be obtained from the image feature descriptor d obtained in the feature extraction module:
[0106] x i =di
[0107] Specifically, the above update module may include two self-attention units corresponding to the first feature point set and the second feature point set respectively, and a cross-attention unit; wherein, the self-attention unit and the cross-attention unit can be used to update the representation state of each feature point by aggregating image messages through a multi-layer perceptron (MLP):
[0108]
[0109] Wherein, represents the representation state of feature point i in image I, represents the image aggregation message originating from image S, and [.|.] represents the superposition of and ;
[0110] Wherein, the image aggregation message can be calculated as the weighted average of the representation states of each feature point j in image S produced by the attention mechanism:
[0111]
[0112] Wherein, W is the projection matrix, represents the attention score between feature points i and j in images I and S;
[0113] In the self-attention unit, each feature point i in image I attends to all feature points in the same image I. At this time, image S = I, and the representation state x of each feature point i is decomposed into a key vector k i and a query vector q i through a linear transformation, and the attention score between feature points i and j is calculated as: i
[0114]
[0115] Wherein, is the rotation encoding of the relative position between feature points, and p i and p i are the positions of feature points i and j respectively;
[0116] In the cross-attention unit, each feature point i in image I attends to all feature points in another image S, and the key vector k of the representation state x of each feature point i is calculated through a linear transformation i i , but do not calculate the query vector, and calculate the attention score between the feature points i and j in images I and S as follows:
[0117]
[0118] Among them, the similarity is calculated only once between the message from image S to image I and the message from image I to image S.
[0119] In the self-attention unit, the rotation encoding R is used to define the attention score a between the feature points i and j ij , to capture the relative position relationship between them. By dividing the space into 2D subspaces and performing rotational projection onto the learnable basis vector b k , the position encoding is achieved. And the rotation encoding R enables the model to retrieve the point j with the learned relative position. This encoding is the same in all operation layers and is calculated only once and cached. While in the cross-attention unit, since the relative position has no meaning between images, there is no need to add position information.
[0120] Specifically, the above prediction module can be used for:
[0121] First, calculate the similarity between each feature point in the two images:
[0122]
[0123] Among them, and are the representation states of the feature points i and j in images A and B respectively, represents the pairwise similarity score matrix between the feature points i and j, M and N are the numbers of feature points in images A and B respectively; the pairwise similarity score matrix S ij represents the affinity of each pair of feature points to become corresponding relationships.
[0124] Then calculate the matchability of each feature point in each image:
[0125] σ i = Sigmoid(Linear(x i )) ∈ [0, 1]
[0126] Among them, σ i represents the matching score of the feature point i; and when the feature point cannot be detected in another image, the feature point is marked as a non-matchable point, and at this time σ i →0;
[0127] Then, according to the similarity and matching degree, calculate the assignment matrix between each feature point in the two images:
[0128]
[0129] Among them, P ij represents the assignment matrix between feature points i and j in image A and image B;
[0130] When both feature point i and feature point j are not non - matchable points, and the similarity between feature point i and feature point j is higher than the similarity between feature point i or feature point j and any other feature point in image B or image A, form a matching point pair for feature point i and feature point j; that is, if and only if both points are predicted to be matchable and their similarity is higher than any other points in the two images, the point pair (i, j) will have a corresponding relationship;
[0131] Then, among all the matching point pairs, select all the matching point pairs where the assignment matrix P ij is greater than the preset assignment threshold τ and greater than any other element in its row and column to form a set of matching point pairs.
[0132] Specifically, the assignment threshold τ can be set according to actual needs to adjust the matching accuracy of the model, and the present invention does not limit this.
[0133] Specifically, each operation layer can also include a classifier module; this classifier module can be used to predict the allocation confidence of the update state of any given layer and help decide whether to stop the inference process; when the image pair is easy to match, if the prediction results of the early layer are the same as those of the later layer and have high confidence, these prediction results can be output in advance and the inference can be stopped, thus avoiding unnecessary calculations; but if only a few points have high confidence, the inference process can continue to the next operation layer, but the calculation amount of the subsequent operation layer can be reduced by discarding (pruning) the feature points with low confidence and non - matchable ones.
[0134] Specifically, this classifier module can be used for:
[0135] Calculate the confidence of each feature point in the first feature point set and the second feature point set through a multi - layer perceptron:
[0136] c i = Sigmoid(MLP(x i )) ∈ [0, 1]
[0137] Among them, x i is the representation state of feature point i, and c i is the confidence of feature point i;
[0138] Determine whether the proportion of all feature points in each operation layer with confidence exceeding the specified layer confidence criterion is greater than the preset reliable proportion threshold, and stop inferring the set of matching points when the proportion is greater than the reliable proportion threshold:
[0139]
[0140] Wherein, is the confidence of feature point i in image A or image B, N and M are the numbers of feature points in image A and image B respectively, α is the reliable proportion threshold, and λ l is the layer confidence criterion for the corresponding operation layer l, and the layer confidence criterion λ l decreases layer by layer according to the verification accuracy of each classifier module;
[0141] If the proportion is not greater than the reliable proportion threshold, considering that the points predicted to be both reliable and non-matchable are likely to be of no help to the matching of other points in subsequent levels, and these points are usually located in the regions that are significantly invisible in the image. Therefore, all feature points with confidence exceeding the layer confidence criterion and non-matchable matching degree can be discarded in each operation layer, and the remaining feature points are passed to the next operation layer. In this way, the computational amount can be significantly reduced, and due to the quadratic complexity of the attention unit in the update module, these discarded feature points will not affect the accuracy of image matching.
[0142] Specifically, the above-mentioned reliable proportion threshold α and the layer confidence criterion λ corresponding to each operation layer l can be set continuously according to actual needs, and the present invention does not limit this.
[0143] Such as Figure 3 shown, Figure 3 is a flowchart of a method for identifying the ROI on the patient's body surface according to an embodiment of the present invention, including the following steps:
[0144] Step S301, obtain the first image and the first three-dimensional body surface contour of the patient during the examination process, and the second image and the second three-dimensional body surface contour of the patient during the treatment process;
[0145] Step S302, select the first body surface ROI on the first three-dimensional body surface contour, and map the first body surface ROI to the first image to generate the first image ROI;
[0146] Step S303, input the first image and the second image into the image matching network model trained by the method for optimizing the image matching network model according to any embodiment of the present invention to generate a set of matching point pairs, and calculate the homography transformation matrix from the first image to the second image through the set of matching point pairs;
[0147] In step S304, the first image ROI is transformed onto the second image through a homography transformation matrix to obtain a second image ROI, and the second image ROI is mapped onto the second three-dimensional body surface contour to generate a second body surface ROI.
[0148] In step S301, the first image and the first three-dimensional body surface contour of the patient obtained during the examination process are a single still image and a three-dimensional body surface contour, which are used as a reference standard for image matching and ROI selection and should contain the complete information of the patient in the standard state. The second image and the second three-dimensional body surface contour of the patient obtained during the treatment process can be either a single image and a three-dimensional body surface contour or a corresponding pair in a continuous series of images and three-dimensional body surface contours. The present invention places no restrictions thereon.
[0149] Similarly, in step S301, the examination process may include the stage of physically examining the patient before treatment, the stage of planning the treatment plan for the patient, and any other preparatory stage before treatment. The present invention places no restrictions thereon; while the treatment process includes any implementation stage of treating the patient, such as the process of a patient receiving radiotherapy in an accelerator treatment room. The present invention places no restrictions thereon.
[0150] Specifically, in step S301, the first image and the first three-dimensional body surface contour of the patient obtained during the examination process may be the first image and the first three-dimensional body surface contour obtained by the patient in a CT scanning room, while the second image and the second three-dimensional body surface contour of the patient obtained during the treatment process may be the second image and the second three-dimensional body surface contour obtained by the patient in an accelerator treatment room. In other embodiments, the patient may also obtain the first image, the first three-dimensional body surface contour, the second image, and the second three-dimensional body surface contour in other ways and in other scenarios. The present invention places no restrictions thereon.
[0151] Specifically, in step S301, in some embodiments, a binocular structured light acquisition system can be used to obtain the patient's image and three-dimensional body surface contour. Among them, the binocular structured light acquisition system can include one binocular structured light camera directly above the treatment couch where the patient lies in the CT scanning room or the accelerator treatment room, or multiple binocular structured light cameras forming a certain angle above the treatment couch. The image of the patient obtained by the binocular structured light acquisition system is the image collected and corrected by the binocular cameras, and the three-dimensional body surface contour is the three-dimensional point cloud of the patient's body surface collected by the binocular cameras scanning the patient. Specifically, the placement position of the binocular cameras can also be set according to the actual scenario as long as it can completely scan and obtain the patient's image and three-dimensional body surface contour, and the present invention does not limit this; in some embodiments, other image and three-dimensional contour acquisition systems or devices that can be used to obtain the patient's image and three-dimensional body surface contour can also be used to obtain the patient's image and three-dimensional body surface contour simultaneously or separately. For example, a conventional camera can be used to obtain the patient's image, and a three-dimensional body sensing photography device such as Kinect can be used to collect the patient's three-dimensional body surface contour, and the present invention does not limit this.
[0152] In step S302, the most suitable body surface contour ROI, that is, the first body surface ROI, can be selected on the first three-dimensional body surface contour. The first body surface ROI is a part of the point cloud area in the first three-dimensional body surface contour. Specifically, the first three-dimensional body surface contour can be displayed on the display interface of the specified device, and then the therapist manually controls the mouse to drag and draw the ROI through the ROI drawing tool according to the actual situation of the patient to obtain the first body surface ROI. Specifically, the first body surface ROI can also be obtained by other means, such as artificial intelligence selection, etc., and the present invention does not limit this. Specifically, the first body surface ROI can be circular, square or of any shape, and the first body surface ROI can be of any size, and the present invention does not limit this.
[0153] After the first body surface ROI is selected, according to the mapping relationship between the first three-dimensional body surface contour and the first image, the first body surface ROI can be projected onto the first image to obtain the first image ROI.
[0154] Specifically, when mapping the first body surface ROI onto the first image to generate the first image ROI in step S302, it may include:
[0155] Read the camera parameter file to obtain the internal parameter matrix M of the camera 3×4 and the external parameter matrix N 4×4 ;
[0156] According to the calculation formula for converting world coordinates to pixel coordinates:
[0157]
[0158]
[0159] Among them, X w , Y w and Z w are the x, y, and z-axis coordinates of the three-dimensional point in the world coordinate system respectively; X c , Y c and Z c are the x, y, and z-axis coordinates of the three-dimensional point after being transformed into the camera coordinate system respectively; x and y are the x and y-axis coordinates of the three-dimensional point after being transformed into the pixel coordinate system;
[0160] Each three-dimensional point of the first three-dimensional body surface contour in the world coordinate system is transformed onto the first image through the above formula to generate corresponding two-dimensional pixel points. Among them, the closed area composed of the pixel points projected by all the three-dimensional points in the first body surface ROI is the first image ROI.
[0161] In step S303, the image pair composed of the first image and the second image obtained in step S301 is input into a pre-trained image matching network model for inference and recognition. The image matching network model can infer and recognize the matching point pairs in the first image and the second image, and generate a set of matching point pairs.
[0162] Specifically, the image matching network model optimization method described in any embodiment of the present invention can be used to obtain the optimal image matching network model, and then the optimal image matching network model is used to perform image matching on the first image and the second image of the patient obtained in step S301 to obtain the set of matching point pairs of the first image and the second image.
[0163] Specifically, in step S303, the process of calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs may include:
[0164] Applying each matching point pair (P, Q) in the set of matching point pairs to the homography transformation formula to obtain the homography transformation matrix from the first image to the second image:
[0165]
[0166]
[0167] Among them, H is the homography transformation matrix.
[0168] Specifically, in order to further improve the accuracy of ROI recognition, in step S303, the homography transformation matrix from the first image to the second image is obtained by calculating the matching point pair set, which may also include: first obtaining all matching point pairs corresponding to the feature points of the first image within the range of the first image ROI in the matching point pair set to form a ROI matching point pair set; and then calculating the homography transformation matrix from the first image to the second image by the ROI matching point pair set. That is, when calculating the homography transformation matrix from the first image to the second image, only the matching point pairs whose corresponding positions in the matching point pair set are within the ROI area are used for calculation, thereby avoiding the influence of other areas on the human body that are not selected by the ROI on the homography transformation matrix, so that the accuracy of the second image ROI obtained by transforming the homography transformation matrix is higher, thereby making the accuracy of the second body surface ROI finally identified higher. Specifically, the process of calculating the homography transformation matrix from the first image to the second image by the ROI matching point pair set can refer to the aforementioned process of calculating the homography transformation matrix from the first image to the second image by the matching point pair set.
[0169] In step S304, the second image ROI on the second image can be identified using the homography transformation matrix obtained in step S303, and finally the second body surface ROI in the second three-dimensional body surface contour can be identified based on the mapping relationship between the second image and the second three-dimensional body surface contour, and the second body surface ROI is a part of the point cloud area in the second three-dimensional body surface contour. The implementation process of converting the second image ROI into the second body surface ROI can refer to the process of converting the first body surface ROI into the first image ROI, and will not be repeated again.
[0170] Specifically, the first body surface ROI selected on the first three-dimensional body surface contour can be the target area targeted by radiotherapy in the patient's breathing state, and the second body surface ROI finally inferred from the first body surface ROI on the first three-dimensional body surface contour is also the real-time position of the target area targeted by radiotherapy in the patient's breathing state in the accelerator treatment room, so that the radiation treatment process in the accelerator treatment room can be guided according to the second body surface ROI. Specifically, the first body surface ROI and the second body surface ROI can also be a smaller area than the target area of treatment to improve the fault tolerance of the second body surface ROI in treatment.
[0171] Specifically, in order to determine the recognition effect of the body surface ROI, it can be quantified by two indicators: the position error and the time consumption of the body surface ROI.
[0172] Specifically, first, the first three-dimensional body surface contour A and the first image a of the patient in the CT scan room, and the second three-dimensional body surface contour B and the second image b of the patient in the accelerator treatment room can be obtained respectively. Then, a three-dimensional marker point P is selected on the first three-dimensional body surface contour A (keeping the marker point in place) as the center point of the first body surface ROI for determining the first three-dimensional body surface contour A. Next, the coordinates Q(x, y, z) of this marker point on the second three-dimensional body surface contour B are obtained. This coordinate point is the physical coordinate of the second body surface ROI on the second three-dimensional body surface contour B. The second body surface ROI on the second three-dimensional body surface contour B is identified using the above patient body surface ROI recognition method, and the center point coordinates Q2(x2, y2, z2) of this second body surface ROI are obtained. Then, the position error R of the body surface ROI recognition can be calculated according to the following formula:
[0173] R = sqrt((x2 - x) 2 +(y2 - y) 2 +(z2 - z) 2 )
[0174] Specifically, when identifying the body surface ROIs of n different patients, the i-th recognition error is R i , then its average error can be calculated according to the following formula:
[0175]
[0176] The standard deviation s of its recognition error r can be calculated according to the following formula:
[0177]
[0178] Specifically, when recording the time T taken for the i-th body surface ROI recognition in the body surface ROI recognition of n different patients i , the average time taken is and the standard deviation s of the time taken t is calculated according to the following formula:
[0179]
[0180]
[0181] Specifically, the present invention also verifies the patient body surface ROI recognition method of the present invention through an experiment. The hardware used in this experiment is an Intel(R) Xeon(R) Silver 4210 CPU @ 2.20GHz 2.19GHz (2 processors), RAM 128GB, NVIDIA RTX 4000, Windows 10 (64-bit). Among them, the image matching network model used in the present invention is trained on the Pytorch 1.12.1 (GPU) framework, and three body surface ROI recognition methods are integrated and implemented on VS2019, including the body surface ROI recognition method based on ICP (Iterative Closest Point) and the body surface ROI recognition method based on SIFT, as well as the patient body surface ROI recognition method proposed by the present invention. The data used in the experiment is the test data collected by simulating the radiotherapy clinical scenario in the radiotherapy laboratory. Among them, 50 males and 50 females of the yellow race are obtained respectively. For males: age (40 ± 10) years old, height (175 ± 15) cm, weight (65 ± 15) kg; for females: age (40 ± 10) years old, height (165 ± 10) cm, weight (60 ± 15) kg. Then, the obtained body surface contour data (including the body surface contours of the CT scanning room and the accelerator treatment room) are tested and verified using the above three recognition methods. The recognition error and time consumption of each body surface ROI are calculated respectively according to the above formula, and a statistical analysis is performed on the 100 body surface recognition results. The statistical results are shown in Table 1:
[0182] Table 1 Comparison of 100 sets of test data by three different recognition methods
[0183]
[0184] The results show that the patient body surface ROI recognition method proposed by the present invention has the following characteristics: 1) The average value of the position error is the smallest, which is 5% of the ICP registration recognition method, significantly smaller than the ICP and SIFT methods, verifying its high recognition accuracy; 2) The standard deviation of the position error is the smallest, verifying its best recognition stability; 3) The average time consumption is the least, only 11% of the SIFT method, significantly smaller than the ICP and SIFT methods, verifying its best recognition efficiency. In summary, the patient body surface ROI recognition method proposed by the present invention has better performance in terms of accuracy, stability and efficiency.
[0185] In addition, the present invention also shows the body surface ROI results recognized by the present invention through an embodiment. As Figures 4a - 4c and Figures 5a - 5c shown, the present invention respectively shows the comparison diagrams of the automatic recognition results of three groups of image ROIs and the automatic recognition results of three groups of body surface ROIs.
[0186] Among them, Figure 4a is a comparison schematic diagram of the automatic recognition results of the image ROI of a patient's left abdomen shown according to an embodiment of the present invention. Among them, Figure 4a On the left is the first image ROI on the recognized patient's left abdomen in the first image obtained in the CT scanning room; on the right is the second image ROI on the recognized patient's left abdomen in the second image obtained in the accelerator treatment room.
[0187] Figure 4b is a comparison schematic diagram of the automatic recognition results of the image ROI of a patient's right abdomen shown according to an embodiment of the present invention. Among them, Figure 4b On the left is the first image ROI on the recognized patient's right abdomen in the first image obtained in the CT scanning room; on the right is the second image ROI on the recognized patient's right abdomen in the second image obtained in the accelerator treatment room.
[0188] Figure 4c is a comparison schematic diagram of the automatic recognition results of the image ROI of a patient's lower abdomen shown according to an embodiment of the present invention. Among them, Figure 4c On the left is the first image ROI on the recognized patient's lower abdomen in the first image obtained in the CT scanning room; on the right is the second image ROI on the recognized patient's lower abdomen in the second image obtained in the accelerator treatment room.
[0189] From Figures 4a - 4c it can be seen that the accelerator treatment room can accurately identify the same image ROI as selected in the CT scanning room.
[0190] Figure 5a is a comparison schematic diagram of the automatic recognition results of the body surface ROI of a patient's left abdomen shown according to an embodiment of the present invention. Among them, Figure 5a On the left is the first body surface ROI on the recognized patient's left abdomen in the first three-dimensional body surface contour obtained in the CT scanning room; on the right is the second body surface ROI on the recognized patient's left abdomen in the second three-dimensional body surface contour obtained in the accelerator treatment room.
[0191] Figure 5b is a comparison schematic diagram of the automatic recognition results of the body surface ROI of a patient's right abdomen shown according to an embodiment of the present invention. Among them, Figure 5b On the left is the first body surface ROI on the recognized patient's right abdomen in the first three-dimensional body surface contour obtained in the CT scanning room; on the right is the second body surface ROI on the recognized patient's right abdomen in the second three-dimensional body surface contour obtained in the accelerator treatment room.
[0192] Figure 5c is a comparison schematic diagram of the automatic recognition results of the body surface ROI of a patient's lower abdomen shown according to an embodiment of the present invention. Among them,Figure 5c On the left is the first surface ROI on the identified lower abdomen of the patient on the first three-dimensional body surface contour obtained in the CT scan room; on the right is the second surface ROI on the identified lower abdomen of the patient on the second three-dimensional body surface contour obtained in the accelerator treatment room.
[0193] From Figures 5a - 5c It can be seen that the accelerator treatment room can also accurately identify the same surface ROI as selected in the CT scan room.
[0194] Specifically, after obtaining the second surface ROI in step S304, the above method can also perform motion tracking on the second surface ROI through a motion tracking system. For example, it can track the undulating motion of the patient's surface caused by breathing, and monitor the treatment situation by tracking this motion. For example, if the motion is too large, the emission of radiation for treatment will be stopped to avoid unexpected damage to the patient. The method for identifying the patient's surface ROI described in the present invention is a means to assist motion tracking, and the identified ROI can be used to monitor the motion of the ROI area, making it more meaningful in treatment. The experimental results show that the method described in the present invention has good recognition performance, providing strong technical support for accurately identifying the surface ROI. The proposed method can assist in surface tracking in tumor radiotherapy, improve the consistency from planning (CT scan room) to implementation (accelerator treatment room), and achieve a more accurate radiotherapy effect.
[0195] In recent years, optical surface monitoring systems have gradually been used as image guidance systems in clinical radiotherapy. The motion of tumors in the chest and abdomen has a strong correlation with the surface ROI. Through the surface breathing motion and related ROIs, the breathing amplitude can be monitored and the real-time breathing state can be tracked to judge the approximate motion of the tumor. By optically monitoring the patient's ROI, the setup accuracy can be improved, thereby reducing the number of other image verifications such as X-rays or CBCT, reducing the additional medical radiation of the patient, and improving the comfort of the patient. For example, in the whole breast radiotherapy after breast-conserving surgery for left breast cancer, by monitoring the patient's ROI and cooperating with the deep inspiration breath-holding technique, the protection of the patient's heart can be achieved. An increase in the heart dose will significantly increase the probability of complications and mortality.
[0196] Corresponding to the embodiment of the method for optimizing the image matching network model described above, the present invention also provides an apparatus for optimizing the image matching network model.
[0197] As Figure 6 shown, Figure 6 is a schematic structural diagram of an apparatus for optimizing an image matching network model according to an embodiment of the present invention, including the following modules:
[0198] The image dataset acquisition module 610 is used to obtain image pairs of several patients to form an image dataset; among them, the image pair includes a first image during the examination process and a second image during the treatment process of the patient.
[0199] The image matching network model training module 620 is used to input the image dataset into the image matching network model for multi-round iterative optimization training to obtain the optimal image matching network model; the image matching network model is used to perform feature matching on the first image and the second image in the image pair to generate a set of matching point pairs for the image pair.
[0200] Specifically, the above-mentioned image matching network model may include a feature extraction module and a feature matching module; among them, the feature extraction module may be used to extract a first feature point set of the first image and a second feature point set of the second image respectively; the feature matching module may be used to calculate the similarity and matching degree between each feature point in the first feature point set and the second feature point set, and generate a set of matching point pairs.
[0201] Specifically, the above-mentioned feature matching module may include several operation layers; each operation layer may include an update module and a prediction module; among them, the update module may be used to update the representation state of each feature point in the first feature point set and the second feature point set; the prediction module may be used to calculate the similarity and matching degree between each feature point according to the representation state of each feature point in the first feature point set and the second feature point set, and generate a set of matching point pairs.
[0202] Specifically, the above-mentioned update module may include two self-attention units corresponding to the first feature point set and the second feature point set respectively, and a cross-attention unit; among them, the self-attention unit and the cross-attention unit may be used to update the representation state of each feature point by aggregating image messages through a multi-layer perceptron:
[0203]
[0204] Among them, represents the representation state of the feature point i in the image I, represents the image aggregation message from the image S, and [.|.] represents the superposition of and ;
[0205] Among them, the image aggregation message can be calculated by the weighted average of the representation states of each feature point j in the image S produced by the attention mechanism:
[0206]
[0207] Among them, W is the projection matrix, represents the attention score between feature points i and j in images I and S;
[0208] In the self-attention unit, each feature point i in image I pays attention to all feature points in the same image I. At this time, image S = I, and the representation state x of each feature point i is transformed into i Decompose into key vector k i and the query vector q i , and calculate the attention score between feature points i and j as:
[0209]
[0210] in, is the rotation encoding of the relative position between feature points, p i and p i are the positions of feature points i and j respectively;
[0211] In the cross attention unit, each feature point i in image I pays attention to all feature points in another image S, and the representation state x of each feature point i is calculated by linear transformation i The key vector k i , and calculate the attention score between feature points i and j in images I and S as:
[0212]
[0213] The similarity between the message from image S to image I and the message from image I to image S is calculated only once.
[0214] Specifically, the above prediction module can be used for:
[0215] First, calculate the similarity between each feature point in the two images:
[0216]
[0217] in, and are the representation states of feature points i and j in image A and image B respectively, represents the pairwise similarity score matrix between feature points i and j, where M and N are the number of feature points in image A and image B respectively;
[0218] Then calculate the matching degree of each feature point in each image:
[0219] σ i =Sigmoid(Linear(x i ))∈[0,1]
[0220] Among them, σ iRepresents the matching score of feature point i; and when the feature point cannot be detected in another image, mark the feature point as an unmatched point, at this time σ i →0;
[0221] Then, according to the similarity and matching degree, calculate the assignment matrix between each feature point in the two images:
[0222]
[0223] Among them, P ij Represents the assignment matrix between feature points i and j in image A and image B;
[0224] When both feature point i and feature point j are not unmatched points, and the similarity between feature point i and feature point j is higher than the similarity between feature point i or feature point j and any other feature point in image B or image A, form a matching point pair for feature point i and feature point j;
[0225] Then, among all the matching point pairs, select all the matching point pairs where the assignment matrix P ij is greater than the preset assignment threshold τ and greater than any other element in its row and column to form a set of matching point pairs.
[0226] Specifically, each operation layer may further include a classifier module; the classifier module can be used for:
[0227] Calculate the confidence of each feature point in the first feature point set and the second feature point set through a multi-layer perceptron:
[0228] c i = Sigmoid(MLP(x i )) ∈ [0, 1]
[0229] Among them, x i is the representation state of feature point i, and c i is the confidence of feature point i;
[0230] Judge whether the proportion of the confidence of all feature points in each operation layer that exceeds the specified layer confidence standard is greater than the preset credible proportion threshold, and stop the inference of the matching point set when the proportion is greater than the credible proportion threshold:
[0231]
[0232] Among them, is the confidence of feature point i in image A or image B, N and M are the numbers of feature points in image A and image B respectively, α is the credible proportion threshold, and λ l is the layer confidence standard corresponding to operation layer l, and the layer confidence standard λ lDecrease layer by layer according to the verification accuracy of each classifier module;
[0233] If the ratio is not greater than the credible ratio threshold, discard all feature points in the operation layer whose confidence exceeds the layer confidence standard and whose matching degree is non-matchable, and pass the remaining feature points to the next operation layer.
[0234] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, and will not be elaborated here.
[0235] Corresponding to the embodiment of the patient body surface ROI recognition method described above, the present invention also provides a patient body surface ROI recognition device.
[0236] Such as Figure 7 shown, Figure 7 is a schematic structural diagram of a patient body surface ROI recognition device shown by the present invention according to an embodiment, including the following modules:
[0237] A data acquisition module 710, configured to acquire a first image and a first three-dimensional body surface contour of a patient during an examination process, and a second image and a second three-dimensional body surface contour of the patient during a treatment process;
[0238] An ROI selection module 720, configured to select a first body surface ROI on the first three-dimensional body surface contour, and map the first body surface ROI to the first image to generate a first image ROI;
[0239] An image matching module 730, configured to input the first image and the second image into an image matching network model obtained by training according to the image matching network model optimization method described in any embodiment of the present invention, generate a set of matching point pairs, and calculate a homography transformation matrix from the first image to the second image through the set of matching point pairs;
[0240] An ROI recognition module 740, configured to transform the first image ROI to the second image through the homography transformation matrix to obtain a second image ROI, and map the second image ROI to the second three-dimensional body surface contour to generate a second body surface ROI.
[0241] Specifically, when acquiring the first image and the first three-dimensional body surface contour of the patient during the examination process, and the second image and the second three-dimensional body surface contour of the patient during the treatment process in the data acquisition module 710, it may include:
[0242] Acquire a first image and a first three-dimensional body surface contour of the patient in a CT scanning room, and a second image and a second three-dimensional body surface contour of the patient in an accelerator treatment room.
[0243] Specifically, when mapping the first body surface ROI to the first image in the ROI selection module 720 to generate the first image ROI, it may include:
[0244] Read the camera parameter file to obtain the internal parameter matrix M of the camera 3×4 and the external parameter matrix N 4×4 ;
[0245] According to the calculation formula for converting world coordinates to pixel coordinates:
[0246]
[0247]
[0248] where X w 、Y w and Z w are the x, y, and z-axis coordinates of the three-dimensional point in the world coordinate system respectively; X c 、Y c and Z c are the x, y, and z-axis coordinates of the three-dimensional point converted to the camera coordinate system respectively; x and y are the x and y-axis coordinates of the three-dimensional point converted to the pixel coordinate system;
[0249] Convert each three-dimensional point of the first three-dimensional body surface contour in the world coordinate system to the first image through the above formula to generate corresponding two-dimensional pixel points, where the closed area composed of the pixel points projected by all the three-dimensional points in the first body surface ROI is the first image ROI.
[0250] Specifically, when calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs in the image matching module 730, it may include:
[0251] Apply each matching point pair (P, Q) in the set of matching point pairs to the homography transformation formula to obtain the homography transformation matrix from the first image to the second image:
[0252]
[0253]
[0254] where H is the homography transformation matrix.
[0255] Specifically, when calculating the homography transformation matrix from the first image to the second image through the set of matching point pairs in the image matching module 730, it may also include:
[0256] Obtain the matching point pairs within the range of the first image ROI for all the feature points corresponding to the first image in the set of matching point pairs, and form a set of ROI matching point pairs;
[0257] The homography transformation matrix from the first image to the second image is calculated through the set of ROI matching point pairs.
[0258] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0259] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0260] The present invention also provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements any one of the image matching network model optimization method and the patient body surface ROI recognition method described in any of the foregoing embodiments, or can simultaneously implement the image matching network model optimization method and the patient body surface ROI recognition method described in any of the foregoing embodiments.
[0261] Figure 8 Figure 14 shows a more specific schematic diagram of the hardware structure of the computing device provided by the present invention. The device may include: a processor 801, a memory 802, an input / output interface 803, a communication interface 804, and a bus 805. Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.
[0262] The processor 801 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided by the present invention. The processor 801 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0263] The memory 802 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 802 can store the operating system and other application programs. When implementing the technical solutions provided by the present invention through software or firmware, the relevant program codes are stored in the memory 802 and are called and executed by the processor 801.
[0264] The input / output interface 803 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0265] The communication interface 804 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or can achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0266] The bus 805 includes a path for transmitting information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804).
[0267] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image matching network model optimization method or the patient body surface ROI recognition method described in any one of the foregoing embodiments.
[0268] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media such as modulated data signals and carrier waves.
[0269] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0270] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0271] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solution of the present invention, the functions of the modules can be realized in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0272] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for identifying the ROI on the patient's body surface, characterized in that, The method includes: Obtaining a first image and a first three-dimensional body surface contour of the patient during the examination process, and a second image and a second three-dimensional body surface contour of the patient during the treatment process; Selecting a first body surface ROI on the first three-dimensional body surface contour and mapping the first body surface ROI to the first image to generate a first image ROI, including: Read the camera parameter file to obtain the internal parameter matrix M of the camera 3×4 and the external parameter matrix N 4×4 ; According to the calculation formula for converting world coordinates to pixel coordinates: Among them, X w , Y w and Z w are the x, y, and z-axis coordinates of the three-dimensional point in the world coordinate system respectively; X c , Y c and Z c are the x, y, and z-axis coordinates of the three-dimensional point after being transformed into the camera coordinate system respectively; x and y are the x and y-axis coordinates of the three-dimensional point after being transformed into the pixel coordinate system; Converting each three-dimensional point of the first three-dimensional body surface contour in the world coordinate system to the first image through the above formula to generate corresponding two-dimensional pixel points, where the closed area formed by the pixel points projected by all three-dimensional points in the first body surface ROI is the first image ROI; Inputting the first image and the second image into an optimal image matching network model to generate a set of matching point pairs, and calculating and obtaining a homography transformation matrix from the first image to the second image through the set of matching point pairs; Transforming the first image ROI to the second image through the homography transformation matrix to obtain a second image ROI, and mapping the second image ROI to the second three-dimensional body surface contour to generate a second body surface ROI.
2. The method according to claim 1, wherein The method for obtaining the optimal image matching network model includes: Obtaining a set of image pairs of several patients to form an image data set; where the image pair includes a first image of the patient during the examination process and a second image during the treatment process; Inputting the image data set into the image matching network model for multiple rounds of iterative optimization training to obtain an optimal image matching network model; the image matching network model is used to perform feature matching on the first image and the second image in the image pair to generate a set of matching point pairs of the image pair.
3. The method according to claim 2, wherein The image matching network model includes a feature extraction module and a feature matching module; The feature extraction module is used to extract a first set of feature points of the first image and a second set of feature points of the second image respectively; The feature matching module is used to calculate the similarity and matching degree between each feature point in the first set of feature points and the second set of feature points, and generate a set of matching point pairs.
4. The method according to claim 3, characterized in that The feature matching module includes several operation layers; each operation layer includes an update module and a prediction module; The update module is used to update the representation state of each feature point in the first set of feature points and the second set of feature points; The prediction module is used to calculate the similarity and matching degree between each feature point according to the representation state of each feature point in the first set of feature points and the second set of feature points, and generate a set of matching point pairs.
5. The method according to claim 4, characterized in that, The update module includes two self-attention units corresponding to the first set of feature points and the second set of feature points respectively, and a cross-attention unit; The self-attention unit and the cross-attention unit are used to update the representation state of each feature point by aggregating image messages through a multi-layer perceptron: Among them, represents the representation state of feature point i in image I, represents the image aggregation message originating from image S, and [.|.] represents the and are superimposed; Among them, the image aggregation message is calculated by the weighted average of the representation states of each feature point j in the image S produced by the attention mechanism: where W is the projection matrix, represents the attention score between the feature points i and j in the images I and S; In the self-attention unit, each feature point i in the image I attends to all feature points in the same image I. At this time, the image S = I, and the representation state x of each feature point i is decomposed into a key vector k i and a query vector q i through a linear transformation, and the attention score between feature points i and j is calculated as: i Among them, is the rotation encoding of the relative position between feature points, p i and p i are the positions of feature points i and j respectively; In the cross-attention unit, each feature point i in the image I attends to all feature points in another image S, and the representation state x of each feature point i is calculated through a linear transformation i for the key vector k i , and the attention score between the feature points i and j in the images I and S is calculated as follows: Among them, the similarity is only calculated once between the message from image S to image I and the message from image I to image S.
6. The method according to claim 4, characterized in that The prediction module is used for: Calculating the similarity between each feature point in the two images: Among them, and are the representation states of feature points i and j in images A and B respectively, represents the pairwise similarity score matrix between feature points i and j, and M and N are the numbers of feature points in the image A and the image B respectively; Calculate the matching degree of each feature point in each image: σ i = Sigmoid(Linear(x i )) ∈ [0, 1] Among them, σ i represents the matching score of feature point i; and when the feature point cannot be detected in another image, the feature point is marked as a non-matchable point, at this time σ i → 0; Calculate and generate an assignment matrix between each feature point in the two images according to the similarity and the matching degree: where P ij represents the assignment matrix between the feature points i and j in the image A and the image B; When feature point i and feature point j are both non-unmatchable points, and the similarity between feature point i and feature point j is higher than the similarity between feature point i or feature point j and any other feature point in image B or image A, form a matching point pair of the feature point i and the feature point j; Among all pairs of matching points, select all the assignment matrices P ij such that they are greater than a preset assignment threshold τ and greater than any other elements in their respective rows and columns. These pairs of matching points form the set of pairs of matching points.
7. The method according to claim 4, characterized in that Each of the operation layers further includes a classifier module; the classifier module is used for: Calculate the confidence of each feature point in the first feature point set and the second feature point set through a multi-layer perceptron: c i = Sigmoid(MLP(x i )) ∈ [0, 1] where x i is the representation state of feature point i, and c i is the confidence of feature point i; Judge whether the proportion of all feature points in each operation layer whose confidence exceeds the specified layer confidence standard is greater than the preset credible proportion threshold, and stop reasoning about the matching point set when the proportion is greater than the credible proportion threshold: wherein, is the confidence of feature point i in image A or image B, N and M are the numbers of feature points in image A and image B respectively, and α is the credible ratio threshold, is the layer confidence criterion for the corresponding operation layer , and the layer confidence criterion decreases layer by layer according to the verification accuracy of each classifier module; If the proportion is not greater than the credible proportion threshold, discard all feature points in the operation layer whose confidence exceeds the layer confidence standard and whose matching degree is unmatchable, and pass the remaining feature points to the next operation layer.
8. The method according to claim 1, wherein The obtaining the first image and the first three-dimensional body surface contour of the patient during the examination process, and the second image and the second three-dimensional body surface contour of the patient during the treatment process includes: Obtain the first image and the first three-dimensional body surface contour of the patient in the CT scanning room, and the second image and the second three-dimensional body surface contour of the patient in the accelerator treatment room.
9. The method according to claim 1, characterized in that, The obtaining the homography transformation matrix from the first image to the second image by calculating through the set of matching point pairs includes: Apply each matching point pair (P, Q) in the set of matching point pairs to the homography transformation formula to obtain the homography transformation matrix from the first image to the second image: Where, H is the homography transformation matrix.
10. The method according to claim 1, wherein The obtaining the homography transformation matrix from the first image to the second image by calculating through the set of matching point pairs includes: Obtain all matching point pairs corresponding to the feature points of the first image within the range of the first image ROI in the set of matching point pairs to form a set of ROI matching point pairs; Calculate and obtain the homography transformation matrix from the first image to the second image through the set of ROI matching point pairs.
11. A patient body surface ROI recognition device, characterized in that, The device includes: A data acquisition module, configured to acquire the first image and the first three-dimensional body surface contour of the patient during the examination process, and the second image and the second three-dimensional body surface contour of the patient during the treatment process; An ROI selection module, configured to select a first body surface ROI on the first three-dimensional body surface contour and map the first body surface ROI to the first image to generate a first image ROI, including: Read the camera parameter file to obtain the internal parameter matrix M of the camera 3×4 and the external parameter matrix N 4×4 ; According to the calculation formula for converting world coordinates to pixel coordinates: wherein, X w , Y w and Z w are respectively the x, y, and z axis coordinates of the three-dimensional point in the world coordinate system; X c , Y c and Z c are respectively the x, y, and z axis coordinates of the three-dimensional point after being transformed into the camera coordinate system; x and y are the x and y axis coordinates of the three-dimensional point after being transformed into the pixel coordinate system; Convert each three-dimensional point of the first three-dimensional body surface contour in the world coordinate system to the first image through the above formula to generate corresponding two-dimensional pixel points, where the closed area formed by the pixel points projected by all three-dimensional points in the first body surface ROI is the first image ROI; An image matching module, configured to input the first image and the second image into an optimal image matching network model, generate a set of matching point pairs, and calculate a homography transformation matrix from the first image to the second image through the set of matching point pairs; An ROI recognition module, configured to transform the first image ROI onto the second image through the homography transformation matrix to obtain a second image ROI, and map the second image ROI onto the second three-dimensional body surface contour to generate a second body surface ROI.
Citation Information
Patent Citations
Patient motion tracking system configured for automatic roi generation
CN112132860A
Method used for registering video stream into scene in three-dimensional geographic information space
CN105118061A
Positioning device for radiotherapy body surface optical tracking
CN111558174A
Image registration method and device, equipment and storage medium
CN116958211A