Image recognition method and system for tumor three-dimensional positioning
Through technical means such as dual-flow quantum convolutional coding and spin deformation field generation network, the accuracy and robustness of multimodal image registration and tumor positioning are solved, and high-precision three-dimensional tumor positioning and dynamic navigation are achieved.
Patent Information
- Application Number
- CN202510416377.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve high-precision and robust multimodal image registration and tumor localization, especially when facing nonlinear deformation, structural differences between modes and motion effects within the tumor area, traditional methods are difficult to meet the real-time navigation requirements.
The two-flow quantum convolutional coding method is used to extract CT and MRI image features, weighted fusion is performed after manifold space alignment, and deformation field is generated by spin deformation field generation network SpinNet and fractional-order ordinary differential equations, and tumor region segmentation is performed by combining convolution kernels with non-exchange curvature tensor modulation, and tumor center coordinates and motion trajectories are calculated using Mean Shift clustering and quaternary spline interpolation.
The multimodal feature alignment accuracy is improved, the accuracy and robustness of tumor area segmentation are enhanced, and the precise modeling of tumor space-time motion is realized, providing a solid foundation for dynamic navigation and motion compensation.
Smart Images

Figure CN120374842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically to an image recognition method and system for three-dimensional tumor localization. Background Art
[0002] In recent years, with the rapid development of medical imaging technology, multi-modal images such as CT and MRI have been widely used in tumor diagnosis and treatment planning. However, due to the large differences in physical principles, signal-to-noise ratios, resolutions, and reflections of tissue structures among different imaging modalities, how to achieve high-precision and robust multi-modal image registration and tumor localization has become a major problem in clinical diagnosis and surgical navigation.
[0003] Traditional image recognition methods usually rely on matching algorithms based on low-dimensional features such as grayscale, texture, or shape. Although relatively stable registration and segmentation can be achieved in some scenarios, it is often difficult to ensure sufficient accuracy and stability when facing the non-linear deformations, inter-modal structural differences, and motion effects within the tumor region. In addition, when existing methods process dynamic image sequences to capture the motion information of tumors over time, they often lack effective non-local feature extraction and spatio-temporal dynamic modeling mechanisms, resulting in the final generated three-dimensional recognition images being difficult to meet the requirements of real-time navigation. Summary of the Invention
[0004] Based on the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide an image recognition method and system for three-dimensional tumor localization to solve the above technical problems.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: An image recognition method for three-dimensional tumor localization, including:
[0006] S1: Obtain the CT image and MRI image to be localized, extract modal features using the dual-flow quantum convolution encoding method to obtain CT modal features and MRI modal features, and align them in the manifold space;
[0007] S2: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spinor representation at each position through the spinor deformation field generation network SpinNet, and generate a deformation field using the fractional ordinary differential equation;
[0008] S3: Map the MRI image to the CT image space through the deformation field to generate a registered and fused image, and construct a fractional U-Net through a convolution kernel modulated by a non-commutative curvature tensor to segment the tumor region and generate a tumor probability map;
[0009] S4: Calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image.
[0010] The present invention is further configured that, in step S1, for the CT image and the MRI image to be located, two parallel quantum convolution branches are used to extract the CT modality features and the MRI modality features respectively;
[0011] When aligning the CT modality features and the MRI modality features in the manifold space, the alignment error is optimized according to a preset manifold alignment loss function.
[0012] The present invention is further configured that, the CT modality features MRI modality features I C and I M are the CT image and the MRI image respectively, QConv(I; {U i}) is a convolution operation based on a parameterized quantum circuit, I is the input image, including I C and I M , {U i} is a set of parameters, including and and U i is the quantum convolution kernel, U i = R z (α i )R x (β i )R z (γ i ), R z and R x are the rotation matrices around the z-axis and the x-axis, α i , β i and γ i are the rotation angles;
[0013] Manifold alignment loss function Ω is the image domain, is the CT modality feature at position p, is the MRI modality feature at position φ(p), φ(p) is the correspondence of the manifold space to be optimized between the CT image and the MRI image, <·,·> is the inner product operation, and ||·||2 is the Euclidean norm.
[0014] The present invention is further configured that, in step S2, the aligned modality features are weighted and fused according to the quantum entanglement attention fusion method to obtain a fusion feature tensor, and the calculation logic is: F CM (p) = G QEG (p)⊙FC (p)+(1 - G QEG (p))⊙F M (p), F CM (p) is the fused feature tensor at position p, F C (p) and F M (p) are the CT - modality feature and MRI - modality feature at position p, ⊙ is the element - wise product, G QEG (p) is based on the joint density matrix ρ CM of the quantum entanglement attention gate, Tr(·) is the trace operation of a matrix, γ is the control parameter, and τ is the threshold.
[0015] The present invention is further set such that the spinor representation ψ(p) = SpinNet(F CM (p); W spin ) ∈ Spin(3), ψ(p) is the spinor representation at position p, SpinNet(·) is the spinor deformation field generation network SpinNet, based on the Spin(3) group representation, W spin is the network parameter;
[0016] The deformation field φ = exp(D a v), satisfying φ is the deformation field, D a is the Caputo fractional - order derivative operator, v is the velocity field, β is the control parameter, ψ(φ t ) is the spinor information at the current deformation state φ t , and v(φ t ) is the velocity - field information at the current deformation state φ t .
[0017] The present invention is further set such that in step S3, the MRI image is mapped to the CT - image space by the deformation field to generate a registered and fused image;
[0018] The registered and fused image is convolved by a convolution kernel modulated by a non - commutative curvature tensor;
[0019] A fractional - order U - Net is constructed to process the convolution result to obtain an entangled - flow feature map;
[0020] The entangled - flow feature map is converted to a single - channel feature map through a mapping layer, and the single - channel feature map is normalized using the Sigmoid activation function to generate a tumor probability map.
[0021] The present invention is further set such that the convolution kernel is the non - commutative curvature modulation factor at the spatial position (i, j, k), κ Ais the curvature at the spatial position (i, j, k), and W base is the learnable weight of the basic convolution kernel, and b is the bias term;
[0022] The entangled flow feature map F out = σ(D b (K * F in ))), D b is the fractional differential operator, K is the convolution kernel, and F in is the registered and fused image, and σ(·) is the activation function;
[0023] The mapping logic is: Z = W p * F out + b p , where Z is the single-channel feature map, W p is the mapping weight, and b p is the mapping bias term. The Sigmoid activation function is used to normalize the single-channel feature map Z to obtain the tumor probability map P tumor = Sigmoid(Z).
[0024] The present invention is further configured such that in step S4, for each position in the image domain of the tumor probability map, in the hyperbolic space, the local centroid within the local neighborhood is calculated;
[0025] After calculating the local centroids of all positions over the entire image domain, the Mean Shift clustering algorithm is used to cluster all the centroids;
[0026] After the clustering algorithm converges, the clustering center with the highest density is selected as the tumor center;
[0027] After determining the tumor center, the motion trajectory is fitted by the quaternion spline interpolation method to determine the motion of the tumor center over time in the dynamic image sequence.
[0028] The present invention is further configured such that the calculation logic of the local centroid is: m(p) is the local centroid at position p, N p is the local neighborhood, P tumor (q) is the tumor probability at position q, K Π (d Π (p, q)) is the kernel function, and weights are assigned to each position within the neighborhood according to the hyperbolic distance d Π (p, q), ∥·∥ 2 is the Euclidean distance;
[0029] The motion trajectory n is the number of tumor centers in the dynamic image sequence, B k (t) is the B-spline basis function, ψk is the spinor information at the k-th tumor center, p k is the position vector of the k-th tumor center in space.
[0030] The present invention also provides an image recognition system for three-dimensional tumor localization, which is used to implement the above-mentioned image recognition method for three-dimensional tumor localization. The system includes:
[0031] Feature alignment module: Obtain the CT image and MRI image to be localized, extract modal features using the dual-flow quantum convolution coding method to obtain CT modal features and MRI modal features, and align them in the manifold space;
[0032] First generation module: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spinor representation at each position through the spinor deformation field generation network SpinNet, and generate a deformation field using a fractional ordinary differential equation;
[0033] Second generation module: Map the MRI image to the CT image space through the deformation field to generate a registered fusion image, and construct a fractional U-Net through a convolution kernel modulated by a non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map;
[0034] Third generation module: Calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image.
[0035] The present invention provides an image recognition method and system for three-dimensional tumor localization. The method includes obtaining the CT image and MRI image to be localized, extracting modal features using the dual-flow quantum convolution coding method to obtain CT modal features and MRI modal features, and aligning them in the manifold space; performing weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spinor representation at each position through the spinor deformation field generation network SpinNet, and generate a deformation field using a fractional ordinary differential equation; map the MRI image to the CT image space through the deformation field to generate a registered fusion image, and construct a fractional U-Net through a convolution kernel modulated by a non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map; calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image. The beneficial effects generated include:
[0036] 1. Improve the accuracy of multi-modal feature alignment: Use a dual-flow quantum convolutional encoding network to extract deep features of CT and MRI modalities respectively, and introduce a manifold alignment loss function. By minimizing the geodesic distance between modalities, precise alignment of multi-modal features in a non-Euclidean space is achieved, effectively alleviating the feature mismatch problem caused by modality differences in traditional methods;
[0037] 2. Improve the accuracy and robustness of tumor region segmentation: Construct a fractional-order U-Net network by non-commutative geometry modulating convolutional kernels, and combine fractional-order calculus convolution to extract high-dimensional features with stronger non-locality and structure perception ability, improving the accuracy of tumor boundary extraction and the adaptability to complex backgrounds;
[0038] 3. Achieve precise modeling of tumor spatio-temporal motion: Calculate the spatial density center of the tumor region based on the Mean Shift clustering algorithm in hyperbolic space, and use quaternion spline interpolation to fit the motion trajectory of the tumor in the dynamic image sequence, which can accurately depict the spatial position and rotation behavior of the tumor over time, providing a basis for dynamic navigation and motion compensation.
[0039] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically presents the specific embodiments of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0041] Figure 1 It is a flowchart of an image recognition method for tumor three-dimensional positioning shown in an exemplary embodiment of the present invention;
[0042] Figure 2 It is a schematic structural diagram of an image recognition system for tumor three-dimensional positioning shown in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the protection scope of the present invention.
[0044] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0045] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0046] Embodiment 1
[0047] An image recognition method for three-dimensional tumor localization, as Figure 1 shown, includes:
[0048] S1: Obtain the CT image and MRI image to be localized, extract modal features using the dual-flow quantum convolution encoding method to obtain CT modal features and MRI modal features, and align them in the manifold space;
[0049] S2: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spin representation at each position through the spin deformation field generation network SpinNet, and generate a deformation field using the fractional ordinary differential equation;
[0050] S3: Map the MRI image to the CT image space through the deformation field to generate a registered and fused image, and construct a fractional U-Net through a convolution kernel modulated by a non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map;
[0051] S4: Calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, and fit the motion trajectory through the quaternion spline interpolation method to generate a three-dimensional recognition image.
[0052] The present invention is further configured such that, in step S1, for the CT image and the MRI image to be located, two parallel quantum convolution branches are used to extract CT modality features and MRI modality features respectively;
[0053] When aligning the CT modality features and the MRI modality features in the manifold space, the alignment error is optimized according to a preset manifold alignment loss function.
[0054] The present invention is further configured such that the CT modality features MRI modality features I C and I M are the CT image and the MRI image respectively, QConv(I; {U i}) is a convolution operation based on a parameterized quantum circuit, I is the input image, including I C and I M , {U i} is a set of parameters, including and and U i is the quantum convolution kernel, U i = R z (α i )R x (β i )R z (γ i ), R z and R x are rotation matrices around the z-axis and the x-axis, α i , β i and γ i are rotation angles; specifically, the above calculation logic uses a convolution operation based on a parameterized quantum circuit (abbreviated as "QConv") to extract the deep features of the CT and MRI images respectively. Specifically, for the CT image I C and the MRI image I M to be processed, through the operation the CT modality feature F C is obtained, and through the operation the MRI modality feature F M is obtained, where QConv(I; {U i}) represents an operation of performing convolution using a parameterized quantum circuit, aiming to utilize the high-dimensional mapping and interference effects of quantum gates to extract deep feature information that is difficult to capture by traditional convolution; the CT and MRI modal features extracted by this method can fully utilize the high-dimensional mapping and interference effects of quantum gates in the quantum convolution operation, achieve accurate multi-modal feature extraction and alignment, use the parameterized quantum circuit to obtain deeper and more discriminative features, and provide a solid foundation for subsequent image registration, segmentation, and motion trajectory modeling;
[0055] Manifold alignment loss function Ω is the image domain, is the CT modal feature at position p, is the MRI modal feature at position φ(p), where φ(p) is the correspondence of the manifold space to be optimized between the CT image and the MRI image, <·,·> is the inner product operation, and ||·||2 is the Euclidean norm; specifically, the manifold alignment loss function aims to align the features from the CT and MRI images in the manifold space, thereby reducing the representation difference between different modalities. This loss measures the consistency by calculating the angle between the feature vectors of the two modalities at the corresponding positions (or the corresponding positions after mapping). Specifically, the cosine similarity of two vectors is calculated using the inner product and the Euclidean norm, and then the angle is obtained through the inverse cosine (arccos). Finally, the angles at each position in the entire image domain are accumulated as the alignment error; this loss function effectively alleviates the feature mismatch problem between CT and MRI due to the difference in imaging principles, thereby improving the accuracy of multi-modal feature fusion; when the different modal images have similar anatomical structures in the local area, the angle between the corresponding feature vectors will be smaller, and the loss function tends to zero; while in the inconsistent areas, the larger angle will prompt the network to adjust the parameters, thereby enhancing the robustness of the overall system to anomalies or noises.
[0056] The present invention is further configured such that in step S2, the aligned modal features are weighted and fused according to the quantum entanglement attention fusion method to obtain a fused feature tensor, and the calculation logic is: F CM (p) = G QEG (p) ⊙ F C (p) + (1 - G QEG (p)) ⊙ F M (p), F CM (p) is the fused feature tensor at position p, F C (p) and F M (p) are the CT modal feature and the MRI modal feature at position p, ⊙ is the element-wise product, and G QEG (p) is the quantum entanglement attention gate based on the joint density matrix ρ CM ; Tr(·) is the trace operation of a matrix, γ is a control parameter, and τ is a threshold. Specifically, the above calculation logic performs weighted fusion on the aligned CT and MRI modal features, aiming to adaptively determine the contribution ratio of different modal features by using the Quantum Entangled Attention Gate (abbreviated as G QEG (p)) to generate a fused feature tensor; the Quantum Entangled Attention Gate G QEG (p) is a quantum state description based on the joint density matrix ρ CM . After calculating the quantization information uncertainty of the von Neumann entropy, it outputs a weight in the form of a Sigmoid function. This weight adaptively reflects the consistency or complementarity of CT and MRI features at the current position, thereby determining the weight distribution during fusion; the joint density matrix ρ CM is a density matrix constructed by fusing CT and MRI modal features to describe the overall state of the two modal features at this position. This matrix satisfies the physical properties of positive definiteness and trace equal to 1; the control parameter γ is used to adjust the slope of the Sigmoid function. Its positive value determines the steepness of the weight change from close to 1 to close to 0, and its value range is [1, 10]. The threshold τ is used to determine at what entropy value level the attention weight begins to decrease significantly. When the entropy value of the joint density matrix is lower than τ, it indicates that the two modal features are relatively consistent, and at this time, the value of G QEG (p) approaches 1; when the entropy value is higher than τ, it indicates that the features are chaotic or inconsistent, and G QEG (p) approaches 0. Its specific value depends on the prior analysis of the data statistical distribution and is optimized and determined during the network training process; by dynamically adjusting the weights of CT and MRI features according to the entropy value of the joint density matrix through the Quantum Entangled Attention Gate, the fine fusion of different modal information is realized, thereby alleviating the feature mismatch problem caused by modal differences.
[0057] The present invention is further set such that the spinor representation ψ(p) at each position is ψ(p) = SpinNet(F CM (p); W spin ) ∈ Spin(3), ψ(p) is the spinor representation at position p, SpinNet(·) is the spinor deformation field generation network SpinNet, based on the Spin(3) group representation, and W spin is the network parameter; specifically, in the previous steps, the fused feature tensor F CM(p), the feature tensor contains the deep information of both CT and MRI modalities. To capture the complex rotation and deformation information in the local space, the present invention adopts the spinor representation based on the Spin(3) group. Using the spinor deformation field generation network SpinNet, the fused feature is mapped to the spinor representation ψ(p), which belongs to the Spin(3) group (equivalent to the quaternion form) and can effectively express the local rotation information, thus providing accurate rotation parameters for subsequent deformation field generation and motion compensation. Spin(3) is the double covering group of the three-dimensional rotation group SO(3) and is commonly represented by quaternions. Compared with SO(3), the representation of Spin(3) is smoother and suitable for end-to-end learning in neural networks. The spinor representation ψ(p) represents the rotation information at position p and is usually represented by the quaternion w + xi + yj + zk. Its connotation includes not only the rotation angle but also the rotation axis information. SpinNet is a neural network based on the Spin(3) group structure for generating spinor representations. The network performs a non-linear mapping on the fused feature tensor and outputs a spinor conforming to the Spin(3) structure, and the network parameters W spin The parameter set contains all the trainable weights in SpinNet and is automatically updated through the training process to ensure that the network can accurately capture and express local rotation and deformation information. By mapping the fused feature to the Spin(3) group, the rotation and deformation characteristics of the local region in the image can be captured, providing a fine rotation description for subsequent deformation field generation; using the spinor representation is more continuous and smoother than the traditional vector form, which helps the model to remain stable when dealing with complex non-linear deformations and improves the accuracy of image registration and motion tracking;
[0058] The deformation field φ = exp(D a v), satisfying φ is the deformation field, D a is the Caputo fractional derivative operator, v is the velocity field, β is the control parameter, ψ(φ t ) is the spinor information at the current deformation state φ t , and v(φ t ) is the velocity field information at the current deformation state φ t . Specifically, in the multi-modal image registration process, in order to accurately align different modal images, a continuous, reversible and memory-effect deformation field φ needs to be constructed. This method uses a fractional ordinary differential equation (ODE) to generate the deformation field, in which the Caputo fractional derivative operator D a and the control parameter β are introduced to capture long-term dependencies (memory effects) and non-local information. The deformation field φ refers to the continuous transformation that maps each point in the source image to the target image space and is the key variable for image registration; the velocity field v represents the motion rate and direction of each point during the deformation process, and the Caputo fractional derivative operator D aIt is a fractional calculus operator, characterized by its ability to capture non-local effects and historical dependencies. The parameter a (order) determines the "memory length" and non-locality of the fractional derivative, with a value range of (0, 1]; the control parameter β is used for order control in the fractional derivative, reflecting the influence degree of historical states on the generation of the current deformation field along the integration path, with a value range of (0, 1]. Through fractional ordinary differential equations and exponential mapping, the generated deformation field is not only continuous and smooth but also ensures homeomorphism (reversibility), which helps to achieve accurate image registration.
[0059] The present invention is further configured such that, in step S3, the MRI image is mapped to the CT image space through the deformation field to generate a registered and fused image; specifically, the MRI image is "distorted" using the deformation field obtained by the above calculation logic, that is, using the mapping operator to map the MRI image I M to the CT image space, thereby obtaining the registered and fused image: Here, represents the function composition operation, that is, for the position q of each pixel or voxel in the MRI image, the corresponding position p = φ(q) in the CT image space is obtained using the deformation field φ, and the gray level or multi-channel information of the MRI image is repositioned to this position;
[0060] The registered and fused image is subjected to a convolution operation using a convolution kernel modulated based on the non-commutative curvature tensor; the present invention is further configured such that the convolution kernel is a non-commutative curvature modulation factor at the spatial position (i, j, k), κ A is the curvature at the spatial position (i, j, k), W base is the learnable weight of the basic convolution kernel, and b is the bias term; specifically, in order to more accurately extract the local geometric features in the registered and fused image, the present invention introduces a convolution kernel modulated based on the non-commutative curvature tensor. Its core idea is: based on the standard convolution kernel, it is dynamically modulated according to the local curvature information of each spatial position in the image, so that the convolution operation can adaptively capture the local complex geometric structure; the non-commutative curvature modulation factor depends on the local curvature κ A at the spatial position (i, j, k), which is obtained from the structure of the image or through geometric analysis, including second-order derivative or curvature tensor calculation, and is not limited here, reflecting the degree of bending or shape characteristics of the local area; by integrating the local curvature information into the convolution kernel, it can be dynamically modulated according to the geometric characteristics of different regions, which helps to extract more discriminative local features, especially for the edge and complex-shaped tumor regions, and can better capture the local structural details;
[0061] Construct a fractional U-Net to process the convolution result and obtain an entangled flow feature map; the entangled flow feature map F out =σ(D b (K * F in ))), where D b is a fractional differential operator, K is a convolution kernel, and F in is the registered and fused image, and σ(·) is an activation function; specifically, the fractional U-Net refers to introducing a fractional differential operator into the U-Net network, aiming to enhance the network's ability to capture non-local and long-range dependence characteristics. The fractional differential operator D b is an operator that generalizes the traditional differential. Its order b determines the strength of non-locality and memory effect, can capture richer local and global feature information, and its value range is (0, 2]. The fractional differential operator can capture the long-range dependence and historical information in the image, making the features after convolution retain both local details and global semantics;
[0062] Convert the entangled flow feature map into a single-channel feature map through a mapping layer, and normalize the single-channel feature map using the Sigmoid activation function to generate a tumor probability map; the mapping logic is: Z = W p *F out +b p , where Z is the single-channel feature map, W p is the mapping weight, and b p is the mapping bias term. Normalize the single-channel feature map Z using the Sigmoid activation function to obtain the tumor probability map P tumor = Sigmoid(Z).
[0063] The present invention is further configured such that in step S4, for each position in the image domain of the tumor probability map, in the hyperbolic space, calculate the local centroid within the local neighborhood; the present invention is further configured such that the calculation logic of the local centroid is: m(p) is the local centroid at position p, N p is the local neighborhood, P tumor (q) is the tumor probability at position q, and K Π (d Π (p, q)) is the kernel function, and assign weights to each position within the neighborhood according to the hyperbolic distance d Π (p, q), ∥·∥ 2 is the Euclidean distance; specifically, in the image domain of the tumor probability map, in order to determine the spatial center of the tumor, the present invention uses the local centroid calculation method in the hyperbolic space. Specifically, for each position p in the image domain, within its local neighborhood N p , according to the tumor probability P of each position q tumor(q) and the hyperbolic distance K between q and p Π (d Π (p, q)) to assign corresponding weights; adopt the kernel function K Π (d Π (p, q)) to perform weighted averaging on each point in the neighborhood, and calculate the local centroid m(p) of the position p; perform subsequent processing on all local centroids in the entire image to determine the final tumor center coordinates; by using the weighted centroid calculation in hyperbolic space, it can better adapt to the image distribution characteristics under non-Euclidean geometry, making the calculated local centroid more accurately reflect the actual distribution of the tumor region;
[0064] On the entire image domain, after calculating the local centroids of all positions, use the Mean Shift clustering algorithm to cluster all centroids; specifically, the Mean Shift clustering algorithm is a prior art and will not be elaborated here;
[0065] After the clustering algorithm converges, select the clustering center with the highest density as the tumor center;
[0066] After determining the tumor center, fit the motion trajectory through the quaternion spline interpolation method to determine the motion of the tumor center over time in the dynamic image sequence; the motion trajectory n is the number of tumor centers in the dynamic image sequence, B k (t) is the B-spline basis function, ψ k is the spinor information at the k-th tumor center, p k is the position vector of the k-th tumor center in space; specifically, through the previous steps, multiple tumor centers at multiple time points are obtained in the dynamic image sequence, and each center is represented by its spatial position vector p k , and at the same time, the spinor information ψ k at each center is obtained, which is used to describe the local rotation and pose. In order to describe the continuous motion of the tumor center over time, the quaternion spline interpolation method is adopted. Specifically, use the B-spline basis function B k (t) to perform smooth interpolation on the tumor centers at each time point, and combine the spinor information ψ k to modulate the rotation change in the motion trajectory, and finally obtain a continuous trajectory function. This trajectory not only includes position changes but also reflects the rotation information during the motion process. Through B-spline interpolation, the contributions of each control point to the motion trajectory are smoothly weighted, making the trajectory continuous and smooth in time, effectively reducing the influence of noise and local errors; fuse the spinor information ψ kIt can not only describe the displacement changes of the tumor center, but also reflect its rotation or pose changes during movement, providing a more complete geometric description for dynamic navigation and motion compensation. Through the quaternion spline interpolation method, the discrete tumor center and its rotation information are smoothly fused in time into a continuous motion trajectory, effectively reflecting the spatial position and rotation changes of tumor motion, and significantly improving the accuracy and robustness of dynamic tumor localization and subsequent navigation.
[0067] Embodiment 2
[0068] Please refer to Figure 2 , the exemplary image recognition system for three-dimensional tumor localization, which is used to implement the above-mentioned image recognition method for three-dimensional tumor localization. The system includes:
[0069] Feature alignment module: Obtain the CT image and MRI image to be localized, extract modal features using the dual-flow quantum convolution encoding method to obtain CT modal features and MRI modal features, and align them in the manifold space;
[0070] First generation module: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spinor representation at each position through the spinor deformation field generation network SpinNet, and generate a deformation field using the fractional ordinary differential equation;
[0071] Second generation module: Map the MRI image to the CT image space through the deformation field to generate a registered fusion image, and construct a fractional U-Net through a convolution kernel modulated by the non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map;
[0072] Third generation module: Calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image.
[0073] It should be noted that the above-mentioned image recognition system for three-dimensional tumor localization provided by the above embodiment and the above-mentioned image recognition method for three-dimensional tumor localization belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the above-mentioned image recognition system for three-dimensional tumor localization provided by the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the above-described functions. This is not limited here either.
[0074] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0075] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0076] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0077] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0078] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0079] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0080] In several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0081] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0082] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0083] When the above-described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0084] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An image recognition method for three-dimensional tumor localization, characterized in that Including: S1: Obtain the CT image and MRI image to be located, extract modal features using the dual-flow quantum convolution encoding method to obtain CT modal features and MRI modal features, and align them in the manifold space; S2: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spin representation at each position through the spin deformation field generation network SpinNet, and generate a deformation field using the fractional ordinary differential equation; S3: Map the MRI image to the CT image space through the deformation field to generate a registered fusion image, and construct a fractional U-Net through a convolution kernel modulated by the non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map; S4: Calculate the tumor center coordinates according to the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image.
2. The image recognition method for three-dimensional tumor localization according to claim 1, wherein In step S1, for the CT image and MRI image to be located, two parallel quantum convolution branches are used to extract CT modal features and MRI modal features respectively; When aligning the CT modal features and MRI modal features in the manifold space, optimize the alignment error according to the preset manifold alignment loss function.
3. The image recognition method for three-dimensional tumor localization according to claim 2, characterized in that, CT modal features MRI modal features I C and I M are the CT image and the MRI image respectively. QConv(I; {U i}) is a convolutional operation based on a parameterized quantum circuit. I is the input image, including I C and I M . {U i} is a set of parameters, including and and U i is the quantum convolution kernel. U i = R z (α i )R x (β i )R z (γ i ). R z and R x are the rotation matrices around the z-axis and the x-axis respectively. α i , β i and γ i are the rotation angles; Manifold alignment loss function Ω is the image domain, is the CT modality feature at position p, is the MRI modality feature at position φ(p), where φ(p) is the manifold space correspondence to be optimized between the CT image and the MRI image, <·,·> is the inner product operation, and ||·||2 is the Euclidean norm.
4. The image recognition method for three-dimensional tumor localization according to claim 1, wherein In step S2, the aligned modal features are weighted and fused according to the quantum entanglement attention fusion method to obtain a fused feature tensor. The calculation logic is: F CM (p) = G QEG (p) ⊙ F C (p) + (1 - G QEG (p)) ⊙ F M (p), where F CM (p) is the fused feature tensor at position p, F C (p) and F M (p) are the CT modal feature and MRI modal feature at position p, ⊙ is the element-wise product, and G QEG (p) is based on the quantum entanglement attention gate of the joint density matrix ρ CM . Tr(·) is the trace operation of the matrix, γ is the control parameter, and τ is the threshold.
5. The image recognition method for three-dimensional tumor localization according to claim 4, wherein, The spinor representation ψ(p) = SpinNet(F CM (p); W spin ) ∈ Spin(3), where ψ(p) is the spinor representation at position p, SpinNet(·) is the spinor deformation field generation network SpinNet, based on the Spin(3) group representation, and W spin are the network parameters; The deformation field φ = exp(D a v), satisfies where φ is the deformation field, D a is the Caputo fractional derivative operator, v is the velocity field, β is the control parameter, ψ(φ t ) is the spinor information at the current deformation state φ t , and v(φ t ) is the velocity field information at the current deformation state φ t .
6. The image recognition method for three-dimensional tumor localization according to claim 1, characterized in that In step S3, map the MRI image to the CT image space through the deformation field to generate a registered fusion image; Perform convolution operation on the registered fusion image through a convolution kernel modulated by the non-commutative curvature tensor; Construct a fractional U-Net to process the convolution result to obtain an entangled flow feature map; Convert the entangled flow feature map to a single-channel feature map through a mapping layer, and normalize the single-channel feature map using the Sigmoid activation function to generate a tumor probability map.
7. The image recognition method for three-dimensional tumor localization according to claim 6, wherein Convolution kernel is a non-commutative curvature modulation factor κ at the spatial position (i, j, k). A is the curvature at the spatial position (i, j, k), and W base is the learnable weight of the basic convolution kernel, and b is the bias term; Entangled flow feature map F out = σ(D b (K * F in ))), D b is a fractional differential operator, K is a convolution kernel, F in is a registered and fused image, and σ(·) is an activation function; The mapping logic is: Z = W p *F out +b p , where Z is a single-channel feature map, W p is the mapping weight, and b p is the mapping bias term. The single-channel feature map Z is normalized using the Sigmoid activation function to obtain the tumor probability map P tumor = Sigmoid(Z).
8. The image recognition method for three-dimensional tumor localization according to claim 1, wherein In step S4, for each position in the image domain of the tumor probability map, calculate the local centroid in the local neighborhood in the hyperbolic space; On the entire image domain, after calculating the local centroids of all positions, use the Mean Shift clustering algorithm to cluster all centroids; After the clustering algorithm converges, select the clustering center with the highest density as the tumor center; When the tumor center is determined, fit the motion trajectory through the quaternion spline interpolation method to determine the motion of the tumor center over time in the dynamic image sequence.
9. The image recognition method for three-dimensional tumor localization according to claim 8, characterized in that The calculation logic of the local center of gravity is as follows: m(p) is the local center of gravity at position p, N p is the local neighborhood, P tumor (q) is the tumor probability at position q, K Π (d Π (p,q)) is the kernel function, according to the hyperbolic distance d Π (p,q) assigns weights to each position within the neighborhood, ||·|| 2 is the Euclidean distance; Motion trajectory n is the number of tumor centers in the dynamic image sequence, B k (t) is the B-spline basis function, ψ k is the spinor information at the k-th tumor center, p k is the position vector of the k-th tumor center in space.
10. An image recognition system for three-dimensional tumor localization, which is used to implement the image recognition method for three-dimensional tumor localization according to any one of claims 1-9, characterized in that, Including: Feature alignment module: Obtain the CT image and MRI image to be located, extract modal features using the dual-flow quantum convolution encoding method to obtain CT modal features and MRI modal features, and align them in the manifold space; First generation module: Perform weighted fusion on the aligned modal features to obtain a fused feature tensor. Based on the fused feature tensor, calculate the spin representation at each position through the spin deformation field generation network SpinNet, and generate a deformation field using the fractional ordinary differential equation; Second generation module: Map the MRI image to the CT image space through the deformation field to generate a registered fusion image, and construct a fractional U-Net through a convolution kernel modulated by the non-commutative curvature tensor to perform tumor region segmentation and generate a tumor probability map; Third generation module: Calculate the tumor center coordinates based on the tumor probability map and the dynamic image sequence through the Mean Shift clustering method, fit the motion trajectory through the quaternion spline interpolation method, and generate a three-dimensional recognition image.
Citation Information
Cited By
Intelligent planning and intraoperative navigation system for spinal surgery
CN121421674A
Intelligent planning and intraoperative navigation system for spinal surgery
CN121421674B