Automatic positioning method and system for eardrum hydrops surgery catheterization point and storage medium

By constructing a tympanic membrane segmentation and object detection model based on the potential diffusion model, combined with the geometric model, the problem of insufficient accuracy and explanation of the positioning of the pipe positioning point in the existing technology of hydrops in the eardrum is solved, and efficient and reliable pipe positioning is achieved, reducing the cost of data labeling and improving the applicability of the model.

CN120284466APending Publication Date: 2025-07-11HEILONGJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510354663.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning-based surgical tubing point positioning method of hydrostatic eardrum surgery has problems such as low accuracy, poor interpretation and insufficient generalization ability, and cannot meet clinical needs.

Method used

The tympanic membrane segmentation model and object detection model based on the potential diffusion model are adopted. By constructing a noise reduction generation unit, a perceptual compression unit and a multimodal condition coding unit, combining texture perception, semantic enhancement and residual multi-layer perceptron, high-precision segmentation and target recognition of ear endoscopic images are achieved, and the pipe positioning point is determined in combination with geometric models.

Benefits of technology

It improves the accuracy and interpretability of the positioning of the pipe positioning point, reduces the cost of data labeling, enhances the generalization ability of the model, reduces the doctor's surgical preparation time, and improves the operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120284466A_ABST
    Figure CN120284466A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic positioning method and system for a catheterization point in an eardrum hydrops operation and a storage medium, relates to the technical field of medical image processing, and aims to solve the problem that the existing catheterization point positioning method based on deep learning is generally a black box model, cannot give a basis for recommending the catheterization point by the model, and cannot enable the obtained catheterization point result to be reliable for medical care. The method comprises the following steps: firstly, constructing a tympanic membrane contour data set and an auxiliary point data set; constructing a tympanic membrane segmentation model comprising a noise reduction generation unit, a sensing compression unit and a multi-modal condition coding unit; through dynamic fusion of texture features, a semantic graph and hidden space representation, high-precision segmentation of the tympanic membrane contour of the ear endoscope image is realized. Identifying an in-ear hammer bone handle and an umbilical region of the ear endoscope image through the target detection model; and determining an eardrum hydrops operation catheterization point according to a mirror image relationship between the left ear and the right ear by combining the tympanic membrane profile curve and the identified in-ear hammer bone handle and the umbilical region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and more particularly, to an automatic positioning method, system and storage medium for the catheterization point of tympanic membrane hydrops surgery. Background Art

[0002] Otitis media with effusion (OME) is a frequent non-suppurative middle ear disease in children and adolescents. If not treated properly, it is extremely likely to lead to poor language development or irreversible deafness. At present, there are two relatively effective treatment methods for this disease under the condition of otoscope: tympanostomy and tympanocentesis. The former has been proven to significantly reduce the recurrence rate and has a better postoperative recovery effect, and is used as the preferred method in clinical treatment.

[0003] Tympanostomy is to connect the inside and outside of the tympanic membrane through a tubular object, so that the liquid in the tympanic membrane flows out and the ventilation and drainage are improved (as shown in Figure 6 left). As shown in Figure 6 (right), clinically, the tympanic membrane is divided into anterior and posterior quadrants along the malleus handle (mirror division of left and right ears). The placement position of the tubular object (catheterization point) is selected in the anteroinferior quadrant, at the midpoint of the line segment AB, marked as a blue circle. The selection of the catheterization point is crucial for the success of the operation and the postoperative cure effect: if it is close to the edge of the tympanic membrane, the puncture is likely to damage the mucosa of the inner wall of the tympanic membrane and the temporal bone; if it deviates from the anteroinferior quadrant, the cure effect is not good, and even postoperative infection may occur. In addition, the diverse morphological structures of the affected ear and the superposition of various ear diseases (such as fungal infections in the ear canal, cerumen, etc.) greatly increase the difficulty of determining the catheterization point, and only a few experienced physicians can perform tympanostomy in various situations.

[0004] Currently, in order to solve the problem of difficult catheterization point positioning, some studies have tried to use the emerging artificial intelligence technology in recent years, especially deep learning, to develop a model for automatically positioning the catheterization point. These studies have alleviated the operation burden of physicians to a certain extent, but there are still some obvious disadvantages, including low accuracy and worse performance in complex ear environments than manual operation.

[0005] Catheterization point: refers to the position where the doctor selects to place the drainage tube during tympanotomy catheterization, usually determined according to the anatomical structure of the tympanic membrane and the distribution of effusion to achieve the best drainage effect.

[0006] Tympanic membrane (ear membrane): a thin film tissue located between the outer ear and the middle ear, mainly functioning to transmit sound wave vibrations to the middle ear and serving as a protective barrier for the middle ear.

[0007] Umbilicus: the slightly concave part inside the center of the ear membrane, connected to the end of the malleus handle, which is an important anatomical landmark of the ear membrane.

[0008] Existing methods usually regard the positioning of the catheterization point as an object recognition task: during the training phase, the model relies on a large amount of sample data with catheterization point annotations; during actual use, the model directly outputs the coordinates of the catheterization point based on the input otoscope image. However, this method has the following main problems:

[0009] 1. High data requirements and high construction costs. Training such end-to-end positioning models requires a large amount of high-quality annotated data sets. Building such a data set not only requires collecting a large number of otoscope images, but also consumes a large amount of manpower for professional annotation. Especially in the case of containing multiple pathological features, the construction process is extremely time-consuming and costly.

[0010] 2. Lack of interpretability. Deep learning models are usually "black box" models, and their decision-making processes are difficult to intuitively understand or explain. In clinical applications, doctors and patients need to clearly understand the basis for the model to recommend the catheterization point. However, since the existing models cannot provide clear decision-making logic, it is difficult for both doctors and patients to fully trust the results given by the model under the premise that the accuracy has not reached the level of senior physicians.

[0011] 3. Insufficient generalization ability. The performance of deep learning supervised models highly depends on the quality and distribution of the training data. If the otoscope images input during actual use are significantly different from the training samples (such as changes in lighting conditions, resolution, or lesion characteristics), the performance of the model may drop significantly, showing poor generalization ability. This limitation seriously affects its applicability in different clinical scenarios.

[0012] In summary, although existing artificial intelligence technologies have made certain progress in catheterization point positioning, their deficiencies in accuracy, interpretability, and generalization ability make them unable to fully meet the actual clinical needs. Therefore, there is an urgent need for a new solution that can achieve higher positioning accuracy in complex scenarios, while having good interpretability and generalization ability to meet the requirements of clinical use. Summary of the Invention

[0013] The technical problem to be solved by the present invention is:

[0014] Existing deep learning-based catheterization point positioning methods are usually "black box" models and cannot give the basis for the model to recommend the catheterization point, making the obtained catheterization point results untrustworthy for medical staff.

[0015] The technical solution adopted by the present invention to solve the above technical problems:

[0016] The present invention provides an automatic positioning method for the catheterization point in myringotomy surgery, including the following steps:

[0017] Step 1: Collect otoscope images, annotate the contour of the eardrum, construct a tympanic membrane contour dataset, annotate the positions of the malleus handle and umbo, and construct an auxiliary point dataset;

[0018] Step 2: Construct a tympanic membrane segmentation model based on a latent diffusion model, including a denoising generation unit, a perceptual compression unit, and a multimodal conditional encoding unit;

[0019] The denoising generation unit is used to implement reverse diffusion using a deep U-Net architecture, generate a high-precision segmentation mask from the latent space noise distribution by iterative denoising, and construct a generative discriminative framework to probabilistically model the segmentation results;

[0020] The perceptual compression unit constructs a bidirectional mapping channel based on a variational autoencoder. The encoder is used to reduce the dimensionality of the segmentation results in the high-dimensional pixel space to a compact latent space representation, and the decoder is used to reconstruct the latent space to achieve data fidelity;

[0021] The multimodal conditional encoding unit is used to extract image features, including a texture perception module, a semantic enhancement module, and a conditional segmentation module. The texture perception module is used to extract image edge features using Sobel and Laplacian differential operators and extract directional features using adaptive Gabor convolution; the semantic enhancement module is used to establish a hypergraph data structure for the image and mine the semantic relationships of the image; the multi-condition fusion module is used to adaptively fuse texture features, semantic features, and encoded features by combining spatial and channel attention mechanisms and a residual multi-layer perceptron;

[0022] Use the tympanic membrane contour dataset to train the tympanic membrane segmentation model;

[0023] Step 3: Construct an object detection model; use the auxiliary point dataset to train the object detection model;

[0024] Step 4: Use the trained tympanic membrane segmentation model to segment the tympanic membrane contour of the otoscope image, and use the trained object detection model to identify the malleus handle and umbo in the ear of the otoscope image;

[0025] Step 5: Fit the extracted tympanic membrane contour to obtain a tympanic membrane contour curve, combine the identified malleus handle and umbo in the ear, and determine the catheterization point according to the mirror image relationship of the left and right ears.

[0026] Furthermore, Step 1 includes performing desensitization, cropping, and scaling processing on the collected otoscope images.

[0027] Furthermore, using the tympanic membrane contour dataset to train the tympanic membrane segmentation model in Step 2 includes the following process:

[0028] The perception compression unit uses a variational autoencoder E and D to implement the perception compression function, expressed as where is the eardrum segmentation mask; the multi-modal conditional encoding unit constructs a conditional encoder τ with a structure similar to E θ to implement the conversion from the pixel space to the latent space, expressed as z c = τ θ (I), is the otoscope image, and z c is the conditional feature; the model training is divided into two stages. In the first stage, the variational autoencoders E and D are trained using the dataset composed of masks. In the second stage, the variational autoencoders are frozen, and the denoising generation unit and the conditional encoder τ are trained using the complete training set θ , making the eardrum segmentation model an end-to-end model;

[0029] During training, the denoising generation unit adds noise to the latent space representation z0 of the masked image, and then uses the denoising U-Net to predict the noise to restore and generate the eardrum mask, expressed as ε θ is the U-Net network;

[0030] The noise prediction loss during the denoising process is

[0031] The noise-adding forward process is:

[0032]

[0033] where ε represents Gaussian noise, is a hyperparameter that controls the noise level;

[0034] The reverse denoising process is:

[0035]

[0036] The loss function of the supervised segmentation is introduced to generate discriminative segmentation. The total loss function of the eardrum segmentation model is as follows:

[0037] L = L noise + λL seg

[0038] where λ represents the weight of the segmentation function.

[0039] Furthermore, in step 2, the texture perception module combines standard convolution and multi-scale convolution cascades. The boundary detection operator and the adaptive Gabor convolution extract boundary and orientation features respectively, and then the features are concatenated to form texture features;

[0040] The mathematical expression for the texture perception module to extract image edges using Sobel differential operators and Laplacian operators in the horizontal and vertical directions is as follows:

[0041]

[0042] Furthermore, the functional implementation process of the semantic enhancement module is as follows:

[0043] Use the conditional encoder τ θ Construct a hypergraph from the high-dimensional features of the intermediate layer Where and ε represent the vertex set and the hyperedge set respectively; a threshold-based ε-ball is established for each feature point, representing a hyperedge, which contains all feature points within the specified threshold range from the central feature point. The hyperedge set is defined as:

[0044]

[0045] Among them, represents the neighborhood of vertex v, d represents the cosine distance,

[0046] Adopt spectral domain hypergraph convolution to learn node features:

[0047]

[0048] Among them, and are two adjacent indicator functions, specifically expressed as follows:

[0049]

[0050] Ω is a trainable parameter;

[0051] Define the matrix formula for two-stage hypergraph information aggregation as:

[0052]

[0053] Among them, D v and D e represent the diagonal matrices of vertices and hyperedges respectively; H represents the incidence matrix of the hypergraph of.

[0054] Furthermore, the functional implementation process of the multi-condition fusion module is as follows:

[0055] The spatial attention mechanism SA and the channel attention mechanism CA are respectively used to enhance the semantic feature F S and the texture feature F T , ensuring effective capture of spatial and channel dependencies:

[0056] CA(x) = σ(MLP(AvgPool(x)) + MLP(MaxPool(x)))

[0057] SA(x) = σ(f 7×7 (Concat[AvgPool(x), MaxPool(x)]))

[0058]

[0059] where σ represents the ReLU activation function, MLP(·) represents a multi-layer perceptron, AvgPool(·) and MaxPool(·) represent average pooling and max pooling respectively, f 7×7 (·) represents a 7×7 convolutional layer, and Concat[·, ·] represents a channel concatenation operation;

[0060] The conditional encoder τ θ The encoded feature F obtained is processed by a BCR sub-module composed of batch normalization, a convolutional layer, and ReLU to obtain a refined feature Then and are concatenated to obtain F':

[0061]

[0062] The preliminarily fused feature is fed into the residual multi-layer perceptron IRMLP, which consists of convolutional layers with different kernel sizes and two linear transformation layers with ReLU non-linear transformation for expanding the channel dimension, to fuse the global and local features at each level:

[0063] IRMLP(x) = f 1×1 (f 1×1 (f depth3×3 (x) + x))

[0064] The conditional feature consists of the residual connection of F and IRMLP, and the convolutional layer is used for channel compression:

[0065] z c = Conv(IRMLP(LN(F')) + F).

[0066] Furthermore, the object detection model in step three is constructed based on YOLOv5.

[0067] Further, Step Five specifically includes the following process: fitting the extracted eardrum contour to obtain an eardrum contour curve, and constructing a geometric model in combination with the identified malleus handle and umbo in the ear. First, connect the malleus handle and the umbo and extend the line to intersect with the contour curve to obtain the first intersection point. Then, draw a perpendicular line to the line connecting the malleus handle and the umbo through the umbo, and the intersection point of the perpendicular line and the contour curve is the second intersection point. Connect the first intersection point and the second intersection point to make an auxiliary line, and draw a perpendicular line from the umbo to the auxiliary line to obtain the foot of the perpendicular. Finally, determine the midpoint between the first intersection point and the foot of the perpendicular, and select one of the midpoints as the catheterization point according to the mirror image relationship between the left and right ears.

[0068] The present invention provides an automatic positioning system for the catheterization point in tympanic hydrops surgery. The system has program modules corresponding to the steps of the method described in any one of the above technical solutions, and executes the steps in the automatic positioning method for the catheterization point in tympanic hydrops surgery described above when running.

[0069] The present invention provides a computer-readable storage medium storing a computer program configured to implement the steps in the automatic positioning method for the catheterization point in tympanic hydrops surgery described in any one of the above technical solutions when called by a processor.

[0070] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0071] The tympanic membrane segmentation model adopted by the present invention realizes high-precision segmentation of the eardrum contour in otoscope images through the dynamic fusion of texture features, semantic schemas and latent space representations, and identifies the malleus handle and umbo in the ear through an object detection model. Finally, by combining the geometric modeling of the eardrum contour and the image processing method, the catheterization point can be quickly and accurately located on the otoscope image, which can effectively reduce the surgical preparation time of doctors and improve the operation efficiency.

[0072] The present invention fully demonstrates the specific functions and bases of each step through the way of geometric construction and task decomposition, making the positioning result have higher clinical credibility and interpretability, and providing intuitive understanding support for doctors and patients.

[0073] The present invention reduces the data annotation cost and improves the generalization ability of the model. The present invention does not rely on a large-scale and finely annotated professional database, but by decomposing the catheterization point positioning task into general image processing subtasks, the annotation requirements are greatly reduced. At the same time, the step-by-step processing mode endows the system with higher adaptability, and it can still maintain high stability and accuracy even in otoscope images collected by different devices or complex lesion scenarios.

[0074] The present invention adopts a modular design, splitting the functions of data annotation, sub-task processing, and elliptical coordinate system construction into relatively independent sub-tasks, which facilitates subsequent upgrading, optimization, and integration of more functional modules, including lesion recognition and postoperative effect evaluation, thereby further enhancing the application value and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a flowchart of the automatic positioning method for the tympanostomy tube placement point in the embodiment of the present invention;

[0076] Figure 2 It is the tympanic membrane contour annotation effect (a) and the segmentation network training effect (b) in the embodiment of the present invention;

[0077] Figure 3 It is the YOLO algorithm in the embodiment of the present invention for locating 2 auxiliary points;

[0078] Figure 4 It is to determine the geometric relationship of the tympanostomy tube placement point in the embodiment of the present invention;

[0079] Figure 5 It is the actual positioning effect diagram in the embodiment of the present invention.

[0080] Figure 6 It is the clinical schematic diagram of tympanostomy (a) and the schematic diagram of the tympanostomy tube placement point position (b) in the background technology of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] In order to enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are only part of the embodiments or examples of the present invention, rather than all of them. Based on the embodiments or examples in the present invention, all other embodiments or examples obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0082] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.

[0083] Specific Embodiment 1: As Figure 1 shown, the present invention provides an automatic positioning method for the tympanostomy tube placement point in otitis media with effusion, including the following steps:

[0084] Step 1: Collect otoscope images, annotate the contour of the eardrum, construct a tympanic membrane contour data set, annotate the position of the manubrium mallei and the umbo, and construct an auxiliary point data set;

[0085] Step 2: Construct a tympanic membrane segmentation model TSLDSeg based on the latent diffusion model LDM, including a noise reduction generation unit, a perceptual compression unit, and a multi-modal conditional encoding unit;

[0086] The noise reduction generation unit is used to implement reverse diffusion using a deep U-Net architecture, generate a high-precision segmentation mask from the latent space noise distribution through iterative denoising, and construct a generative discriminative framework to probabilistically model the segmentation results;

[0087] The perceptual compression unit constructs a bidirectional mapping channel based on a variational autoencoder. The encoder is used to reduce the dimension of the segmentation results in the high-dimensional pixel space to a compact latent space representation, and the decoder is used to reconstruct the latent space to achieve data fidelity;

[0088] The multi-modal conditional encoding unit is used to extract image features, including a texture perception module, a semantic enhancement module, and a conditional segmentation module. The texture perception module is used to extract image edge features using Sobel and Laplacian differential operators and extract directional features using adaptive Gabor convolution; the semantic enhancement module is used to establish a hypergraph data structure for the image and mine the semantic relationships of the image; the multi-conditional fusion module is used to adaptively fuse texture features, semantic features, and encoded features by combining spatial and channel attention mechanisms and a residual multi-layer perceptron;

[0089] Use the tympanic membrane contour dataset to train the tympanic membrane segmentation model;

[0090] Step 3: Construct an object detection model; use the auxiliary point dataset to train the object detection model;

[0091] Step 4: Use the trained tympanic membrane segmentation model to segment the tympanic membrane contour of the otoscope image, and use the trained object detection model to identify the incus handle and umbo in the otoscope image;

[0092] Step 5: Fit the extracted tympanic membrane contour to obtain a tympanic membrane contour curve, combine the identified incus handle and umbo in the ear, and determine the catheterization point according to the mirror image relationship between the left and right ears.

[0093] Specific implementation plan 2: Step 1 includes desensitizing, cropping, and scaling the collected otoscope images. Other parts of this implementation plan are the same as those of Specific implementation plan 1.

[0094] Specific implementation plan 3: Use the tympanic membrane contour dataset to train the tympanic membrane segmentation model in Step 2; the process includes the following:

[0095] The perceptual compression unit uses a variational autoencoder E and D to implement the perceptual compression function, expressed as where is the tympanic membrane segmentation mask; the multi-modal conditional encoding unit constructs a conditional encoder τ with a structure similar to E θ realizes the conversion from the pixel space to the latent space, denoted as z c = τ θ (I), is the otoscope image, z c is the conditional feature; the model training is divided into two stages. In the first stage, the variational autoencoders E and D are trained using the dataset composed of masks. In the second stage, the variational autoencoders are frozen, and the denoising generation unit and the conditional encoder τ are trained using the complete training set θ to make the tympanic membrane segmentation model an end-to-end model;

[0096] During training, the denoising generation unit adds noise to the latent space representation z0 of the masked image, and then uses the denoising U-Net to predict the noise restore and generate the tympanic membrane mask, denoted as ε θ is the U-Net network;

[0097] The noise prediction loss in the denoising process is

[0098] The noise-adding forward process is:

[0099]

[0100] where ε represents Gaussian noise, is a hyperparameter that controls the noise level;

[0101] The noise-removing reverse process is:

[0102]

[0103] Introduce the loss function of supervised segmentation to generate discriminative segmentation. The total loss function of the tympanic membrane segmentation model is as follows:

[0104] L = L noise + λL seg

[0105] where λ represents the weight of the segmentation function, which is set to 1 in this implementation. Other parts of this implementation are the same as those of the first specific implementation

[0106] Specific implementation four: The texture perception module described in step two is as shown in Figure 1 (b), combining standard convolution and multi-scale convolution cascades. The boundary detection operator and the adaptive Gabor convolution extract boundary and direction features respectively, and then the features are concatenated to form texture features;

[0107] Perceptual compression and conditional encoding in LDM map images to a low-dimensional feature space, abstracting away high-frequency, imperceptible details. However, these details are crucial for medical image segmentation. TSLDSeg addresses this issue by combining trainable Gabor convolutional kernels, enabling the adaptive capture of meaningful directional features. The expression of the Gabor function is as follows:

[0108]

[0109] where x, y, λ, θ, γ, and ψ are all parameters of the Gabor function;

[0110] The mathematical expression for the texture perception module to extract image edges using Sobel differential operators and Laplacian operators in the horizontal and vertical directions is:

[0111]

[0112] The other parts of this implementation scheme are the same as those of the third specific implementation scheme.

[0113] Specific implementation scheme five: The functional implementation process of the semantic enhancement module is as follows:

[0114] Hyperedges of a hypergraph can connect multiple vertices, supporting simultaneous interactions to facilitate complex relationship modeling. Vertices connected by the same hyperedge exhibit semantic similarity. To enhance semantic information, the semantic enhancement module introduces a hypergraph construction method and its convolutional operation.

[0115] Use the conditional encoder τ θ Construct a hypergraph from the high-dimensional features of the intermediate layer where and ε represent the vertex set and the hyperedge set respectively; a threshold-based ε-ball is established for each feature point, representing a hyperedge that contains all feature points within the threshold range from the central feature point. As shown in Figure 1 (c), the hyperedge set is defined as:

[0116]

[0117] where represents the neighborhood of vertex v, d represents the cosine distance,

[0118] To promote the transmission of semantic information on the hypergraph structure, spectral domain hypergraph convolution is used to learn node features:

[0119]

[0120] where and are two adjacent indicator functions, specifically expressed as follows:

[0121]

[0122] Ω is a trainable parameter;

[0123] For the convenience of calculation, the matrix formula for two-stage hypergraph information aggregation is defined as:

[0124]

[0125] where, D v and D e respectively represent the diagonal matrices of vertices and hyperedges; H is the incidence matrix H representing the hypergraph .

[0126] Other parts of this implementation scheme are the same as those of the first specific implementation scheme.

[0127] Specific implementation scheme six: As shown in Figure 1 (d), the functional implementation process of the multi-condition fusion module is as follows:

[0128] The spatial attention mechanism SA and the channel attention mechanism CA are respectively used to enhance the semantic feature F S and the texture feature F T , ensuring the effective capture of spatial and channel dependencies:

[0129] CA(x) = σ(MLP(AvgPool(x)) + MLP(MaxPool(x)))

[0130] SA(x) = σ(f 7×7 (Concat[AvgPool(x), MaxPool(x)]))

[0131]

[0132] where, σ represents the ReLU activation function, MLP(·) represents the multi-layer perceptron, AvgPool(·) and MaxPool(·) respectively represent average pooling and max pooling, f 7×7 (·) represents a 7×7 convolutional layer, and Concat[·, ·] represents the channel concatenation operation;

[0133] The encoded feature F obtained by the conditional encoder τ θ is processed by a BCR sub-module composed of batch normalization, convolutional layer and ReLU to obtain a refined feature Then and are concatenated to obtain F':

[0134]

[0135] Send the preliminarily fused features into the Residual Perceptual Machine IRMLP, which consists of convolutional layers with different kernel sizes and linear transformation layers with 2 ReLU non-linear transformations for expanding the channel dimension, alleviating gradient vanishing, explosion and network degradation to a certain extent, and fusing the global and local features at each level:

[0136] IRMLP(x) = f 1×1 (f 1×1 (fdepth 3×3 (x) + x))

[0137] Conditional feature Consists of the residual connection of F and IRMLP, and the convolutional layer is used for channel compression:

[0138] z c = Conv(IRMLP(LN(F')) + F).

[0139] Other parts of this implementation scheme are the same as those of the first specific implementation scheme.

[0140] Specific implementation scheme seven: The object detection model described in step three is constructed based on YOLOv5.

[0141] Other parts of this implementation scheme are the same as those of the sixth specific implementation scheme.

[0142] This implementation scheme of YOLOv5 has excellent real-time detection performance, can quickly process high-resolution otoscope images, and still accurately identify target points under complex backgrounds or changing light conditions.

[0143] The object detection model of this implementation scheme supports multiple model sizes, including lightweight versions and enhanced versions, which is convenient for flexible selection and integration according to specific requirements.

[0144] Specific implementation scheme eight: Step five specifically includes the following process: Fit the extracted tympanic membrane contour to obtain the tympanic membrane contour curve, and construct a geometric model by combining the identified malleus handle and umbo in the ear. First, connect the malleus handle and umbo and extend the line to intersect with the contour curve to obtain intersection point one. Then, draw a perpendicular line from the umbo to the line connecting the malleus handle and umbo, and the intersection point of the perpendicular line and the contour curve is intersection point two. Connect intersection point one and intersection point two to make an auxiliary line, draw a perpendicular line from the umbo to the auxiliary line to obtain the foot of the perpendicular. Finally, determine the midpoint between intersection point one and the foot of the perpendicular, and select one midpoint as the catheterization point according to the mirror image relationship between the left and right ears.

[0145] Specifically as Figure 4 shown, fit the extracted tympanic membrane contour to obtain the tympanic membrane contour curve, and determine the catheterization point according to the mirror image relationship between the identified malleus handle and umbo in the ear, as Figure 5 shown.

[0146] The elliptical tympanic membrane is obtained by least - squares fitting, and the standard equation of the ellipse is obtained according to plane geometric relations:

[0147]

[0148] Among them, the center coordinates of the ellipse are (x0, y0), a and b respectively represent the major axis and minor axis of the ellipse, and θ represents the rotation angle of the ellipse.

[0149] As Figure 4 shown is a schematic diagram of the geometric structure abstracted from the otoscope image, where point A is the handle of the malleus, point O is the umbo, connect AO and extend it to B, OC is perpendicular to OA, connect BC, draw a perpendicular line from point O perpendicular to BC at point E, and the mid - point of BE is the catheterization point.

[0150] According to the mirror image relationship between the left and right ears, point C of the other ear is symmetric about AB, that is, Figure 4 the position of point D in

[0151] To determine point C, according to the orientation of the left and right ears in the otoscope image, calculate the cross - product of OA and OC:

[0152]

[0153] If it is the left ear, then C is on the left side of OA, and the cross - product result is negative. Conversely, it is on the right side in the mirror image.

[0154] Other aspects of this implementation scheme are the same as those of the seventh specific implementation scheme.

[0155] The present invention determines the coordinates (pixel coordinates) of the catheterization point in a given otoscope image, and the visualization result is as Figure 5 shown. Taking photos and sampling with an ear endoscope in the Shandong Provincial Center for Public Health Clinical Medicine, the positioning results of more than 2,200 images are obtained, and the success rate can reach 92%.

[0156] An automatic positioning method (algorithm) for the catheterization point in the tympanic membrane hydrops surgery proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on this algorithm.

[0157] Based on the method proposed by the present invention, an automatic positioning system for the catheterization point in the tympanic membrane hydrops surgery is developed using a programming language. This system has program modules corresponding to the steps of the above - mentioned technical solution, and executes the steps in the above - mentioned automatic positioning method for the catheterization point in the tympanic membrane hydrops surgery when running.

[0158] The computer program of the developed system (software) is stored on a computer - readable storage medium, and the computer program is configured to implement the steps of the above - mentioned automatic positioning method for the catheterization point in the tympanic membrane hydrops surgery when called by a processor. That is, the present invention is materialized on a carrier and becomes a computer program product.

[0159] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The computational programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor and can implement these computational programs using high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disks, optical disks, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0161] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. An automatic positioning method for the catheterization point in the surgery of middle ear effusion, characterized in that, It includes the following steps: Step 1: Collect otoscope images, annotate the contour of the eardrum, construct a tympanic membrane contour dataset, annotate the positions of the malleus handle and umbilicus, and construct an auxiliary point dataset; Step 2: Construct a tympanic membrane segmentation model based on the latent diffusion model, including a noise reduction generation unit, a perceptual compression unit, and a multi-modal conditional encoding unit; The noise reduction generation unit is used to implement reverse diffusion using a deep U-Net architecture, generate a high-precision segmentation mask from the latent space noise distribution by iterative denoising, and construct a generative discriminative framework to probabilistically model the segmentation result; The perceptual compression unit constructs a bidirectional mapping channel based on the variational autoencoder. The encoder is used to reduce the dimensionality of the segmentation result in the high-dimensional pixel space to a compact latent space representation, and the decoder is used to reconstruct the latent space to achieve data fidelity; The multi-modal conditional encoding unit is used to extract image features, including a texture perception module, a semantic enhancement module, and a conditional segmentation module. The texture perception module is used to extract image edge features using Sobel and Laplacian differential operators and extract direction features using adaptive Gabor convolution; The semantic enhancement module is used to establish a hypergraph data structure for the image and mine the semantic relationships of the image; the multi-condition fusion module is used to adaptively fuse texture features, semantic features, and encoding features by combining spatial and channel attention mechanisms and residual multi-layer perceptrons; Use the tympanic membrane contour dataset to train the tympanic membrane segmentation model; Step 3: Construct an object detection model; use the auxiliary point dataset to train the object detection model; Step 4: Use the trained tympanic membrane segmentation model to segment the tympanic membrane contour of the otoscope image, and use the trained object detection model to identify the malleus handle and umbilicus in the ear of the otoscope image; Step 5: Fit the extracted tympanic membrane contour to obtain a tympanic membrane contour curve, combine the identified malleus handle and umbilicus in the ear, and determine the catheterization point according to the mirror image relationship between the left and right ears.

2. The automatic positioning method for the tympanostomy point of the otitis media with effusion surgery according to claim 1, wherein Step 1 includes desensitization, cropping, and scaling processing of the collected otoscope images.

3. The automatic positioning method for the catheterization point in the tympanic membrane hydrops surgery according to claim 1, characterized in that, In Step 2, the tympanic membrane contour dataset is used to train the tympanic membrane segmentation model; it includes the following process: The perception compression unit uses a variational autoencoder E and D to implement the perception compression function, denoted as where is the eardrum segmentation mask; the multimodal conditional encoding unit constructs a conditional encoder τ with a structure similar to E θ to implement the conversion from the pixel space to the latent space, denoted as z c = τ θ (I), is the otoscope image, and z c is the conditional feature; The model training is divided into two stages. In the first stage, the variational autoencoders E and D are trained using a dataset composed of masks. In the second stage, the variational autoencoders are frozen, and the denoising generation unit and the conditional encoder τ are trained using the complete training set θ , making the tympanic membrane segmentation model an end-to-end model; During training, the noise reduction generation unit adds noise to the latent space representation z0 of the masked image, and then uses the denoising U-Net to predict the noise Restore and generate the eardrum mask, denoted as ε θ is the U-Net network; The noise prediction loss of the noise reduction process is The noise addition forward process is: where ε represents Gaussian noise, is a hyperparameter that controls the noise level; The reverse noise reduction process is: Introduce the loss function of supervised segmentation For generative discriminative segmentation, the total loss function of the tympanic membrane segmentation model is as follows: L = L noise + λL seg Among them, λ represents the segmentation function weight.

4. The automatic positioning method for the tympanostomy point in the operation for middle ear effusion according to claim 3, wherein, In Step 2, the texture perception module combines standard convolution and multi-scale convolution cascades, and the boundary detection operator and adaptive Gabor convolution extract boundary and direction features respectively, and then splice the features to form texture features; The mathematical expressions for the texture perception module to extract the image edge using the Sobel differential operator and Laplacian operator in the horizontal and vertical directions are:

5. The automatic positioning method for the catheterization point of the eardrum hydrops surgery according to claim 1, characterized in that, The functional implementation process of the semantic enhancement module is: Use the conditional encoder τ θ Construct a hypergraph with high-dimensional features in the intermediate layer where and ε represent the vertex set and the hyperedge set respectively; a threshold-based ε-ball is established for each feature point, representing a hyperedge that contains all feature points within the specified threshold range from the central feature point. The hyperedge set is defined as: Among them, represents the neighborhood of vertex v, and d represents the cosine distance. Use spectral domain hypergraph convolution to learn node features: Among them, and are two adjacent indicator functions, which are specifically represented as follows: Ω is a trainable parameter; Define the matrix formula for two-stage hypergraph information aggregation as: Among them, D v and D e represent the diagonal matrices of vertices and hyperedges respectively; H is the incidence matrix representing the hypergraph .

6. The automatic positioning method for the tympanostomy tube insertion point in the operation of otitis media with effusion according to claim 1, wherein The functional implementation process of the multi-condition fusion module is: The spatial attention mechanism SA and the channel attention mechanism CA are respectively used to enhance the semantic feature F S and the texture feature F T , ensuring the effective capture of spatial and channel dependencies: CA(x) = σ(MLP(AvgPool(x)) + MLP(MaxPool(x))) SA(x) = σ(f 7×7 (Concat[AvgPool(x), MaxPool(x)])) Among them, σ represents the ReLU activation function, MLP(·) represents the multi-layer perceptron, AvgPool(·) and MaxPool(·) represent average pooling and max pooling respectively, and f 7×7 (·) represents a 7×7 convolutional layer, and Concat[·, ·] represents the channel concatenation operation; Conditional encoder τ θ The obtained encoded feature F is processed by a BCR sub-module composed of batch normalization, convolutional layer, and ReLU to obtain a refined feature Then and are concatenated to obtain F': Send the preliminarily fused features into the residual perception machine IRMLP, which consists of convolutional layers with different kernel sizes and linear transformation layers with 2 ReLU non-linear transformations for expanding the channel dimension, to fuse the global and local features at each level: IRMLP(x) = f 1×1 (f 1×1 (f depth3×3 (x) + x)) Conditional Feature It consists of the residual connection of F and IRMLP, and the convolutional layer is used for channel compression: z c = Conv(IRMLP(LN(F′)) + F).

7. The automatic positioning method for the tympanostomy tube insertion point in the operation for serous otitis media according to claim 6, characterized in that, The object detection model in step three is constructed based on YOLOv5.

8. The automatic positioning method for the catheterization point in the operation of middle ear effusion according to claim 7, wherein Step five specifically includes the following process: Fit the extracted eardrum contour to obtain the eardrum contour curve, and construct a geometric model in combination with the identified malleus handle and umbo in the ear. First, connect the malleus handle and the umbo and extend the line to intersect with the contour curve to obtain intersection point one. Then, draw a perpendicular line from the umbo to the line connecting the malleus handle and the umbo, and the intersection point of the perpendicular line and the contour curve is intersection point two. Connect intersection point one and intersection point two to make an auxiliary line, and draw a perpendicular line from the umbo to the auxiliary line to obtain the foot of the perpendicular. Finally, determine the midpoint between intersection point one and the foot of the perpendicular, and select one of the midpoints as the catheterization point according to the mirror image relationship between the left and right ears.

9. An automatic positioning system for the catheterization point in the surgery of middle ear effusion, characterized in that, The system has program modules corresponding to the steps of the method described in any one of claims 1 to 8 above, and executes the steps in the automatic positioning method of the catheterization point for the tympanic hydrops surgery described above when running.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps in the automatic positioning method of the catheterization point for the tympanic hydrops surgery described in any one of claims 1 to 8 when called by a processor.

Citation Information

Cited By

  • Tympanic membrane puncture positioning method and system based on clinical image

    CN121606439A