Ear-nose-throat endoscope image-based lesion segmentation method, system, device and medium

By using time-triggered dual-spectral synchronous acquisition and dynamic mucus displacement modeling, combined with cross-modal feature registration and deformable convolutional layer correction, the problems of mucus interference and multimodal registration in ENT endoscopic images were solved, enabling precise lesion localization and real-time diagnosis and treatment decisions.

CN121481964BActive Publication Date: 2026-05-08HE BEI SHENG ZHONG YI YUAN (FIRST AFFILIATED HOSPITAL OF HEBEI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE HEBEI CENTER FOR PREVENTION & CONTROL OF SCOLIOSIS IN CHILDREN & ADOLESCENTS)
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HE BEI SHENG ZHONG YI YUAN (FIRST AFFILIATED HOSPITAL OF HEBEI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE HEBEI CENTER FOR PREVENTION & CONTROL OF SCOLIOSIS IN CHILDREN & ADOLESCENTS)
Filing Date
2025-11-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In ENT endoscopic image processing, mucus interference leads to a high rate of missed lesion detection, multimodal image registration fails, and separation of diagnosis and treatment processes leads to decision delays and errors.

Method used

By using time-triggered dual-spectral synchronous acquisition, dynamic mucus displacement modeling, cross-modal feature registration, and deformable convolutional layer correction, a lesion probability map is generated and real-time navigation markers and treatment decisions are output.

Benefits of technology

It achieves the elimination of mucus interference and accurate registration of multimodal features, improving the visualization integrity of lesion areas and the real-time performance and accuracy of diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481964B_ABST
    Figure CN121481964B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on ear-nose-throat endoscope image lesion segmentation method, system, equipment and medium, the method is by time sequence trigger synchronous acquisition dual-spectrum image, solve the problem of anatomical structure obstruction caused by mucus flow;Dynamic mucus displacement field is modeled using pixel gradient, eliminate the misjudgment of traditional segmentation method to static lesion and dynamic secretion;Variable deformation convolution layer is used to correct the spatial offset of white light and narrowband image, overcome the mismatch of multimodal features due to optical scattering;Finally, based on the topological properties of lesion probability map, real-time surgical navigation markers and clinical treatment plan are generated simultaneously, realizing the closed-loop link from image analysis to diagnosis and treatment decision-making.This method breaks through the integration of mucus interference suppression, cross-modality precise registration and real-time diagnosis and treatment auxiliary function, significantly improves the lesion recognition accuracy of endoscope image and clinical operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to a method, system, device, and medium for lesion segmentation based on otolaryngological endoscope images. Background Technology

[0002] Otorhinolaryngological endoscopy, a crucial tool for clinical diagnosis, provides physicians with visual information about the mucosal surface and vascular structures through the combined imaging of white light and narrowband spectroscopy. Traditional diagnostic procedures rely on physicians' visual interpretation of real-time endoscopic images, combining the anatomical information from white light imaging with the vascular enhancement characteristics of narrowband imaging to locate and determine the nature of lesions. However, the complex physiological environment of the ENT cavities, especially dynamic mucus coverage, organ peristalsis, and spatial mismatch issues in multimodal imaging, severely restricts diagnostic accuracy and the reliability of real-time surgical assistance.

[0003] In existing technologies, endoscopic image processing faces three major bottlenecks: First, the fluidity and adhesion of mucus within the cavity cause anatomical structures in white light images to be partially or completely obscured, and traditional threshold segmentation methods cannot distinguish between dynamic mucus and static lesions, resulting in a significant increase in the rate of missed lesion detection. Second, due to the difference in optical scattering characteristics between white light and narrowband images, subpixel-level spatial offset occurs, and existing registration algorithms (such as affine transformation) are unable to compensate for the nonlinear misalignment caused by mucosal deformation, leading to the failure of multimodal feature fusion. Third, lesion segmentation results are separate from surgical navigation and treatment plan generation in independent systems, requiring doctors to manually mark lesion locations and consult guidelines, which delays intraoperative decision-making and easily introduces subjective errors. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a method, system, device and medium for lesion segmentation based on ENT endoscope images that can simultaneously solve the problems of mucus interference elimination, multimodal accurate registration and real-time diagnosis and treatment closed loop.

[0005] The objective of this invention is achieved through the following solution:

[0006] In a first aspect, the present invention provides a lesion segmentation method based on endoscopic images of the ear, nose, and throat, comprising the following steps:

[0007] S1: Synchronously acquire the timing trigger signal of the endoscope light source, and capture timing-aligned dual-spectral images by alternately driving the white light source and the narrowband filter device, thereby generating a timing-aligned white light image sequence and a narrowband imaging image.

[0008] S2: Model the mucus displacement of the white light image sequence, calculate the displacement vector of the dynamic mucus by the pixel gradient change of adjacent frames, and generate the mucus motion vector field;

[0009] S3: Masking repair is performed on the mucus motion vector field and white light image sequence. The mucus-covered area is segmented by displacement amplitude threshold and the occluded anatomical structure pixels are reconstructed to generate an optimized white light image with mucus interference removed.

[0010] S4: Perform cross-modal feature registration on the optimized white light image and narrowband imaging image, extract the white light feature map and narrowband feature map respectively, and correct the spatial offset between the feature maps through deformable convolutional layers to generate the registered multimodal feature map;

[0011] S5: Perform lesion segmentation on the multimodal feature map, generate lesion probability map by feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopic video and treatment decision schemes matching clinical guidelines based on the spatial coordinates and area attributes of the connected domains in the lesion probability map.

[0012] Among them, real-time navigation markers are used to indicate the anatomical location of lesions and the probability of malignancy, while treatment decision plans are used to indicate surgical or drug intervention plans that match the type and size of the lesions.

[0013] In one embodiment, S1 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0014] S11: Perform time-division triggering processing on the control signal of the endoscope light source, and generate time-division imaging control commands without spectral crosstalk by alternately switching the phase of the driving current of the white light source and the narrowband filter.

[0015] S12: Input the time-division imaging control command into the image sensor driving circuit of the endoscope imaging system, perform frame synchronization processing on the raw photoelectric conversion data captured by the sensor, and generate a bispectral image with exposure time alignment by compensating for the start time offset between white light and narrowband exposure.

[0016] S13: Motion artifact detection processing is performed on the dual-spectral images. Through cross-correlation analysis of adjacent frames in the white light channel and the narrowband channel, distorted frames caused by organ tremors are removed, and time-aligned white light image sequences and narrowband imaging images are generated.

[0017] In one embodiment, S2 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0018] S21: Perform viscous rheological property analysis on white light image sequences, calculate constitutive equation parameters of viscous shear stress and strain rate through non-rigid transformation relationship of pixel gradient field between adjacent frames, and generate viscous viscoelastic parameters.

[0019] S22: Model the displacement field of the viscoelastic parameters of the viscous fluid, solve the displacement vector equation through viscous fluid dynamic constraints, substitute the viscoelastic parameters into the preset Oldroyd-B model to calculate the stress tensor distribution of the viscous fluid flow, and generate a dynamic viscous fluid displacement field.

[0020] S23: Noise suppression is performed on the dynamic mucus displacement field. The anatomical structure motion and mucus flow components are separated by low-pass filtering. Abnormal pulsation signals with frequencies higher than the physiological motion threshold are filtered out to generate a denoised mucus motion vector field.

[0021] In one embodiment, S3 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0022] S31: Perform amplitude adaptive segmentation on the mucus motion vector field, and generate a dynamic threshold mask boundary associated with the mucus adhesion strength by calculating the spatiotemporal accumulation of motion vectors in the local neighborhood.

[0023] S32: Reconstruct the anatomical structure of the area covered by the dynamic threshold mask boundary, restore the continuity of the mucosal gland direction in the occluded area through the neighborhood healthy tissue texture propagation algorithm, and generate a preliminary reconstructed white light image.

[0024] S33: Optimize the edge consistency of the white light reconstructed image, eliminate step artifacts in the texture transition area through multi-scale gradient fusion, and generate an optimized white light image with continuous anatomical structure.

[0025] In one embodiment, S4 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0026] S41: Perform multi-resolution feature extraction on the optimized white light image, and generate multi-scale feature response maps covering macroscopic anatomical structures to microscopic textures through convolutional pyramid decomposition, thereby generating cross-scale white light feature maps.

[0027] S42: Enhance the frequency domain features of narrowband imaging images by strengthening the frequency domain response of blood vessel morphology and mucosal gland openings corresponding to the 415nm and 540nm bands through bandpass filtering, and generate enhanced narrowband feature maps.

[0028] S43: Perform spatial offset correction on the white light feature map and the enhanced narrowband feature map, learn the feature deformation parameters between the white light and narrowband features through deformable convolution kernels, and generate a registered multimodal feature map.

[0029] In one embodiment, the formula for calculating the characteristic deformation parameters of a lesion segmentation method based on ENT endoscopy images provided by the present invention is as follows:

[0030]

[0031]

[0032] in, For characteristic deformation parameters, The target location coordinates, The sampling grid is 3×3. For bilinear interpolation weights, It is characterized by narrow bands. For learnable offsets, It is characterized by white light.

[0033] In one embodiment, step S5 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0034] S51: Perform feature channel splicing on the multimodal feature map, generate a fused feature tensor by splicing white light and narrowband feature map along the channel axis, and input the fused feature tensor into the U-shaped decoding network for upsampling and feature fusion to generate a lesion probability map;

[0035] S52: Perform connected component attribute parsing on the lesion probability map. By scanning connected regions in the probability map that are greater than a set threshold, calculate the centroid spatial coordinates and pixel area of ​​each connected region to generate quantitative attributes of lesion location and size.

[0036] S53: Real-time navigation marker generation based on quantized attributes, mapping the centroid coordinates of the lesion to the endoscopic video coordinate system through three-dimensional projection transformation, generating real-time navigation markers for the endoscopic video;

[0037] S54: Based on quantitative attributes, clinical decision matching is performed. By querying the pre-set guideline knowledge base for treatment rules that match the location and area of ​​the lesion, a treatment decision plan matching the clinical guidelines is generated.

[0038] Secondly, the present invention provides a lesion segmentation system based on otolaryngological endoscopic images, the system being configured with the following modules:

[0039] The dual-spectral image synchronous acquisition module is used to synchronously acquire the timing trigger signal of the endoscope light source. By alternately driving the white light source and the narrowband filter device, it captures timing-aligned dual-spectral images and generates timing-aligned white light image sequences and narrowband imaging images.

[0040] The slime displacement modeling module is used to model slime displacement in white light image sequences. It calculates the displacement vector of dynamic slime by changing the pixel gradient between adjacent frames and generates a slime motion vector field.

[0041] The slime mask repair and image optimization module is used to perform mask repair on the slime motion vector field and white light image sequence. It segments the slime-covered area by displacement amplitude threshold and reconstructs the occluded anatomical structure pixels to generate an optimized white light image with slime interference removed.

[0042] The cross-modal feature registration module is used to perform cross-modal feature registration on optimized white light images and narrowband imaging images. It extracts white light feature maps and narrowband feature maps respectively, and corrects the spatial offset between feature maps through deformable convolutional layers to generate registered multimodal feature maps.

[0043] The lesion segmentation and treatment plan generation module is used to segment lesions from multimodal feature maps, generate lesion probability maps through feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopic video and treatment decision plans matching clinical guidelines based on the spatial coordinates and area attributes of the connected domains in the lesion probability map.

[0044] Among them, real-time navigation markers are used to indicate the anatomical location of lesions and the probability of malignancy, while treatment decision plans are used to indicate surgical or drug intervention plans that match the type and size of the lesions.

[0045] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the above-mentioned lesion segmentation methods based on ENT endoscope images.

[0046] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned lesion segmentation methods based on ENT endoscope images.

[0047] In summary, the lesion segmentation method based on ENT endoscopy images provided in this application achieves strict frame alignment between white light and narrowband imaging through a time-triggered dual-spectrum synchronous acquisition mechanism, thereby eliminating ghosting interference caused by organ peristalsis. Dynamic mucus displacement modeling based on pixel gradient fields accurately captures the flow vector of mucus, effectively distinguishing the secretion-covered area from the actual anatomical structure. Displacement amplitude threshold segmentation and texture reconstruction strategies restore the continuity of mucosal glands and vascular networks obscured by mucus, improving the visualization integrity of the lesion area. Adaptive learning of deformation parameters of deformable convolutional layers enables sub-pixel spatial registration of white light and narrowband features, overcoming multimodal feature misalignment caused by differences in spectral scattering characteristics. Finally, through topological attribute analysis of the lesion probability map and mapping with clinical guideline rules, anatomical positioning navigation markers and graded treatment decisions superimposed on real-time video are generated synchronously, enabling precise intraoperative lesion localization and immediate output of evidence-based medical plans.

[0048] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0049] Figure 1 A flowchart illustrating a lesion segmentation method based on ENT endoscopic images provided in this application embodiment;

[0050] Figure 2 A schematic diagram illustrating the process of generating multimodal feature maps provided in an embodiment of this application;

[0051] Figure 3 This is a schematic diagram of a lesion segmentation system based on ENT endoscope images, provided as another embodiment of this application. Detailed Implementation

[0052] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0054] In one embodiment, such as Figure 1 As shown, a lesion segmentation method based on ENT endoscopic images is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and furthermore, to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0055] S1: Synchronously acquire the timing trigger signal of the endoscope light source, and capture timing-aligned dual-spectral images by alternately driving the white light source and the narrowband filter device, generating a timing-aligned white light image sequence and a narrowband imaging image.

[0056] Specifically, the system performs time-synchronized acquisition of dual-spectrum images, with timing control as its core. Timing alignment of white light and narrowband images is achieved through signal coordination between hardware components. Preferably, the system can use a field-programmable gate array (FPGA) as the timing controller. This controller establishes connections with the endoscope system's white LED light source, narrowband filter device, and CMOS image sensor (CIS), forming a complete signal transmission and control link. The core function of the timing controller is to output two synchronous trigger signals. One is a light source trigger signal, which drives the white LED light source to turn on and the narrowband filter device to switch in an alternating high-level manner, ensuring that each trigger corresponds to only one spectral mode and avoiding the simultaneous action of two spectra in the same drive. The other is an image acquisition trigger signal, which is synchronized with the rising edge of the light source trigger signal and triggers the CMOS image sensor exposure after a certain delay following the output of the light source trigger signal. The purpose of the delay is to compensate for the response lag during the light source illumination process, ensuring that the light source is already operating stably when the sensor is exposed, thereby avoiding image distortion caused by spectral mixing.

[0057] After image acquisition, the system stitches the acquired images sequentially using the timing controller's caching module, ultimately generating two types of time-aligned image sequences: a white light image sequence and a narrowband imaging image sequence. Images in different spectral bands within the narrowband imaging sequence are distinguished using frame indexes. The system stores both types of image sequences in a high-speed storage device that can handle real-time image writing and reading, providing data support for subsequent image processing. Throughout the acquisition process, the system continuously monitors the signal status of each component through the timing controller to ensure the synchronization and stability of trigger signals, thereby guaranteeing that the dual-spectral images correspond to the same anatomical location within a cavity.

[0058] S2: Model the mucus displacement of the white light image sequence, calculate the displacement vector of the dynamic mucus by the pixel gradient change of adjacent frames, and generate the mucus motion vector field.

[0059] Specifically, the system models the mucus displacement in white light image sequences to generate a mucus motion vector field. Dynamic mucus exhibits fluidity and adhesiveness within cavities, and its pixel grayscale change rate differs from that of static anatomical structures. The system calculates the gradient changes of adjacent frame pixels based on optical flow methods, using this difference to distinguish the motion of mucus from that of anatomical structures, thereby generating a dense displacement vector field for the mucus. During processing, the system preprocesses the white light image sequences, including filtering and grayscale conversion. Filtering removes image noise, while grayscale conversion converts RGB images into single-channel grayscale images to avoid the influence of color channel differences on gradient calculation.

[0060] After preprocessing, the system selects two consecutive grayscale images from the white light image sequence for displacement calculation, calculating the displacement components of each pixel in the x and y directions to form an initial displacement vector field. Since the anatomical structure exhibits minute peristalsis, which may interfere with the mucus displacement calculation, the system uses the RANSAC algorithm to remove outliers from the initial displacement vector field. By randomly sampling some displacement vectors as inlier samples, a motion model of the anatomical structure is fitted. This model assumes that the anatomical structure motion is rigid translation and the mucus motion is nonlinear flow. Subsequently, the system calculates the residual between each pixel's displacement vector and the model, removing anatomical structure displacement vectors whose residuals exceed the allowable range, and retaining the mucus displacement vectors.

[0061] After generating the initial slime displacement vector field, the system verifies the vector field and calculates the average displacement amplitude of the slime motion vector field. If the average displacement amplitude is at a low level, it indicates that the slime motion is weak and may be static slime. At this time, the system automatically switches to the displacement calculation method based on inter-frame difference and re-performs the displacement calculation to ensure the accuracy of slime region identification and finally generates a slime motion vector field that meets the requirements.

[0062] S3: Masking and repairing the mucus motion vector field and white light image sequence, segmenting the mucus-covered area by displacement amplitude thresholding and reconstructing the occluded anatomical structure pixels to generate an optimized white light image free of mucus interference.

[0063] Specifically, the system performs masking and restoration on the mucus motion vector field and white light image sequence to remove mucus interference. Based on the mucus motion vector field, the system locates the mucus-covered area and distinguishes between mucus and non-mucus regions using a displacement amplitude thresholding method. Because the displacement amplitude of mucus is significantly greater than that of anatomical structures, the system can use an adaptive thresholding algorithm to generate a segmentation threshold. This algorithm can dynamically adjust the threshold according to different cavity scenes to ensure segmentation accuracy. The system calculates the displacement amplitude of each pixel in the mucus motion vector field, marking pixels with displacement amplitudes exceeding the segmentation threshold as mucus regions, generating a binary mask image where mucus and non-mucus regions are distinguished by different pixel values.

[0064] After generating a binary mask image, the system optimizes it by filling the internal pores of the slime region and eliminating isolated noise points at the edges through morphological operations. Dilation and erosion operations are performed first to obtain an optimized slime mask with clear edges and a complete internal structure. Subsequently, the system performs image inpainting on the optimized slime mask and the original white light image using a texture consistency-based image inpainting algorithm. This algorithm starts the inpainting process from the edge pixels of the slime mask, calculating the texture priority of each pixel to be repaired. The texture priority is determined by texture similarity and confidence. The system prioritizes repairing areas with high texture consistency with the surrounding area. After repairing each pixel, the system updates the mask and confidence map, iterating until all slime region pixels are repaired.

[0065] After the repair work is completed, the system verifies the repair effect. It evaluates the consistency between the repaired area and the adjacent non-repaired area in the optimized white light image by using the structural similarity index and peak signal-to-noise ratio. If the evaluation result does not meet the requirements, the system adjusts the repair window size and repeats the repair operation until the repair effect meets the clinical application requirements and generates an optimized white light image with mucus interference removed.

[0066] S4: Perform cross-modal feature registration on the optimized white light image and narrowband imaging image, extract the white light feature map and narrowband feature map respectively, and correct the spatial offset between the feature maps through deformable convolutional layers to generate the registered multimodal feature map.

[0067] Specifically, the system performs cross-modal feature registration between optimized white light images and narrowband imaging images to generate registered multimodal feature maps. The white light images emphasize anatomical structural information, while the narrowband imaging images emphasize vascular features. Due to differences in optical scattering characteristics, there is a sub-pixel-level spatial offset between the two, and this offset is non-linear, caused by mucosal deformation. Therefore, the system first extracts high-level semantic features from the dual-modal images, and then corrects the spatial offset between the feature maps using deformable convolutional layers to achieve feature-level registration.

[0068] Preferably, the system can employ a backbone network-based feature extraction structure for feature extraction. Fully connected layers of the backbone network are removed, while a specified number of convolutional layers are retained. Both the optimized white light image and the narrowband imaging image are normalized before being input into the feature extraction structure. Multi-scale feature maps are extracted from both images, with each scale corresponding to a different resolution and number of channels. After feature extraction, the system uses the feature map of the narrowband imaging image as a reference feature map and the feature map of the optimized white light image as the feature map to be registered, performing multi-scale registration in a coarse-to-fine order.

[0069] In the coarse registration stage, the system selects the feature map with the lowest resolution for processing. The feature map to be registered is input into a deformable convolutional layer, which contains branch networks. These branch networks predict the offset of each convolutional kernel sampling point based on the feature differences between the feature map to be registered and the reference feature map. The system adjusts the feature sampling positions of the feature map to be registered using these offsets until the mutual information between the two feature maps meets the requirements. In the fine registration stage, the system enlarges the feature map to be registered from the previous scale to the current scale using interpolation methods. It then aligns the features with the reference feature map at the current scale, repeating the deformable convolutional correction process to gradually improve the registration accuracy until the feature distance of the highest resolution feature map meets the requirements.

[0070] After registration, the system stitches the registered optimized white light image feature map and the narrowband imaging image feature map along the channel dimension to generate a registered multimodal feature map. Preferably, to ensure that the registration accuracy meets clinical requirements, the system performs registration verification by randomly selecting anatomical landmarks, manually marking their coordinates in the optimized white light image and the narrowband imaging image, and calculating the average offset error of the registered landmarks. If the error exceeds the allowable range, the system re-performs the registration operation until the error meets the requirements.

[0071] S5: Perform lesion segmentation on the multimodal feature map, generate lesion probability map by feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopic video and treatment decision schemes matching clinical guidelines based on the spatial coordinates and area attributes of the connected domains in the lesion probability map.

[0072] Among them, real-time navigation markers are used to indicate the anatomical location of lesions and the probability of malignancy, while treatment decision plans are used to indicate surgical or drug intervention plans that match the type and size of the lesions.

[0073] Specifically, the system segments lesions from the multimodal feature map and generates real-time navigation markers and treatment decision schemes based on the segmentation results. Preferably, the system can use a U-shaped decoding network for lesion segmentation. The encoding end of this network is a backbone network convolutional layer used in the feature extraction process, and the decoding end is a transposed convolutional layer. An attention gate is set between the encoding and decoding ends to enhance the feature response of the lesion region and suppress the interference of background region features. The system inputs the registered multimodal feature map into the U-shaped decoding network. The encoding end extracts the deep semantic features of the lesion, and the decoding end restores the feature map resolution to the original image resolution through transposed convolution. The network output layer generates a lesion probability map through an activation function. The system determines the lesion region based on the probability map and marks the region with a probability value exceeding a threshold as the lesion region.

[0074] After the lesion region is determined, the system performs connected component analysis, extracting connected components of independent lesions using the neighborhood connectivity criterion and calculating the spatial attributes of each connected component. The system calculates the lesion center coordinates using the mean of the pixel coordinates of the connected components, and converts the pixel coordinates into actual anatomical coordinates by combining the endoscope camera calibration parameters; it calculates the actual lesion area by counting the number of pixels in the connected components and combining the camera pixel size and working distance; based on the vascular density and lesion boundary irregularity extracted from the narrowband imaging image, it calculates the lesion malignancy probability using the grading formula in the clinical guideline knowledge base, where vascular density is the ratio of the number of vascular pixels to the lesion area, and boundary irregularity is calculated through the relationship between the lesion contour perimeter and the lesion area.

[0075] The system generates real-time navigation markers based on the malignancy probability of lesions. The marker color and content are determined according to the malignancy probability grading, and the navigation markers are overlaid on the corresponding lesion locations in the real-time endoscopic video. The marker positions are dynamically updated as the endoscope lens moves. Simultaneously, the system matches the feature vectors output by the segmentation network with lesion type templates in the clinical guideline knowledge base to determine the lesion type. Then, combining the lesion area with the treatment plan mapping table in the knowledge base, it generates a treatment decision plan. The system records lesion attributes, navigation marker data, and treatment decision plans in a standard format and synchronously outputs them to the hospital's PACS (Picture Archiving and Communication System) and the endoscope monitor, enabling real-time storage and clinical application of diagnostic and treatment data. Lesion attributes include center coordinates, area, and malignancy probability, while treatment decision plans include surgical methods or drug intervention plans.

[0076] Specifically, after the system completes the calculation of lesion attributes and type matching, the diagnosis and treatment logic for different diseases needs to be combined with the subdivision rules in the clinical guideline knowledge base. In this embodiment, nasal polyps are used as an example to illustrate the specific decision-making process. The system completes the nasal polyp type matching through the lesion feature vector output by the segmentation network. This feature vector includes key features such as the morphology of the lesion's mucosal elevation, surface smoothness, and absence of abnormal blood vessel density. The system compares these features with the nasal polyp feature template in the clinical guideline knowledge base. When the feature matching degree reaches a set threshold, the lesion type is determined to be nasal polyp. Subsequently, the corresponding treatment decision plan is generated by combining the lesion area and the probability of malignancy (the probability of malignancy of nasal polyps is usually less than 0.3, which falls into the category of benign lesions).

[0077] If the calculated nasal polyp lesion area is small and no obstruction of the surrounding nasal passages is detected, the system will retrieve the treatment plan corresponding to "small benign nasal polyps" from the knowledge base. This plan includes specific implementation guidelines for drug intervention, clearly specifying the use of nasal corticosteroids as the core treatment drug. The system will indicate the drug administration frequency in the output plan, i.e., a fixed number of nasal spray administrations per day, and also indicate the duration of the treatment course, typically a follow-up evaluation after a set period of continuous use. The criteria for the follow-up evaluation are preset by the system in conjunction with the knowledge base, mainly including the rate of change in lesion volume and the degree of improvement in patient symptoms. If the system detects no significant reduction in lesion volume or no improvement in symptoms during the follow-up, it will automatically trigger the plan adjustment process and switch to the next treatment plan.

[0078] If the calculated nasal polyp lesion area is large, or if the lesion is found to be obstructing the nasal passages and affecting nasal ventilation, the system will retrieve the treatment plan corresponding to "medium to large benign nasal polyps" from the knowledge base. This plan focuses on minimally invasive resection surgery. The system will specify the surgical method as endoscopic polypectomy in the output plan, and highlight key technical points of the procedure, including preserving normal nasal mucosa to maintain nasal physiological function and avoiding damage to the sinus openings. The surgical plan also includes postoperative care guidelines, specifying the frequency and type of nasal irrigation fluid required postoperatively, as well as the consolidation treatment period using nasal corticosteroids. Furthermore, the system will adjust the plan based on whether the lesion is complicated by complications such as sinusitis. If sinusitis is detected, the system will supplement the surgical plan with the course and dosage of postoperative antibiotics to ensure the completeness and specificity of the treatment plan.

[0079] When outputting a treatment decision plan for nasal polyps, the system simultaneously links it to a follow-up management plan, specifying the timeframe for the first postoperative follow-up and the intervals for subsequent regular follow-ups. During the follow-up process, the system needs to re-acquire endoscopic images to assess whether the lesions have recurred. If signs of recurrence are detected, a new treatment plan will be generated based on the new lesion attributes, forming a closed-loop management system for the diagnosis and treatment of nasal polyps.

[0080] In summary, the lesion segmentation method based on ENT endoscopy images provided in this application achieves strict frame alignment between white light and narrowband imaging through a time-triggered dual-spectrum synchronous acquisition mechanism, thereby eliminating ghosting interference caused by organ peristalsis. Dynamic mucus displacement modeling based on pixel gradient fields accurately captures the flow vector of mucus, effectively distinguishing the secretion-covered area from the actual anatomical structure. Displacement amplitude threshold segmentation and texture reconstruction strategies restore the continuity of mucosal glands and vascular networks obscured by mucus, improving the visualization integrity of the lesion area. Adaptive learning of deformation parameters of deformable convolutional layers enables sub-pixel spatial registration of white light and narrowband features, overcoming multimodal feature misalignment caused by differences in spectral scattering characteristics. Finally, through topological attribute analysis of the lesion probability map and mapping with clinical guideline rules, anatomical positioning navigation markers and graded treatment decisions superimposed on real-time video are generated synchronously, enabling precise intraoperative lesion localization and immediate output of evidence-based medical plans.

[0081] The method provided in this application innovatively integrates three major functions: dynamic mucus suppression, cross-modal feature fusion, and closed-loop diagnosis and treatment decision-making. It can significantly reduce the risk of false negatives caused by mucus interference, solve the problem of misdiagnosis of vascular lesions caused by registration misalignment in traditional methods, and avoid decision delays caused by the separation of diagnosis and treatment processes, thereby improving the accuracy, efficiency, and clinical operability of ENT endoscopy.

[0082] In one embodiment, S1 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0083] S11: Perform time-division triggering processing on the control signal of the endoscope light source, and generate time-division imaging control commands without spectral crosstalk by alternately switching the phase of the driving current of the white light source and the narrowband filter.

[0084] Specifically, the system performs time-division triggering of the endoscope light source control signal to generate time-division imaging control commands without spectral crosstalk. This is achieved by alternately switching the driving current phases of the white light source and the narrowband filter. The system first acquires the initial control signal of the endoscope light source, which contains the basic parameters of the white light source drive and the narrowband filter drive. The system inputs the initial control signal to the light source control unit, which then splits and re-plans the driving current phase.

[0085] The phase switching of the driving currents for the white light source and the narrowband filter must adhere to the principle of "non-overlapping operation." Through a phase adjustment module within the light source control unit, the effective operating periods of the two driving currents are completely staggered, ensuring that only one driving current is in the conducting state at any given time. This avoids simultaneous output of both spectra due to current superposition, thereby eliminating spectral crosstalk. During phase switching, the system monitors the phase states of the two driving currents in real time and corrects phase shifts through a feedback adjustment mechanism, ensuring the accuracy and stability of the switching.

[0086] After phase planning is completed, the system generates time-division imaging control commands from the light source control unit. These commands include key parameters such as the conduction duration of the white light source driving current, the conduction duration of the narrowband filter driving current, and the switching interval between the two currents. The system verifies the generated control commands, confirming that there are no phase overlap regions and that the parameter settings meet the light source's operating requirements, before transmitting the commands to the subsequent endoscopic imaging system.

[0087] S12: Input the time-division imaging control command into the image sensor driving circuit of the endoscope imaging system, perform frame synchronization processing on the raw photoelectric conversion data captured by the sensor, and generate a bispectral image with exposure time alignment by compensating for the start time offset between white light and narrowband exposure.

[0088] Specifically, the system inputs time-division imaging control commands into the image sensor driving circuit of the endoscopic imaging system. This circuit controls the image sensor, thereby completing frame synchronization processing of the raw photoelectric conversion data and generating a dual-spectral image with exposure time alignment. After receiving the time-division imaging control commands, the image sensor driving circuit first parses the exposure control parameters in the commands, including the white light exposure start time, white light exposure duration, narrowband exposure start time, and narrowband exposure duration. Based on these parameters, it generates corresponding sensor driving signals.

[0089] The image sensor performs photoelectric conversion under the action of a driving signal, converting the received light signal into an electrical signal to form raw photoelectric conversion data. Due to the physical response difference between the driving of the white light source and the narrowband filter, the start time of white light exposure and narrowband exposure may be offset. The system compensates for this offset through a frame synchronization processing module. The frame synchronization processing module extracts the frame timestamp information from the raw photoelectric conversion data, compares the exposure start timestamps of the white light frame and the narrowband frame, calculates the time offset, and then adjusts the start time of subsequent exposures according to the offset to keep the start time of white light exposure and narrowband exposure consistent.

[0090] After frame synchronization, the system performs preprocessing such as signal amplification and noise reduction on the corrected photoelectric conversion raw data. Then, it converts the analog signal into digital image data through analog-to-digital conversion, generating white light images and narrowband imaging images respectively. The system performs time-series alignment verification on the generated dual-spectral images. By comparing the frame numbers and timestamps of the two images, it confirms that the exposure timing of each set of white light images and narrowband imaging images is completely matched, ensuring that the images correspond to the same cavity scene, and finally forming a set of dual-spectral images with aligned exposure timing.

[0091] S13: Motion artifact detection processing is performed on the dual-spectral images. Through cross-correlation analysis of adjacent frames in the white light channel and the narrowband channel, distorted frames caused by organ tremors are removed, and time-aligned white light image sequences and narrowband imaging images are generated.

[0092] Specifically, the system performs motion artifact detection processing on the time-aligned dual-spectral images. Through cross-correlation analysis of adjacent frames in the white light channel and narrowband channel, it removes distorted frames caused by organ tremors, ultimately generating a time-aligned white light image sequence and narrowband imaging images. The system first separates the dual-spectral images into a white light channel image set and a narrowband channel image set, then sorts the images in both channels according to their acquisition time order to form a preliminary image sequence.

[0093] For the initial image sequence of each channel, the system selects two adjacent frames for cross-correlation analysis to calculate the cross-correlation coefficient between them. The cross-correlation coefficient characterizes the similarity between two frames. When organ tremors occur, adjacent frames will show significant differences due to scene changes, and the cross-correlation coefficient will decrease significantly. The system presets a cross-correlation coefficient threshold based on clinical diagnostic needs, and frames with a cross-correlation coefficient below the threshold are identified as distorted frames. These distorted frames are images with motion artifacts caused by organ tremors.

[0094] After detecting distorted frames in a single channel, the system further verifies the frame correspondence between the white light channel and the narrowband channel. Since the images from both channels need to maintain temporal alignment, if a frame in one channel is identified as a distorted frame, the system must simultaneously check whether the corresponding timestamp frame in the other channel is also a distorted frame. This ensures that after removing distorted frames, the remaining white light frames and narrowband frames still correspond one-to-one. The system then reorders the white light channel image set and the narrowband channel image set after removing distorted frames, adding frame numbers to form a continuous image sequence.

[0095] In one embodiment, S2 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0096] S21: Perform viscous rheological property analysis on white light image sequences. Calculate the constitutive equation parameters of viscous shear stress and strain rate by using the non-rigid transformation relationship of pixel gradient fields between adjacent frames, and generate viscous viscoelastic parameters.

[0097] Specifically, the system performs viscous rheological property analysis on the white light image sequence, preprocesses the white light image sequence by converting each frame of color image into a single-channel grayscale image to eliminate the interference of color channel differences on pixel value changes, and then calculates the pixel gradient field of each frame of grayscale image using the Sobel operator. The x-direction gradient and y-direction gradient of the Sobel operator are calculated according to the following formulas:

[0098]

[0099]

[0100] in, It is a grayscale image. This represents the rate of change of grayscale value of a pixel in the x-direction. The gradient field represents the rate of grayscale change of a pixel in the y-direction, and as a whole, it reflects the boundary between the mucus and the anatomical structure and the difference in grayscale distribution.

[0101] Furthermore, the system selects two adjacent image frames, extracts feature points of the mucus region using a feature point matching algorithm, and then constructs a non-rigid transformation matrix between the feature points using thin-plate spline interpolation. Based on this matrix, the system calculates the shear stress and strain rate of the mucus, with the shear stress... and strain rate Calculate using the formula:

[0102]

[0103]

[0104] in, The pixel density of the slime region. This represents the change in gradient field between adjacent frames; For pixels Directional displacement, For pixels Directional displacement, It reflects the deformation rate of the viscous fluid flow. The system will and Substituting into the constitutive equation of mucus:

[0105]

[0106] in, For elastic modulus, In response, The viscosity coefficient is obtained by fitting a curve using the least squares method. and Generate viscoelastic parameters for the mucus.

[0107] S22: Model the displacement field of the viscoelastic parameters of the viscous fluid, solve the displacement vector equation through viscous fluid dynamic constraints, substitute the viscoelastic parameters into the preset Oldroyd-B model to calculate the stress tensor distribution of the viscous fluid flow, and generate a dynamic viscous fluid displacement field.

[0108] Specifically, the system models the displacement field based on the viscoelastic parameters of the viscous fluid, constructing the displacement vector equations for the viscous flow. The system establishes a mass conservation constraint according to the principles of continuum mechanics, expressed by the following formula:

[0109]

[0110] in, For divergence operators, For the velocity field of the mucus, The velocity component is in the x-direction. Let be the velocity component in the y-direction. This formula ensures that the volume remains constant during the flow of the viscous fluid. Simultaneously, the system establishes a momentum conservation constraint, including the elastic modulus. With viscosity coefficient Substitute these forces into the inertial force, viscous force, and elastic force of the mucus.

[0111] Preferably, the system can discretize the displacement vector equations using the finite difference method, transforming the continuous displacement field into a set of displacement equations for discrete pixels. The system then calls the preset Oldroyd-B model to calculate the viscosity stress tensor distribution according to the formula:

[0112]

[0113] in, For stress tensor, For time, and For relaxation time, The strain rate tensor is used. The system modifies the displacement equations based on the stress tensor distribution to ensure that the displacement field and stress field satisfy the mechanical equilibrium condition. The dynamic displacement value of each pixel is obtained through iterative solution and arranged in spatial coordinates to generate a dynamic viscous displacement field.

[0114] S23: Noise suppression is performed on the dynamic mucus displacement field. The anatomical structure motion and mucus flow components are separated by low-pass filtering. Abnormal pulsation signals with frequencies higher than the physiological motion threshold are filtered out to generate a denoised mucus motion vector field.

[0115] Specifically, the low-frequency signal is mixed with anatomical peristalsis and electronic noise. The system selects a low-pass filter algorithm to process the dynamic mucus displacement field, and the frequency response of the filter algorithm is set according to the formula:

[0116]

[0117] in, The signal angular frequency, The cutoff frequency, Based on the normal movement frequency range of the anatomical structure of the ear, nose, and throat cavities, the low-frequency component of mucus flow is preserved, while the filtering frequency is higher than [the specified frequency]. Abnormal pulsating signals.

[0118] Preferably, the system performs low-pass filtering on the x- and y-direction displacement components of the dynamic mucus displacement field, attenuates high-frequency signal components through convolution operations, and retains low-frequency mucus flow components and anatomical structure peristalsis components. The system extracts the periodic features of the signal through time-domain analysis, identifying signal components with continuous periodic changes as anatomical structure motion components and signal components without a fixed period as mucus flow components, thus completing the separation of the two types of components. The system calculates the deviation of adjacent pixel displacement values ​​within the mucus flow component, confirms the absence of abnormal pulsation signal residue, and then organizes the component according to spatial coordinates and time series to generate a denoised mucus motion vector field.

[0119] In one embodiment, S3 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0120] S31: Perform amplitude adaptive segmentation on the mucus motion vector field, and generate a dynamic threshold mask boundary associated with the mucus adhesion strength by calculating the spatiotemporal accumulation of motion vectors in the local neighborhood.

[0121] Specifically, the system performs amplitude adaptive segmentation on the slime motion vector field. The core of this process involves calculating the spatiotemporal accumulation of motion vectors within a local neighborhood to generate a dynamic threshold mask boundary associated with the slime adhesion strength. The system first acquires the denoised slime motion vector field, which contains the displacement components of each pixel in the x and y directions. For each pixel in the vector field, the system delineates a local neighborhood centered on that pixel. The neighborhood size is determined based on the image resolution and slime distribution density, ensuring that the neighborhood contains a sufficient number of motion vector samples while avoiding cross-regional interference.

[0122] Furthermore, the system calculates the spatiotemporal accumulation of motion vectors within the local neighborhood. This spatiotemporal accumulation is used to quantify the overall activity level of motion vectors within the neighborhood, and its calculation formula is as follows:

[0123]

[0124] in, For pixels The spatiotemporal accumulation of the local neighborhood. For pixels The local neighborhood range, The starting frame time, The time window length, and They are respectively The displacement components of pixel (x, y) in the x and y directions at time step (x, y). The system uses this formula to iterate through all pixels and obtain the spatiotemporal cumulative distribution of the entire image.

[0125] Preferably, the system generates a dynamic threshold based on the spatiotemporal accumulation distribution. The determination of the dynamic threshold is related to the slime adhesion strength; regions with higher adhesion strength typically have larger spatiotemporal accumulations. The system employs an adaptive thresholding algorithm, using the mean and standard deviation of the spatiotemporal accumulation of all pixels in the image as a basis to calculate the dynamic threshold for each local region, ensuring that the threshold can distinguish between slime and non-slime regions. The system compares the spatiotemporal accumulation of each pixel with the dynamic threshold of the corresponding local region. Pixels with accumulations higher than the threshold are marked as slime regions, and pixels with accumulations lower than the threshold are marked as non-slime regions, thereby generating a dynamic threshold mask boundary. This boundary adaptively adjusts with changes in slime adhesion strength, accurately delineating the slime coverage area.

[0126] S32: Reconstruct the anatomical structure of the area covered by the dynamic threshold mask boundary, restore the continuity of the mucosal glands in the occluded area through the neighborhood healthy tissue texture propagation algorithm, and generate a preliminary reconstructed white light image.

[0127] Specifically, the system reconstructs the anatomical structure of the area covered by the dynamic threshold mask boundary. The core of this process involves restoring the continuity of the mucosal glands in the occluded area using a neighborhood healthy tissue texture propagation algorithm, generating a preliminary reconstructed white-light image. Based on the dynamic threshold mask boundary, the system locates the mucus-occluded area in the image, i.e., the set of pixels enclosed by the mask boundary. Simultaneously, it identifies the surrounding neighborhood healthy tissue, which must meet the conditions of being free of mucus coverage and having complete texture features, providing original texture samples for texture propagation.

[0128] Furthermore, the system can employ a neighborhood healthy tissue texture propagation algorithm for reconstruction, gradually propagating the texture features of healthy tissue to the occluded area. The system defines the pixel update formula for texture propagation as follows:

[0129]

[0130] in, pixels in the occluded area The reconstructed pixel values, For pixels The corresponding set of healthy tissue pixels in the neighborhood, For healthy tissue pixels For reconstructed pixels The weighting coefficients, For healthy tissue pixels The original pixel values. Weighting coefficients. The weighting coefficient is determined based on the spatial distance and texture similarity between healthy tissue pixels and reconstructed pixels. The closer the distance and the higher the texture similarity, the larger the weighting coefficient, ensuring the consistency between the reconstructed texture and the healthy tissue texture.

[0131] During texture propagation, the system focuses on maintaining the continuity of the mucosal gland orientation. Preferably, the system can extract the edge contours of mucosal glands in healthy tissue using an edge detection algorithm, determine the direction vector of the gland orientation, and adjust the pixel value distribution according to this direction vector when the texture propagates to the occluded area. This ensures a smooth transition between the gland edges in the reconstructed area and the healthy tissue, avoiding any breaks in the orientation. After traversing all pixels in the occluded area and completing pixel value reconstruction, the system integrates the pixel information from the reconstructed area and the non-occluded area to generate a preliminary white light reconstructed image. This image has restored most of the anatomical structures obscured by mucus, but may still exhibit inconsistencies in texture transition areas.

[0132] S33: Optimize the edge consistency of the white light reconstructed image, eliminate step artifacts in the texture transition area through multi-scale gradient fusion, and generate an optimized white light image with continuous anatomical structure.

[0133] Specifically, the system optimizes the edge consistency of the white light reconstructed image by eliminating step artifacts in texture transition areas through multi-scale gradient fusion, generating an optimized white light image with continuous anatomical structures. The system acquires a preliminary white light reconstructed image and performs multi-scale gradient calculations on it. These multi-scale gradients are used to capture edge features at different resolutions. The system employs a Gaussian pyramid decomposition algorithm to decompose the preliminary reconstructed image into multiple image layers at different scales. Each scale corresponds to a different Gaussian filter standard deviation. A larger standard deviation indicates a lower image layer resolution and captures more macroscopic edge features; a smaller standard deviation indicates a higher image layer resolution and captures more refined edge features.

[0134] Furthermore, the system calculates gradient maps for each image layer at each scale. These gradient maps are then calculated using the Sobel operator to obtain the x-axis and y-axis gradients of the image at each scale. Preferably, the system can employ a multi-scale gradient fusion algorithm to fuse gradient maps from different scales into a unified gradient map. The fusion formula is as follows:

[0135]

[0136] in, The gradient value after fusion. The number of scale layers, The fusion weights of the gradient map at the k-th scale satisfy the following conditions: , and The first Gradient values ​​of each scale image layer in the x and y directions. Fusion weights. The weight of the gradient map is determined based on the edge sharpness of each scale. The higher the edge sharpness, the greater the weight, ensuring that the fused gradient map can both preserve the continuity of macro edges and reflect the details of fine edges.

[0137] Preferably, the system performs edge correction on the preliminary white light reconstructed image based on the fused gradient map, eliminating step artifacts in texture transition areas through gradient-guided pixel value adjustment. The system traverses all pixels in texture transition areas of the image, fine-tuning pixel values ​​according to the gradient direction and magnitude of the fused gradient map, ensuring that pixel value changes in transition areas conform to the gradient trend and avoiding sudden pixel value jumps. After edge correction, the system performs overall image smoothing to further eliminate local minor artifacts, ultimately generating an optimized white light image with continuous anatomical structures. In this image, the direction of mucosal glands is consistent and there is no mucus obstruction interference.

[0138] In one embodiment, S4 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0139] S41: Perform multi-resolution feature extraction on the optimized white light image, and generate multi-scale feature response maps covering macroscopic anatomical structures to microscopic textures through convolutional pyramid decomposition, thereby generating cross-scale white light feature maps.

[0140] Specifically, after acquiring continuous optimized white-light images of the anatomical structures, the system initiates a multi-resolution feature extraction process. The system first constructs a convolutional pyramid, using the optimized white-light image as the bottom layer input. It generates images for each layer of the pyramid by performing convolution operations and downsampling processing layer by layer. The bottom layer image retains the microscopic texture information of the original image, such as details of gland openings on the mucosal surface. As the pyramid level increases, the convolution operations gradually extract more abstract features, while the downsampling processing expands the feature receptive field to cover macroscopic anatomical structures, such as cavity boundaries and mucosal fold morphology. Each convolution operation uses a fixed-size convolution kernel and calculates the feature response values ​​of the pixel neighborhood through a sliding window to generate a feature response map for the corresponding layer. The downsampling processing selects the maximum or average value of the feature responses within the neighborhood, reducing the image resolution while retaining key feature information. The system integrates the feature response maps of each layer of the convolutional pyramid along the channel dimension, upsamples the feature response maps of different layers to the same resolution, and then stitches the features together to form a multi-scale feature set covering macroscopic anatomical structures to microscopic textures. Finally, it generates a cross-scale white light feature map that contains both global structural information and retains local detail features.

[0141] S42: Enhance the frequency domain features of the narrowband imaging image by using bandpass filtering to enhance the frequency domain response of the vascular morphology and mucosal gland openings corresponding to the 415nm and 540nm bands, and generate an enhanced narrowband feature map.

[0142] Specifically, after acquiring time-aligned narrowband imaging images, the system initiates a frequency domain feature enhancement process targeting the vascular morphology and mucosal surface glandular opening features corresponding to the 415nm and 540nm bands. The system performs frequency domain transformation on the narrowband imaging images, converting the spatial domain image into a frequency domain image. It decomposes the image into frequency components through mathematical transformations, where different frequency components correspond to different structural features of the image: low-frequency components correspond to the overall grayscale distribution of the image, mid-frequency components correspond to target structures such as vascular morphology and glandular openings, and high-frequency components correspond to noise or minute textures. Based on the optical characteristics of the 415nm band (emphasizing mucosal vascular endothelial features) and the 540nm band (emphasizing mucosal capillary network), the system designs a dedicated bandpass filter function. This function only allows mid-frequency components matching vascular morphology and glandular opening features to pass through, filtering out low-frequency redundant information and high-frequency noise. The system applies the bandpass filter function to the frequency domain image, enhancing the frequency domain response value of the target structure, and then performs an inverse frequency domain transformation to restore the processed frequency domain image to a spatial domain image. The system extracts features from the inverse-converted spatial image, extracts enhanced blood vessel morphology and gland opening features through convolution operations, and generates an enhanced narrowband feature map. The contrast of the target structure in this feature map is significantly improved, providing a clear narrowband feature benchmark for subsequent spatial offset correction.

[0143] S43: Perform spatial offset correction on the white light feature map and the enhanced narrowband feature map, learn the feature deformation parameters between the white light and narrowband features through deformable convolution kernels, and generate a registered multimodal feature map.

[0144] Specifically, after acquiring cross-scale white light feature maps and enhanced narrowband feature maps, the system learns feature deformation parameters through deformable convolutional kernels and initiates the spatial offset correction process. The system first defines the white light feature map as... Enhanced narrowband feature map ,by As a benchmark Offset correction is performed. The system first calculates the learnable offset. Using a 3×3 convolution kernel and The formula for joint convolution is:

[0145]

[0146] in, This represents a 3×3 convolution operation. As a learnable offset, reflecting Compared to The spatial offset trend at each pixel location. The system then bases this on... Calculate characteristic deformation parameters , with target location coordinates Construct a 3×3 sampling grid around the center. Each sampling point within the grid is calculated using bilinear interpolation. weight , and then combine The formula for the positional features after offset is:

[0147]

[0148] in, The characteristic deformation parameters at the target position p. The target location coordinates, The sampling grid is 3×3. The weights are determined by bilinear interpolation; the closer the sampling point is to the target location, the greater the weight. It is characterized by narrow bands. For learnable offsets, Characterized by white light; for Zhongyu Target location The corresponding corrected position. The system optimizes through iterative learning. and ,make Feature distribution and To maintain spatial consistency, the final corrected and Feature stitching is performed to generate a registered multimodal feature map, which integrates the core features of white light and narrowband without spatial offset error.

[0149] In one embodiment, step S5 of the lesion segmentation method based on ENT endoscopic images provided by the present invention specifically includes the following steps:

[0150] S51: Perform feature channel splicing on the multimodal feature map. Generate a fused feature tensor by splicing white light and narrowband feature map along the channel axis. Input the fused feature tensor into the U-shaped decoding network for upsampling and feature fusion to generate a lesion probability map.

[0151] Specifically, the system performs feature channel stitching on the multimodal feature maps. The core process involves integrating white light and narrowband feature information along the channel axis to generate a fused feature tensor, which is then processed by a U-shaped decoding network to generate a lesion probability map. The system first acquires the registered multimodal feature map, which includes a white light feature map and an enhanced narrowband feature map. Both types of feature maps have undergone spatial offset correction, and the anatomical positions of corresponding pixels are consistent. The system determines the axis direction for feature channel stitching as the channel dimension, maintaining the height and width dimensions of the feature maps unchanged, and only superimposing the number of channels from both types of feature maps along the channel dimension to form the fused feature tensor. For example, if the white light feature map contains C1 channels and the enhanced narrowband feature map contains C2 channels, the number of channels in the fused feature tensor is C1 + C2, and the height and width are the same as the original feature map. Figure 1 This ensures that the semantic information of both types of features is fully preserved in the concatenated feature tensor.

[0152] Furthermore, the system integrates feature tensors into a U-shaped decoding network, which consists of an encoder and a decoder. The encoder has already performed feature compression in the earlier feature extraction stage, while the decoder uses transposed convolutional layers to upsample, gradually restoring the feature map resolution to match the original endoscopic image. During upsampling, the system uses skip connections to fuse the feature maps from each level of the decoder with the corresponding level of the encoder, supplementing the lost details, especially the microscopic texture information of the lesion edges. The network output layer uses the sigmoid activation function to map feature values ​​to the 0-1 range, generating a lesion probability map. The value of each pixel in the probability map represents the probability that the location belongs to a lesion region; a higher value indicates a greater likelihood of a lesion.

[0153] S52: Perform connected component attribute parsing on the lesion probability map. By scanning connected regions in the probability map that are greater than a set threshold, calculate the centroid spatial coordinates and pixel area of ​​each connected region to generate quantitative attributes of lesion location and size.

[0154] Specifically, the system performs connected component attribute parsing on the lesion probability map. The core process involves scanning connected regions in the probability map that meet threshold conditions, calculating the centroid coordinates and pixel area of ​​each region, and generating quantitative attributes for lesion location and size. The system first calls a preset probability threshold, set based on the sensitivity and specificity requirements for lesion identification in clinical diagnosis, to distinguish lesion areas from background areas. The system traverses the lesion probability map using a combination of row and column scans, marking pixels with values ​​greater than the preset threshold as candidate lesion pixels. Then, it uses the 8-neighborhood connectivity criterion to determine whether candidate lesion pixels constitute a connected region. That is, if two candidate lesion pixels are adjacent in the horizontal, vertical, or diagonal direction, they are classified as belonging to the same connected region, avoiding misclassification of independent noise points as lesions.

[0155] Furthermore, the system calculates the quantization attributes for each confirmed connected region, calculates the centroid spatial coordinates, and uses a pixel coordinate weighted average formula:

[0156]

[0157] in, Let R be the centroid coordinates of the connected region R. For pixels The system calculates the probability value of lesions and uses a weighted average of these probability values ​​to ensure the centroid is closer to the core area of ​​the lesion. Then, it counts the number of pixels of all candidate lesions within the connected regions to obtain the pixel area. This pixel area is then converted into the actual anatomical area by combining the pixel size calibration parameters of the endoscopic image, thus quantifying the lesion size. The system filters the calculated quantified attributes, removing connected regions with pixel areas smaller than a set minimum threshold, as these regions are often noise interference. Finally, it retains the valid lesion localization and size quantification attributes.

[0158] S53: Real-time navigation marker generation based on quantized attributes. The centroid coordinates of the lesion are mapped to the endoscopic video coordinate system through three-dimensional projection transformation to generate real-time navigation markers for the endoscopic video.

[0159] Specifically, the system generates real-time navigation markers based on quantified attributes. The core mechanism involves mapping the centroid coordinates of the lesion to the endoscopic video coordinate system via 3D projection transformation, generating real-time navigation markers that can be superimposed on the video. The system first obtains the centroid spatial coordinates of the lesion, which belong to the image coordinate system. The origin is the upper left corner of the image, with the x-axis pointing horizontally to the right and the y-axis pointing vertically downwards. Simultaneously, it calls upon the endoscope system's intrinsic and extrinsic parameters. The intrinsic parameters include parameters such as camera focal length and pixel pitch, while the extrinsic parameters include the position and orientation parameters of the endoscope lens relative to the patient's anatomical structures. These parameters are obtained through preoperative calibration and stored in the system.

[0160] Furthermore, the system constructs a three-dimensional projection transformation model, transforming the centroid coordinates in the image coordinate system... Convert to coordinates in the endoscopic video coordinate system The video coordinate system uses the top-left corner of the monitor screen as its origin, with the x-axis pointing horizontally to the right and the y-axis pointing vertically downwards. The transformation process must consider the resolution matching between the image coordinate system and the video coordinate system, as well as the correction for endoscope lens distortion. The system completes the coordinate mapping using a projection transformation formula, ensuring that the mapped centroid coordinates accurately correspond to the actual location of the lesion in the video frame. Subsequently, the system generates navigation markers, which include a bounding box surrounding the lesion area and text annotating key lesion information. The size of the bounding box is determined based on the pixel area of ​​the lesion; the larger the area, the larger the bounding box. The text content includes quantitative attributes such as the lesion's centroid coordinates and actual area.

[0161] The system overlays the generated navigation markers onto the endoscopic video stream in real time. Before each video frame is updated, the system recalculates the quantitative attributes and projection coordinates of the lesion in the current frame to ensure that the navigation markers dynamically adjust with the video frames and always accurately cover the lesion area. At the same time, the system sets the transparency of the navigation markers to a level that does not affect the doctor's view of the original video image, balancing marker clarity and image readability. Finally, it generates an endoscopic video with real-time navigation markers, providing lesion localization guidance for intraoperative procedures.

[0162] S54: Based on quantitative attributes, clinical decision matching is performed. By querying the pre-set guideline knowledge base for treatment rules that match the location and area of ​​the lesion, a treatment decision plan matching the clinical guidelines is generated.

[0163] Specifically, the system performs clinical decision matching based on quantitative attributes. Its core mechanism involves querying treatment rules in a pre-defined guideline knowledge base to generate treatment plans that conform to clinical guidelines. The system first constructs a guideline knowledge base, which is stored in a structured format and includes lesion location classifications, area range divisions, and corresponding treatment rules. For example, lesions are categorized by location (nasal cavity, ear canal, pharynx, etc.) and by area (small, medium, large, etc.). Each location-area combination corresponds to a unique treatment rule, the content of which is developed with reference to the latest clinical practice guidelines and updated regularly.

[0164] Furthermore, the system extracts location and area information from the quantitative attributes of the lesion. Location information is determined by combining centroid coordinates with an anatomical zoning map pre-endoscopically defined. For example, if the centroid coordinates fall within the middle nasal meatus region, the lesion is determined to be located in the middle nasal meatus. Area information is obtained by matching the actual anatomical area with an area range in the knowledge base. The system employs a rule-matching algorithm, first filtering the corresponding location category subset in the knowledge base based on the lesion location, and then matching the corresponding treatment rule within the subset based on the area range. For example, small lesions in the middle nasal meatus are matched with drug intervention rules, medium-sized lesions with minimally invasive resection rules, and large lesions with open surgery rules.

[0165] The system generates treatment decision plans based on the matched treatment rules. These plans include treatment methods, specific operational guidelines, and postoperative follow-up requirements. The system outputs the plans in text format, simultaneously displaying them in the sidebar of the endoscope monitor and storing them in the patient's medical database for doctors to view in real time and for subsequent medical record organization. If multiple lesions or quantitative attributes are at the boundary of an interval, the system outputs priority-ranked treatment plans, noting the applicable basis for each plan to assist doctors in making the final decision.

[0166] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0167] Based on the same inventive concept, this application also provides a lesion segmentation system based on otolaryngological endoscope images for implementing the aforementioned lesion segmentation method based on otolaryngological endoscope images. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the lesion segmentation system based on otolaryngological endoscope images provided below can be found in the limitations of the lesion segmentation method based on otolaryngological endoscope images described above, and will not be repeated here.

[0168] Preferably, such as Figure 3 As shown, the present invention provides a lesion segmentation system 600 based on otolaryngological endoscopy images, which is configured with the following modules:

[0169] The dual-spectrum image synchronous acquisition module 610 is used to synchronously acquire the timing trigger signal of the endoscope light source. By alternately driving the white light source and the narrowband filter device, it captures timing-aligned dual-spectrum images and generates a timing-aligned white light image sequence and a narrowband imaging image.

[0170] The slime displacement modeling module 620 is used to model slime displacement in white light image sequences. It calculates the displacement vector of dynamic slime by changing the pixel gradient between adjacent frames and generates a slime motion vector field.

[0171] The slime mask repair and image optimization module 630 is used to perform mask repair on the slime motion vector field and white light image sequence. It segments the slime-covered area by displacement amplitude threshold and reconstructs the occluded anatomical structure pixels to generate an optimized white light image with slime interference removed.

[0172] The cross-modal feature registration module 640 is used to perform cross-modal feature registration on optimized white light images and narrowband imaging images. It extracts white light feature maps and narrowband feature maps respectively, and corrects the spatial offset between feature maps through deformable convolutional layers to generate registered multimodal feature maps.

[0173] The lesion segmentation and treatment plan generation module 650 is used to segment lesions from multimodal feature maps, generate lesion probability maps through feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopic video and treatment decision plans matching clinical guidelines based on the spatial coordinates and area attributes of the connected domains in the lesion probability map.

[0174] Among them, real-time navigation markers are used to indicate the anatomical location of lesions and the probability of malignancy, while treatment decision plans are used to indicate surgical or drug intervention plans that match the type and size of the lesions.

[0175] Preferably, the dual-spectral image synchronous acquisition module 610 provided in this application is configured with the following units:

[0176] The light source time-division triggering unit is used to perform time-division triggering processing on the control signal of the endoscope light source. By alternately switching the phase of the driving current of the white light source and the narrowband filter, a time-division imaging control command without spectral crosstalk is generated.

[0177] The frame synchronization processing unit is used to input the time-division imaging control command into the image sensor driving circuit of the endoscope imaging system, perform frame synchronization processing on the raw photoelectric conversion data captured by the sensor, and generate a bispectral image with exposure time alignment by compensating for the start time offset between white light and narrowband exposure.

[0178] The artifact detection and removal unit is used to perform motion artifact detection processing on dual-spectral images. By analyzing the cross-correlation between adjacent frames of the white light channel and the narrowband channel, it removes distorted frames caused by organ tremors and generates a time-aligned white light image sequence and a narrowband imaging image.

[0179] Preferably, the viscous displacement modeling module 620 provided in this application is configured with the following units:

[0180] The mucus rheology analysis unit is used to analyze the mucus rheological properties of white light image sequences. By using the non-rigid transformation relationship of pixel gradient fields between adjacent frames, it calculates the constitutive equation parameters of mucus shear stress and strain rate, and generates mucus viscoelastic parameters.

[0181] The displacement field modeling unit is used to model the displacement field of the viscoelastic parameters of the viscous fluid. It solves the displacement vector equation through viscous fluid dynamics constraints, substitutes the viscoelastic parameters into the preset Oldroyd-B model to calculate the stress tensor distribution of the viscous fluid flow, and generates a dynamic viscous fluid displacement field.

[0182] The displacement field noise reduction unit is used to suppress noise in the dynamic mucus displacement field. It separates the anatomical structure motion and mucus flow components through low-pass filtering, filters out abnormal pulsation signals with frequencies higher than the physiological motion threshold, and generates a denoised mucus motion vector field.

[0183] Preferably, the mucus mask repair and image optimization module 630 provided in this application is configured with the following units:

[0184] The amplitude adaptive segmentation unit is used to perform amplitude adaptive segmentation of the mucus motion vector field. By calculating the spatiotemporal accumulation of motion vectors in the local neighborhood, a dynamic threshold mask boundary associated with the mucus adhesion strength is generated.

[0185] The anatomical structure reconstruction unit is used to reconstruct the anatomical structure of the area covered by the dynamic threshold mask boundary. It restores the continuity of the mucosal gland direction in the occluded area through the neighborhood healthy tissue texture propagation algorithm and generates a preliminary reconstructed white light image.

[0186] The edge consistency optimization unit is used to optimize the edge consistency of the white light reconstructed image. It eliminates step artifacts in the texture transition area through multi-scale gradient fusion and generates an optimized white light image with continuous anatomical structure.

[0187] Preferably, the cross-modal feature registration module 640 provided in this application is configured with the following units:

[0188] The white light feature extraction unit is used to extract multi-resolution features from optimized white light images. It generates multi-scale feature response maps covering macroscopic anatomical structures to microscopic textures through convolutional pyramid decomposition, thus generating cross-scale white light feature maps.

[0189] The narrowband feature enhancement unit is used to enhance the frequency domain features of narrowband imaging images. It enhances the frequency domain response of blood vessel morphology and mucosal surface gland openings corresponding to the 415nm and 540nm bands through bandpass filtering, and generates enhanced narrowband feature maps.

[0190] The feature offset correction unit is used to perform spatial offset correction on the white light feature map and the enhanced narrowband feature map. It learns the feature deformation parameters between the white light and narrowband features through deformable convolution kernels to generate a registered multimodal feature map.

[0191] Preferably, the lesion segmentation and treatment plan generation module 650 provided in this application is configured with the following units:

[0192] The feature splicing and decoding unit is used to perform feature channel splicing processing on the multimodal feature map. It generates a fused feature tensor by splicing white light and narrowband feature map along the channel axis, and inputs the fused feature tensor into the U-shaped decoding network for upsampling and feature fusion to generate a lesion probability map.

[0193] The connected component parsing unit is used to perform connected component attribute parsing on the lesion probability map. By scanning the connected regions in the probability map that are greater than a set threshold, the centroid spatial coordinates and pixel area of ​​each connected component are calculated to generate quantitative attributes of lesion location and size.

[0194] The navigation marker generation unit is used to generate real-time navigation markers based on quantized attributes. It maps the centroid coordinates of the lesion to the endoscopic video coordinate system through three-dimensional projection transformation to generate real-time navigation markers for the endoscopic video.

[0195] The clinical decision matching unit is used to perform clinical decision matching based on quantitative attributes. It generates a treatment decision plan that matches the clinical guidelines by querying the treatment rules in the preset guideline knowledge base that match the location and area of ​​the lesion.

[0196] In one embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described lesion segmentation method based on ENT endoscope images.

[0197] In one embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described lesion segmentation method based on ENT endoscope images.

[0198] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0199] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0200] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A lesion segmentation method based on otolaryngological endoscopy images, characterized in that, Includes the following steps: S1: Synchronously acquire the timing trigger signal of the endoscope light source, and capture timing-aligned dual-spectral images by alternately driving the white light source and the narrowband filter device, thereby generating a timing-aligned white light image sequence and a narrowband imaging image; S2: Model the mucus displacement of the white light image sequence, calculate the displacement vector of the dynamic mucus by the pixel gradient change of adjacent frames, and generate the mucus motion vector field; S3: Perform mask repair on the mucus motion vector field and the white light image sequence, segment the mucus-covered area by displacement amplitude threshold and reconstruct the occluded anatomical structure pixels to generate an optimized white light image with mucus interference removed. S4: Perform cross-modal feature registration on the optimized white light image and the narrowband imaging image, extract the white light feature map and the narrowband feature map respectively, and correct the spatial offset between the feature maps through a deformable convolutional layer to generate a registered multimodal feature map; S5: Perform lesion segmentation on the multimodal feature map, generate a lesion probability map by feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopy video based on the spatial coordinates and area attributes of the connected domains in the lesion probability map. The real-time navigation markers are used to indicate the anatomical location of the lesion and the probability grading of malignancy. The masking and repair of the mucus motion vector field and the white light image sequence, segmenting the mucus-covered area using a displacement amplitude threshold and reconstructing the occluded anatomical structure pixels to generate an optimized white light image free of mucus interference, includes: S31: Perform amplitude adaptive segmentation on the mucus motion vector field, and generate a dynamic threshold mask boundary associated with the mucus adhesion strength by calculating the spatiotemporal accumulation of motion vectors in the local neighborhood. S32: Reconstruct the anatomical structure of the covered area of ​​the dynamic threshold mask boundary, restore the continuity of the mucosal gland direction in the occluded area through the neighborhood healthy tissue texture propagation algorithm, and generate a preliminary reconstructed white light image. S33: The edge consistency of the white light reconstructed image is optimized by multi-scale gradient fusion to eliminate step artifacts in the texture transition area and generate an optimized white light image with continuous anatomical structure.

2. The method according to claim 1, characterized in that, S1 includes: S11: Perform time-division triggering processing on the control signal of the endoscope light source, and generate time-division imaging control command without spectral crosstalk by alternately switching the phase of the driving current of the white light source and the narrowband filter. S12: Input the time-division imaging control command into the image sensor driving circuit of the endoscope imaging system, perform frame synchronization processing on the raw photoelectric conversion data captured by the sensor, and generate a bispectral image with exposure time alignment by compensating for the start time offset between white light and narrowband exposure. S13: Perform motion artifact detection processing on the dual-spectral image, and remove distorted frames caused by organ tremors by cross-correlation analysis of adjacent frames in the white light channel and the narrowband channel, and generate a time-aligned white light image sequence and narrowband imaging image.

3. The method according to claim 1, characterized in that, S2 includes: S21: Perform viscous rheological property analysis on the white light image sequence, calculate the constitutive equation parameters of viscous shear stress and strain rate through the non-rigid transformation relationship of pixel gradient fields between adjacent frames, and generate viscous viscoelastic parameters. S22: Model the displacement field of the viscoelastic parameters of the viscous fluid, solve the displacement vector equation through viscous fluid dynamic constraints, substitute the viscoelastic parameters into the preset Oldroyd-B model to calculate the stress tensor distribution of the viscous fluid flow, and generate a dynamic viscous fluid displacement field. S23: Noise suppression is performed on the dynamic mucus displacement field by separating the anatomical structure motion and mucus flow components through low-pass filtering, filtering out abnormal pulsation signals with frequencies higher than the physiological motion threshold, and generating a denoised mucus motion vector field.

4. The method according to claim 1, characterized in that, S4 includes: S41: Perform multi-resolution feature extraction on the optimized white light image, and generate a multi-scale feature response map covering macroscopic anatomical structures to microscopic textures through convolutional pyramid decomposition, thereby generating a cross-scale white light feature map. S42: Enhance the frequency domain features of the narrowband imaging image by bandpass filtering to strengthen the frequency domain response of blood vessel morphology and mucosal surface gland openings corresponding to the 415nm and 540nm bands, and generate an enhanced narrowband feature map. S43: Perform spatial offset correction on the white light feature map and the enhanced narrowband feature map, learn the feature deformation parameters between the white light and the narrowband features through deformable convolution kernels, and generate a registered multimodal feature map.

5. The method according to claim 4, characterized in that, The formula for calculating the characteristic deformation parameter is as follows: ; ; in, For characteristic deformation parameters, The target location coordinates, The sampling grid is 3×3. For bilinear interpolation weights, It is characterized by narrow bands. For learnable offsets, It is characterized by white light.

6. The method according to any one of claims 1-5, characterized in that, S5 includes: S51: Perform feature channel splicing processing on the multimodal feature map, generate a fused feature tensor by splicing white light and narrowband feature map along the channel axis, and input the fused feature tensor into a U-shaped decoding network for upsampling and feature fusion to generate a lesion probability map; S52: Perform connected component attribute parsing processing on the lesion probability map. By scanning connected regions in the probability map that are greater than a set threshold, calculate the centroid spatial coordinates and pixel area of ​​each connected region to generate quantitative attributes of lesion location and size. S53: Based on the quantified attributes, real-time navigation markers are generated. The centroid coordinates of the lesion are mapped to the endoscopic video coordinate system through three-dimensional projection transformation to generate real-time navigation markers for the endoscopic video.

7. A lesion segmentation system based on otolaryngological endoscopic images, characterized in that, The system includes: The dual-spectral image synchronous acquisition module is used to synchronously acquire the timing trigger signal of the endoscope light source. By alternately driving the white light source and the narrowband filter device, it captures timing-aligned dual-spectral images and generates timing-aligned white light image sequences and narrowband imaging images. The slime displacement modeling module is used to model the slime displacement of the white light image sequence, calculate the displacement vector of the dynamic slime by the pixel gradient change of adjacent frames, and generate a slime motion vector field. The slime mask repair and image optimization module is used to perform mask repair on the slime motion vector field and the white light image sequence, segment the slime-covered area by displacement amplitude threshold and reconstruct the occluded anatomical structure pixels to generate an optimized white light image with slime interference removed. The cross-modal feature registration module is used to perform cross-modal feature registration on the optimized white light image and the narrowband imaging image, extract the white light feature map and the narrowband feature map respectively, and correct the spatial offset between the feature maps through a deformable convolutional layer to generate a registered multimodal feature map; The lesion segmentation and treatment plan generation module is used to segment lesions from the multimodal feature map, generate a lesion probability map through feature channel splicing and U-shaped decoding network, and output real-time navigation markers superimposed on the endoscopic video based on the spatial coordinates and area attributes of the connected domains in the lesion probability map. The real-time navigation markers are used to indicate the anatomical location of the lesion and the probability grading of malignancy. The masking and repair of the mucus motion vector field and the white light image sequence, segmenting the mucus-covered area using a displacement amplitude threshold and reconstructing the occluded anatomical structure pixels to generate an optimized white light image free of mucus interference, includes: The amplitude adaptive segmentation of the mucus motion vector field is performed, and a dynamic threshold mask boundary associated with the mucus adhesion strength is generated by calculating the spatiotemporal accumulation of motion vectors in the local neighborhood. Anatomical structure reconstruction is performed on the coverage area of ​​the dynamic threshold mask boundary, and the continuity of the mucosal gland direction in the occluded area is restored by the neighborhood healthy tissue texture propagation algorithm to generate a preliminary reconstructed white light image. Edge consistency optimization is performed on the white light reconstructed image, and step artifacts in the texture transition area are eliminated by multi-scale gradient fusion to generate an optimized white light image with continuous anatomical structure.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal medical image registration optimization method and system

    CN119832032A

  • Multi-modal multi-view-angle endoscopic image segmentation method, device and equipment and medium

    CN119991621A

  • Fiber endoscope image focus detection method and system

    CN120543553A