Early laryngeal cancer grading method based on laryngoscope image deep learning

By constructing a priori spectrum of high-brightness pseudo-peaks and a phase consistency spectrum verification mechanism, the specular reflection artifacts in laryngoscope images are separated from the features of real lesions. Combined with causal verification maps and three-domain coupling control commands, the problem of misjudgment of specular reflection artifacts in laryngeal cancer grading is solved, and the accuracy and reliability of laryngeal cancer grading are achieved.

CN121921289APending Publication Date: 2026-04-24THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV
Filing Date
2026-01-13
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, laryngeal cancer grading methods based on laryngoscopy images are prone to misjudgment when facing laryngeal secretions, membranes, or postoperative scar tissue due to specular reflection, which can lead to bright artifacts and affect the accuracy and reliability of early laryngeal cancer grading. The risk of misjudgment is especially high in dynamic scenarios.

Method used

By constructing a cross-timescale optical and temporal joint observation baseline, a high-brightness pseudo-peak prior spectrum is generated. Combined with phase consistency spectrum and curvature gradient verification mechanism, specular reflection artifacts and real lesion features are separated, an edge evidence set is constructed, dynamic optical perturbation conditions are simulated, a lesion fingerprint database is generated, and the identification and acquisition closed loop is realized through causal verification diagram and three-domain coupled control command, thereby weakening specular reflection interference and enhancing lesion edge response.

Benefits of technology

It effectively avoids high-level misjudgments caused by optical artifacts, improves the grading capability of laryngoscope images in complex dynamic scenarios, and provides more reliable support for early laryngeal cancer screening and assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921289A_ABST
    Figure CN121921289A_ABST
Patent Text Reader

Abstract

The invention discloses an early laryngeal cancer grading method based on laryngoscope image deep learning, and relates to the technical field of medical image intelligent diagnosis, and the method comprises the following steps: S1, in an image collection process, establishing a cross-time-scale optical and time sequence combined observation baseline, injecting a micro-amplitude polarization mark through multi-frame image collection, constructing a specular reflection path dictionary, and obtaining a specular reflection path dictionary; generating a highlight false peak prior map, and forming a space-time anchor point for positioning a reflection area; and S2, performing image matching operation based on the highlight false peak prior map, extracting a brightness anomaly core region by applying a phase congruency spectrum and curvature gradient combined verification mechanism, separating specular reflection artifacts from real lesion features, and constructing an edge evidence set to express a lesion structure contour. According to the method, the reflection prior map and a causal checking mechanism are constructed, phase conjugate regulation and control and light field remodeling are combined, accurate inhibition of specular reflection artifacts and stable extraction of focus characteristics are achieved, and grading accuracy and diagnosis reliability of laryngoscope images in a dynamic scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent diagnostic technology for medical imaging, specifically to an early laryngeal cancer grading method based on deep learning of laryngoscopy images. Background Technology

[0002] Early laryngeal cancer grading based on deep learning of laryngoscopy images involves using images of internal laryngeal tissue acquired through a laryngoscopy. A deep learning model performs image matching, feature extraction, and intelligent analysis on these images to automatically identify potential early lesion areas. The degree of laryngeal cancer development is then graded based on imaging features such as lesion morphology, texture, color, and vascular distribution. This method establishes an image matching model, enabling the system to find feature patterns similar to typical lesion images within a large dataset of laryngoscopy images, achieving automatic differentiation between different stages from healthy tissue to cancer. This not only improves the accuracy and objectivity of doctors in early laryngeal cancer diagnosis but also reduces subjective bias caused by manual observation, providing intelligent auxiliary data for early screening, disease assessment, and individualized treatment of laryngeal cancer.

[0003] Existing technologies have the following shortcomings: In existing technologies, laryngeal cancer grading based on laryngoscopy images typically relies on deep convolutional networks to extract and analyze features such as brightness, edge contours, and texture gradients. However, when there is a secretory membrane or postoperative scar tissue in the larynx, a highly reflective specular layer easily forms on its surface. Light waves undergo dynamic reflection and multiple energy superpositions during incident and reflection processes, forming localized high-brightness pseudo-feature areas. These high-brightness areas exhibit transient enhancement, clear boundaries, and abrupt grayscale changes in time-series images. Their appearance in the deep convolutional feature space is extremely similar to the edge of a real tumor, causing the model to be unable to accurately distinguish between optical reflection artifacts and real lesion features during image matching and feature recognition. This easily leads to misclassification of normal tissue as high-grade malignant lesions, resulting in severely overestimation of early laryngeal cancer grading results. This affects the medical credibility and clinical diagnostic reliability of the recognition model, especially in dynamic scenarios such as patient breathing, humidity fluctuations, or changes in lighting angles, where this risk of misjudgment is further amplified.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide an early laryngeal cancer grading method based on deep learning of laryngoscopy images, in order to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an early laryngeal cancer grading method based on deep learning of laryngoscope images, comprising the following steps: S1. During the image acquisition process, a cross-timescale optical and temporal joint observation baseline is established. Micro-amplitude polarization markers are injected through multi-frame image acquisition to construct a specular reflection path dictionary, generate a high-brightness pseudo-peak prior map, and form a spatiotemporal anchor point for locating the reflection area. S2, based on the prior spectrum of high-brightness pseudo-peaks, performs image matching operation, applies the phase consistency spectrum and curvature gradient joint verification mechanism to extract the core region of brightness abnormality, separates specular reflection artifacts from real lesion features, and constructs an edge evidence set to express the lesion structure contour; S3 utilizes edge evidence sets to perform temporal counterfactual playback, simulates respiratory rhythm and light source angle changes to reconstruct dynamic image sequences, examines the stability of edge features under optical perturbation conditions, eliminates unstable texture regions, and generates a reliable lesion fingerprint database. S4. Based on the lesion fingerprint database, a causal verification map is constructed. The displacement trajectory of the reflection source and its influence on the depth feature space are analyzed. Three-domain coupled control commands for exposure parameters, polarization angle and image acquisition posture are generated. Control signals are sent back to the image acquisition link to construct a closed loop for recognition and acquisition. After the closed-loop identification and acquisition is running stably, S5 initiates a dynamic control mechanism, performs phase conjugate polarization rotation scanning, and links the time grid migration and energy return suppression processes to perform time-series inversion and real-time reshaping of the incident conical light field, weakening specular reflection interference and enhancing lesion edge response, thereby realizing a dynamic and adaptive laryngeal cancer grading control process.

[0007] Preferably, step S1 includes: By setting a joint observation baseline and defining parameters such as the imaging viewing angle range, the light source incident angle range, and the imaging time interval, a cross-timescale observation frame sequence structure is constructed, enabling image acquisition to have controllable temporal variations and light field incident differences. Multi-frame image acquisition was performed under the cross-timescale observation frame sequence structure, and micro-amplitude polarization markers were injected into each frame image to make the reflection area form a regular polarization response trajectory in the continuous frames. Based on the differences in polarization response in the image sequence, abnormal regions of light intensity change are extracted, the light intensity peak drift trajectory is tracked with the time axis as the index, and a specular reflection path dictionary is constructed. A prior map of high-brightness pseudo-peaks is generated based on a specular reflection path dictionary, and the reflection behavior trajectory is mapped according to the image spatial coordinate system to form a spatial-temporal prior distribution map of the reflection area.

[0008] Preferably, step S2 includes: Regions with probability values ​​higher than a threshold in the prior spectrum of high-brightness pseudo-peaks are selected as candidate regions for image matching. Pixel-level sliding window matching is performed in the original image frame to extract the spatial response region. Extract the phase consistency spectrum features of the matching candidate region in the frequency domain, and determine its structural synergy in multiple directions and multiple frequency bands to identify the core region of brightness anomaly; By combining phase consistency spectrum features, curvature gradient feature analysis is performed on candidate regions to extract curvature continuity and gradient transition to screen out edge artifacts formed by optical reflection; The regions preserved after double verification are used to generate an edge evidence set, forming a continuous boundary set that expresses the structural contour of the lesion, and then undergoing post-processing.

[0009] Preferably, the edge evidence set is only included in the generation range when the phase consistency spectrum features of the image matching candidate region meet the conditions of stable phase response and high synergy in multiple directions.

[0010] Preferably, step S3 includes: Using edge evidence sets as input, and combining image acquisition parameters, a dynamic image sequence with respiratory rhythm and light source angle perturbation features is constructed to simulate the edge response trajectory in dynamic acquisition scenarios. The response changes of edge evidence points in dynamic image sequences are tracked along the time axis to evaluate their brightness gradient, texture continuity and spatial stability, and low-confidence edge points are screened out. The set of highly stable edge points is spatially reorganized to construct a coherent edge graphic structure, which serves as the basis for representing the boundary of the lesion region. Based on stable edge structures, texture features, dynamic response and spatial shape information are extracted, a structure-time response triple mapping model is constructed and a lesion fingerprint database is generated.

[0011] Preferably, when constructing the lesion fingerprint database, the dynamic response of the edge structure is judged by the combined criteria of brightness change amplitude, texture direction consistency and boundary continuity, and only the edge regions that maintain stable response in all image frames are retained as the basis for fingerprint generation.

[0012] Preferably, step S4 includes: Based on the edge structure and brightness response information in the lesion fingerprint database, the spatial distribution trajectory of the optical interference source is deduced in reverse, a displacement model of the reflection source is constructed, and the interference behavior is calibrated. A causal verification graph is constructed based on the displacement model of the reflection source to quantify the causal relationship between the behavior nodes of the reflection source and the perturbation of image features, and the interference path and intensity are recorded. Based on the causal verification map, a three-domain coupled control command containing exposure parameters, polarization angle and image acquisition posture is generated and used to optimize image acquisition behavior; The control commands are sent back to the image acquisition chain to build a dynamic closed-loop control between the recognition chain and the acquisition chain, enabling real-time adjustment and feedback updates of image acquisition parameters.

[0013] Preferably, in the three-domain coupled control command, the exposure parameters are adjusted synchronously by controlling the light source illuminance and exposure time, the polarization angle is dynamically adjusted according to the polarization sensitive range of the reflection source, and the image acquisition posture controls the lens incident angle according to the reflection path direction, so as to minimize optical interference and maximize the clarity of lesion edges.

[0014] Preferably, step S5 includes: A phase conjugate modulation strategy is set up to adjust the emission phase of the light wave based on the reflection source trajectory data in order to achieve interference reduction and reduce high brightness interference response; The polarization rotation scanning operation is performed based on the reflection path dictionary, and the polarization direction of the light source is continuously changed within the image acquisition cycle, thereby enhancing the feature separation capability between the lesion edge and the reflection area. A time grid migration and energy reflection suppression mechanism is introduced. By synchronously adjusting the light source start-up delay and acquisition time, the incident cone-shaped light field is reconstructed to weaken the reflected echo energy. The dynamic response evaluation mechanism is activated to analyze the texture clarity and brightness balance of the lesion edge area in real time under the current light field configuration, and to complete the adaptive optimization and adjustment of the light field.

[0015] Preferably, the phase conjugate modulation strategy introduces a phase adjustment unit at the light source emission end to set the phase of the emitted light wave to the conjugate state of the reflected wave of the previous cycle, so that the interference effect forms a local energy cancellation region on the tissue surface, which is used to weaken the specular reflection high brightness response.

[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention accurately identifies optical artifacts and constructs an edge evidence set by building a high-brightness pseudo-peak prior map and a specular reflection path dictionary, combined with a joint discrimination strategy of phase consistency spectrum and curvature gradient. Temporal counterfactual verification is performed under simulated respiratory rhythm and illumination changes to screen out lesion structural features with stable performance and construct a fingerprint database. Furthermore, the interference effect of reflection sources on the feature space is precisely quantified through a causal check map, generating three-dimensional coupled control instructions to achieve reverse closed-loop control of the acquisition behavior based on the identification results. Finally, through phase conjugate polarization scanning and real-time reshaping of the conical light field, active suppression of reflected energy and enhanced response to lesion edges are achieved, establishing a dynamic adaptive control mechanism that deeply integrates identification and acquisition. This scheme effectively avoids the high-level misjudgment problem caused by optical artifacts in traditional methods, improves the intelligent grading capability of laryngoscope images in complex dynamic scenes, and provides more reliable technical support for accurate screening, objective assessment, and clinical decision-making for early laryngeal cancer. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0018] Figure 1 This is a flowchart of the method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to the present invention. Detailed Implementation

[0019] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0020] This invention provides, for example Figure 1 The early laryngeal cancer grading method based on deep learning of laryngoscopy images, as shown, includes the following steps: S1. During the image acquisition process, a cross-timescale optical and temporal joint observation baseline is established. Micro-amplitude polarization markers are injected through multi-frame image acquisition to construct a specular reflection path dictionary, generate a high-brightness pseudo-peak prior map, and form a spatiotemporal anchor point for locating the reflection area. This step addresses the severe interference from specular reflection artifacts during laryngoscope image acquisition. It proposes an image acquisition and preprocessing method combining optical and temporal observation techniques. By constructing a reflection path dictionary and a priori map of pseudo-peaks, potential reflection interference regions in the image are accurately located, enhancing the stability and accuracy of subsequent feature recognition. The specific steps are as follows:

[0021] Before image acquisition, a joint observation baseline is preset, and the imaging field of view, incident angle range of the light source, and imaging time interval parameters of the laryngoscope are spatially and temporally calibrated to construct a cross-timescale observation frame sequence structure. Under this frame sequence structure, each group of consecutive image acquisition frames is sequentially excited according to a specific time period and incident direction, enabling controllable temporal variations and light field incident differences in the imaging process. The core of this step lies in achieving diverse presentations of the same laryngeal region under different illumination angles during image acquisition through artificially designed temporal scheduling and optical path fine-tuning, providing basic data support for subsequent tracking of reflection behavior. To ensure the alignment of the tracking path, a slight deformation of the glottis is induced by setting a respiratory rhythm simulation device during the pre-acquisition stage, enhancing the observability of the target area in the temporal dimension, thereby improving the tracking extension capability of the reflection area in the image sequence along the time axis.

[0022] Under the aforementioned cross-timescale observation frame sequence structure, multiple frames of images were acquired, and a slight change in polarization angle was introduced within the imaging cycle of each frame. Specifically, the polarizer angle at the light source emission end was rotated by a preset amplitude between each exposure. The rotation amplitude was controlled within a perceptible but non-destructive range of tissue optical properties, resulting in subtle differences in the polarization response of the reflection characteristics formed by the incident light on the tissue surface. Because the specular reflection region is more sensitive to changes in polarization direction, it will exhibit regular oscillations in brightness response in consecutive frames, while real tissue will exhibit a stable or low-fluctuation response. This polarization marker injection behavior persists throughout the entire frame sequence image acquisition cycle, giving each reflection point a unique trajectory of optical behavior in consecutive frames. Based on this difference in optical behavior, a discrimination criterion is provided for subsequently extracting reflection path features from the image sequence.

[0023] Based on the image sequence after polarization marker injection, regions with abnormal light intensity changes in all frames were extracted. Highlighted regions exhibiting specular reflection characteristics were located using the continuous brightness variation trend and the brightness difference relationship between adjacent frames. During the localization process, the peak shift behavior of each pixel's light intensity in consecutive frames was analyzed using the time axis as an index. Regions exhibiting periodic enhancement or grayscale jumps were classified as candidate high-brightness pseudo-peak regions. Combining the previously set imaging angle and incident direction records, the positional information of each high-brightness region was traced back to its source, constructing its change trajectory in image space and observation time axis. Furthermore, high-brightness regions with similar response characteristics and spatial trajectory patterns were clustered and normalized to establish a specular reflection path dictionary. This dictionary records the mapping relationship and behavioral patterns of each artifact region in different frames, forming a structural description of specular reflection behavior. This dictionary is a crucial reference basis for distinguishing between real lesions and optical reflections in subsequent identification.

[0024] Based on the established specular reflection path dictionary, a corresponding high-brightness pseudo-peak prior map is generated. This map, based on the path information and brightness variation features recorded in the dictionary, remaps all reflection trajectories according to the image spatial coordinate system, forming a spatial-temporal prior distribution map covering the entire set of image frames. The grayscale value of each pixel in the map represents its probability level of being identified as a specular reflection artifact in multiple image frames. The statistical stability of the map is further enhanced by overlaying multiple sets of image acquisition results. This map, as a probabilistic representation model of reflection regions, not only possesses spatial distribution accuracy but also temporal evolution clues, providing a localization reference for subsequent image feature recognition. By using this map, optical feature masking or weighted correction can be applied to suspected reflection regions before entering the image recognition stage, effectively reducing the risk of misidentification caused by reflection interference and improving the deep learning model's ability to identify real lesion structures.

[0025] S2, based on the prior spectrum of high-brightness pseudo-peaks, performs image matching operation, applies the phase consistency spectrum and curvature gradient joint verification mechanism to extract the core region of brightness abnormality, separates specular reflection artifacts from real lesion features, and constructs an edge evidence set to express the lesion structure contour; After constructing the prior map of the high-brightness pseudo-peaks, this step further conducts specular reflection artifact recognition based on joint discrimination of image space and frequency domain. A multi-dimensional feature cross-validation mechanism separates the optical reflection region from the real lesion region, and finally extracts a high-confidence edge information set to represent the lesion structure boundary. The specific steps are as follows:

[0026] Based on the prior map of the high-brightness pseudo-peaks, regions in the current time frame with a probability value higher than a set confidence threshold are selected as candidate regions for image matching, and image patch-level similarity matching is performed. This matching operation uses the original image frame as the retrieval space and the highly reflective regions marked in the prior map as guiding templates. Pixel-level sliding window matching is performed in the image space to compare the local similarity of adjacent regions in brightness distribution, texture directionality, and gray-level hierarchical structure. During the matching process, considering the slight structural shifts caused by polarization rotation, viewing angle adjustment, and tissue deformation during image acquisition, a multi-scale sampling strategy is adopted to evaluate the hierarchical distribution of similar patterns within the region, thereby improving the adaptability of the matching results to local image distortions. Through this operation, a spatial response region corresponding to the prior map of the high-brightness pseudo-peaks can be established in the current frame image, providing a positional reference and boundary framework for subsequent frequency domain feature extraction.

[0027] After image matching is completed, candidate regions from the matching results are used as the analysis object, and their phase consistency spectrum features in the frequency domain are extracted. Specifically, a frequency domain transformation is performed on each candidate region to extract its phase information distribution across multiple directions and frequency bands, and its phase consistency level is calculated in each direction. This feature measures the structural coherence of a local image across different frequency components and is highly sensitive to changes in image edges and structure. Due to their optical origins, specular reflection regions often exhibit high-intensity but low-structure random textures; therefore, their phase consistency spectrum distribution typically shows characteristics such as large directional differences, discrete frequency distribution, and concentrated intensity response. Conversely, real lesion regions exhibit high directional consistency and multi-spectral coherence in the frequency domain, characterized by a stable distribution of phase response and high coherence spectral values. Therefore, by comparing the response patterns of the phase consistency spectrum in multiple directions, it is possible to preliminarily determine whether lesion information with real structural edge features exists within the region.

[0028] Building upon the extraction of phase features in the frequency domain, a joint verification analysis is further performed by incorporating curvature gradient features in the spatial domain. In this step, edge contour extraction is performed on each candidate region, and higher-order derivatives are used to fit the curvature of the brightness variation surface, extracting the rate of curvature change and gradient direction fluctuation values ​​in the local region of the image. Specular reflection artifacts typically appear in images as sharp boundaries with abrupt curvature changes and poor edge extensibility, while the edges of genuine lesion structures usually possess a certain degree of continuous curvature change and smooth gradient transition. Therefore, by analyzing the continuity of edge curvature, the consistency of curvature direction, and the distribution of gradient intensity, edge artifacts caused by optical reflection can be further eliminated, enhancing the confidence in identifying genuine lesion boundaries. During this process, curvature gradient features are cross-referenced with the phase consistency spectrum features extracted in the previous step, forming complementary enhancements in the spatial and frequency domains, thereby further improving the robustness of the judgment results.

[0029] The image regions retained through dual verification in both the frequency and spatial domains are used as the final identification results for the core regions of brightness anomalies, generating a centralized edge evidence set. This evidence set is composed of pixel-level boundary points, forming continuous closed or semi-closed structural contours in image space to represent the possible anatomical edges of lesions. To improve the continuity and accuracy of edge representation, post-processing operations are performed on the edge evidence set, including edge point connection interpolation, morphological closure processing, and interference edge removal, to ensure the spatial consistency and structural coherence of the evidence set. Simultaneously, combined with the polarization angle and incident direction sequence set during image acquisition, the edge evidence set is mapped and tracked across multiple frames of images, eliminating unstable edge fragments under different lighting conditions and retaining edge sets with stable response patterns as reliable lesion structure representation results. The final constructed edge evidence set will serve as a key input for subsequent lesion fingerprint database generation and counterfactual playback verification, achieving accurate separation of specular reflection artifacts and effective representation of the true lesion contour.

[0030] S3 utilizes edge evidence sets to perform temporal counterfactual playback, simulates respiratory rhythm and light source angle changes to reconstruct dynamic image sequences, examines the stability of edge features under optical perturbation conditions, eliminates unstable texture regions, and generates a reliable lesion fingerprint database. This step, building upon the completed edge evidence set, introduces a temporal perturbation playback mechanism to simulate the image acquisition process under varying respiratory rhythms and light source angles. This dynamically evaluates the optical response stability of edge features, thereby screening stable lesion feature regions and constructing a reliable lesion fingerprint database. The specific steps are as follows:

[0031] Using the generated set of edge evidence as input, the edge positions are temporally calibrated according to the recording parameters at the time of image acquisition, and a replayable dynamic image sequence structure is constructed. This sequence uses image frames as basic units, and a time index is used to map the spatial positions of edge points in different frames to the lighting environment. To enhance the realism of the simulated scene, the range of respiratory rhythm changes set during the acquisition phase is mapped to a geometric perturbation model in the image sequence. This model applies periodic displacement and slight scaling in the vertical and forward / backward directions, respectively, to simulate the tissue deformation of the glottis region during respiration. Simultaneously, the incident angle of the light field in each frame is reconstructed according to the light source incident direction control parameters at the image capture time, thereby restoring the realistic scene representation of the laryngoscope image under different lighting perturbation conditions. Through the joint modeling of respiratory geometric perturbation and lighting changes, a dynamic image sequence with temporal coherence and optical perturbation characteristics is synthesized based on the original image frames, to represent the trajectory of edge evidence points under dynamic conditions.

[0032] Frame-by-frame response tracking is performed on all edge evidence points in the dynamic image sequence along the time axis to establish the response change curve of each edge point across multiple frames, and its stability is quantitatively evaluated. In this step, the grayscale gradient value, brightness response intensity, boundary continuity, and local texture consistency of the edge points in image space are used as evaluation indicators. Response values ​​are extracted from each frame and a time-series data structure is constructed. Based on this, the spatial positional drift amplitude and texture expression consistency of the edge points at different time points are compared longitudinally to determine whether the edge point maintains a high degree of image expression stability under perturbation conditions. For edge points exhibiting abrupt brightness changes, texture disappearance, or edge breakage in some frames, their stability level is set to low confidence, thus being filtered out in subsequent processing. This process not only examines the brightness stability of edge points but also the temporal coherence of their structural features to minimize the misidentification of specularly bright areas caused by illumination changes as actual lesion edges.

[0033] Based on the stability assessment of edge points, the entire edge evidence set is screened, removing edge point sets that fail to maintain continuous expression in most frames, have severe texture information interruptions, or exhibit spatial drift exceeding a preset threshold. This screening operation is based on the geometric continuity and visual consistency of the lesion structure as the judgment criteria, ensuring that the retained edge features have highly overlapping expression forms under different observation conditions. Simultaneously, the performance records of edge points in the dynamic sequence are retained during the edge point removal process to verify the removal criteria in subsequent retrospective analysis. For the retained edge feature point set, further spatial reconstruction is performed. A structural map is constructed according to the spatial connectivity of the edge points, isolated edge points are deleted, and continuous edge segments are retained to form complete contour lines. Morphological expansion methods are used to smoothly fill the gaps between edges to generate a complete and clearly defined edge graphic structure. This structure serves as the boundary representation basis for the possible lesion region, providing a geometric framework for the final feature extraction.

[0034] Based on the selected stable edge structures, a reliable lesion fingerprint database is constructed. The fingerprint database construction process involves three dimensions of content extraction and integration: First, texture description information within the corresponding region of the edge structure is extracted, including directional gradient distribution, grayscale level structure, and detail distribution density, to characterize the local image features of the lesion region; second, the brightness response pattern and boundary dynamic change trajectory exhibited by this region in a time-series image sequence are extracted as a behavioral description of the lesion region under dynamic optical perturbation conditions; finally, the location information, shape contour, and stable expression time span of the region in the entire image space are integrated to establish a structure-time-response triple mapping model. By summarizing and encoding the above information, each stable lesion region is formed into an independent fingerprint entry and stored in the lesion fingerprint database. This fingerprint database will serve as an important basis for subsequent lesion identification, causal inference, and acquisition control, realizing the construction of a stable and reliable lesion image identification system starting from edge features, and providing a highly reliable structural reference template for the image recognition process.

[0035] S4. Based on the lesion fingerprint database, a causal verification map is constructed. The displacement trajectory of the reflection source and its influence on the depth feature space are analyzed. Three-domain coupled control commands for exposure parameters, polarization angle and image acquisition posture are generated. Control signals are sent back to the image acquisition link to construct a closed loop for recognition and acquisition. After establishing the lesion fingerprint database, this step further focuses on reflection interference control and image acquisition optimization, constructing a dynamic modeling and control feedback mechanism for reflection sources. This achieves closed-loop control between image recognition and image acquisition, enhancing the stability of lesion recognition and the consistency of image acquisition. The specific steps are as follows:

[0036] Based on the edge structure, brightness response, and texture representation of lesion regions recorded in the lesion fingerprint database within image sequences, the distribution trajectory of optical interference sources highly correlated with these regional features in image space is deduced in reverse, constructing a reflection source displacement model. In the specific implementation, the time-series representation of lesion edge points is used as the core clue to extract the light spot drift trajectory, brightness peak drift path, and polarization response fluctuation range of each lesion region across multiple frames. Combining the relative positional distribution of lesion regions in image space, spatial interpolation is used to connect the optical center points represented by artifacts in each frame, forming a complete reflection source motion path map. Based on this, the dynamic coupling relationship between the reflection source displacement trajectory and the lesion region feature representation is further identified, clarifying the degree of interference of reflection source position changes on edge feature sharpness, grayscale stability, and texture boundary integrity. This reverse modeling process not only reconstructs the dynamic behavior of reflection interference sources during acquisition but also establishes a causal relationship between them and image quality changes, providing a baseline for subsequent control command generation.

[0037] Based on the reflection source displacement model and interference impact analysis, a causal verification graph is constructed to quantify the actual impact of interference sources on the feature space of depth images. In this graph, the image acquisition frame sequence time axis is used as the horizontal axis, and the feature change index of the lesion region in the image space is used as the vertical axis to establish the connection channel between each reflection source behavior node and the feature perturbation index. For example, when a change in polarization angle within a specific time period leads to a decrease in edge sharpness or a disruption of texture continuity, the causal verification graph will record the connection between the reflection source behavior node and the edge distortion event, and mark its influence intensity level. By traversing all reflection source interference events in the entire image sequence and their corresponding responses in the depth feature space, a causal network map between reflection behavior and feature changes is gradually drawn. This map not only characterizes which reflection perturbations substantially interfere with the image feature expression, but also depicts the entire causal path structure of how feature stability evolves under multiple superimposed interference conditions.

[0038] Based on the interference path and behavioral influence intensity information quantified in the causal verification graph, the optical parameters involved in the image acquisition process are optimized and adjusted in reverse, and a three-domain coupled control command is generated accordingly. This control command covers three main dimensions: exposure parameters, polarization angle, and image acquisition posture. Regarding exposure parameter adjustment, for periods of localized overexposure due to enhanced reflection, the light source illuminance is appropriately reduced or the exposure time is shortened to reduce the interference expression of brightness peak areas in the image. Regarding polarization angle control, considering the relative spatial relationship between the reflection source and the lesion structure, the angle combination that minimizes polarization-sensitive response is selected to stabilize or weaken the brightness response of the specular reflection area. Regarding image acquisition posture control, the lens incident angle is adjusted according to the projection direction of the reflection path, creating an angular offset between the reflection direction and the lens receiving angle, thereby avoiding the direct incident path and weakening the specular reflection incident signal. These three types of control commands form a multi-dimensional coupled feedback adjustment framework, which is dynamically generated and sent to the acquisition device interface before the image acquisition process enters the next frame or cycle, guiding real-time optimization and adjustment of the acquisition behavior.

[0039] After the control instructions are generated, the aforementioned coupling instructions are transmitted back to the image acquisition chain as a synchronization signal, establishing a real-time interactive closed loop between the acquisition chain and the recognition chain. During this closed-loop control process, the acquisition chain actively adjusts the acquisition scheme for the next image frame based on the transmitted exposure adjustment values, polarization angle settings, and attitude transformation instructions. This ensures that the effect of specular reflection is minimized while fully preserving the lesion edges and texture information. Simultaneously, the recognition chain receives parameter response data from the acquisition chain in real time, which is used to correct the feature evolution path of the lesion region and update the dynamic response template in the lesion fingerprint database. Thus, in continuous multi-frame acquisition, the recognition results not only affect the acquisition strategy, but the feedback from the acquisition behavior also optimizes the recognition path, ultimately achieving dynamic linkage, real-time control, and performance enhancement between recognition and acquisition. Throughout the entire image processing process, the lesion region remains in an optimal acquisition response state, and the image data quality tends to be consistent and optimized under the linkage of multi-dimensional parameters, ensuring that subsequent diagnostic and grading tasks are based on stable, realistic, and reliable images.

[0040] S5, after the closed-loop identification and acquisition is running stably, starts the dynamic control mechanism, performs phase conjugate polarization rotation scanning operation, links the time grid migration and energy return suppression process, performs time-series inversion and real-time reshaping of the incident conical light field, weakens specular reflection interference and enhances lesion edge response, and realizes the dynamic adaptive laryngeal cancer grading control process. This step, based on the stable operation of the closed-loop identification and acquisition, further activates a dynamic light field modulation mechanism. Through the synergistic effect of phase control, polarization rotation, time mapping, and energy reflection suppression, it achieves dynamic shaping of the image incident path to enhance the response of the actual lesion edge and reduce specular reflection interference, thereby constructing a dynamic and adaptive laryngeal cancer grading control process. The specific steps are as follows:

[0041] With the closed-loop control between image recognition and acquisition chains operating stably, a phase conjugate modulation strategy is established based on reflection source trajectory data and reflection behavior patterns in the causal verification diagram to achieve inverse optical field reconstruction. In the specific implementation, the phase mismatch region formed during the reflection of light waves on the tissue surface during image acquisition is first identified. Based on the light wave propagation path records in historical image frames, the location of phase shift in the reflection path is modeled. This modeling is based on the initial phase of the incident light wave and the phase perturbation data in the return path, calculating the conjugate deviation between it and the ideal reflected wave. Subsequently, by introducing a phase adjustment unit at the light source emission end, the phase of each emitted light wave is adjusted to the conjugate state of the reflected wave in the previous cycle, thereby causing the two beams to produce mutually canceling interference effects on the tissue surface. This phase conjugate strategy can effectively reduce the high-brightness interference response caused by reflection from tissue fluid surfaces or irregular scars in local areas, forming a spatially oriented optical interference reduction mechanism, providing a stable background reference for subsequent polarization control and energy management.

[0042] While activating the phase conjugation strategy, a polarization rotation scan operation is executed concurrently. This enhances the adaptability of the light field to specific reflection angles by periodically changing the polarization direction of the light source. During implementation, a polarization rotation timing table is constructed based on the intersection information of high-incidence reflection areas and polarization-sensitive angles recorded in the specular reflection path dictionary. This timing table arranges the polarization angles to be used in each cycle according to the image acquisition time window. Subsequently, within the image acquisition cycle, the polarization direction of the light source in each frame is continuously rotated, so that the polarization state of the light field during the incident process covers multiple directional distributions, thereby generating optical responses with different polarization sensitivities on the tissue surface. Since the specular reflection area is highly sensitive to specific polarization directions, rotation scanning can suppress the response in multiple directions in this area, thus significantly reducing the expression frequency of high-brightness artifacts in the image. At the same time, the real lesion area, due to its internal scattering characteristics, has strong polarization stability and can maintain clear texture and continuous edges under multiple polarization directions, thereby achieving feature separation between reflection and lesion in terms of contrast intensity.

[0043] Building upon phase control and polarization modulation, a time grid migration mechanism and an energy reflection suppression strategy are further introduced to comprehensively reconstruct the incident conical light field distribution during image acquisition. Specifically, by adjusting the time synchronization relationship between the light source's emitting element and the lens's incident optical channel, each frame's image acquisition window is divided into multiple time grids according to the periodic rhythm of respiratory rhythm and dynamic lesion behavior. Within each grid, by fine-tuning the light source's start-up delay and the image acquisition trigger time, incident synchronization of light waves at different phase states on the tissue surface is achieved, forming a conical light field sequence with time-jumping characteristics. Simultaneously, combined with the energy accumulation model of reflection interference behavior in the reflection path, an energy suppression threshold is set in the light field design. Angle shifting and intensity reduction operations are implemented for reflected beams exceeding the interference energy level, thereby preventing the formation of high-brightness clusters in the image center region. This step, through the synergy of time management and optical redirection, achieves active attenuation of specular reflection energy echoes and forms a more uniform brightness distribution structure in the image space.

[0044] Simultaneously with the completion of dynamic light field reconstruction, a dynamic response evaluation mechanism is activated to perform real-time quality analysis on image frames after phase control, polarization adjustment, temporal reconstruction, and energy suppression, in order to determine whether the enhancement effect of lesion edge response has reached a stable standard. Specifically, high-confidence edge regions stored in the lesion fingerprint database are used as reference regions, comparing their texture contrast, edge sharpness, and grayscale balance under the reconstructed light field, and simultaneously comparing them with the performance in standard acquisition frames. If the current light field configuration significantly improves edge sharpness and reduces high-brightness artifact interference, the current acquisition parameters are written into the acquisition control path for reference in subsequent frames; if the expected results are not achieved, the next control cycle begins, and the phase, polarization, and temporal configurations are readjusted. This cycle repeats, gradually approaching the optimal light field configuration state through multi-cycle acquisition and recognition feedback, ultimately achieving full-process adaptive control that reverses the acquisition behavior based on recognition requirements. Through the above-mentioned multi-level collaborative mechanism, the edge expression of the lesion region in the image sequence is continuously enhanced, while the specular reflection region is dynamically suppressed, thereby significantly improving the stability of the image structure and the accuracy of hierarchical judgment, and constructing a new model for laryngeal cancer image recognition under dynamic adaptive control.

[0045] This invention accurately identifies optical artifacts and constructs an edge evidence set by building a high-brightness pseudo-peak prior map and a specular reflection path dictionary, combined with a joint discrimination strategy of phase consistency spectrum and curvature gradient. Temporal counterfactual verification is performed under simulated respiratory rhythm and illumination changes to screen out lesion structural features with stable performance and construct a fingerprint database. Furthermore, the interference effect of reflection sources on the feature space is precisely quantified through a causal check map, generating three-dimensional coupled control instructions to achieve reverse closed-loop control of the acquisition behavior based on the identification results. Finally, through phase conjugate polarization scanning and real-time reshaping of the conical light field, active suppression of reflected energy and enhanced response to lesion edges are achieved, establishing a dynamic adaptive control mechanism that deeply integrates identification and acquisition. This scheme effectively avoids the high-level misjudgment problem caused by optical artifacts in traditional methods, improves the intelligent grading capability of laryngoscope images in complex dynamic scenes, and provides more reliable technical support for accurate screening, objective assessment, and clinical decision-making for early laryngeal cancer.

[0046] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for early laryngeal cancer grading based on deep learning of laryngoscopy images, characterized in that, Includes the following steps: S1. During the image acquisition process, a cross-timescale optical and temporal joint observation baseline is established. Micro-amplitude polarization markers are injected through multi-frame image acquisition to construct a specular reflection path dictionary, generate a high-brightness pseudo-peak prior map, and form a spatiotemporal anchor point for locating the reflection area. S2, based on the prior spectrum of high-brightness pseudo-peaks, performs image matching operation, applies the phase consistency spectrum and curvature gradient joint verification mechanism to extract the core region of brightness abnormality, separates specular reflection artifacts from real lesion features, and constructs an edge evidence set to express the lesion structure contour; S3 utilizes edge evidence sets to perform temporal counterfactual playback, simulates respiratory rhythm and light source angle changes to reconstruct dynamic image sequences, examines the stability of edge features under optical perturbation conditions, eliminates unstable texture regions, and generates a reliable lesion fingerprint database. S4. Based on the lesion fingerprint database, a causal verification map is constructed. The displacement trajectory of the reflection source and its influence on the depth feature space are analyzed. Three-domain coupled control commands for exposure parameters, polarization angle and image acquisition posture are generated. Control signals are sent back to the image acquisition link to construct a closed loop for recognition and acquisition. After the identification and acquisition closed-loop operation is stable, S5 starts the dynamic control mechanism, performs phase conjugate polarization rotation scanning operation, links the time grid migration and energy return suppression process, performs time-series inversion and real-time reshaping of the incident conical light field, weakens specular reflection interference and enhances the response of lesion edge.

2. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 1, characterized in that, Step S1 includes: By setting a joint observation baseline and defining parameters such as the imaging viewing angle range, the light source incident angle range, and the imaging time interval, a cross-timescale observation frame sequence structure is constructed, enabling image acquisition to have controllable temporal variations and light field incident differences. Multi-frame image acquisition was performed under the cross-timescale observation frame sequence structure, and micro-amplitude polarization markers were injected into each frame image to make the reflection area form a regular polarization response trajectory in the continuous frames. Based on the differences in polarization response in the image sequence, abnormal regions of light intensity change are extracted, the light intensity peak drift trajectory is tracked with the time axis as the index, and a specular reflection path dictionary is constructed. A prior map of high-brightness pseudo-peaks is generated based on a specular reflection path dictionary, and the reflection behavior trajectory is mapped according to the image spatial coordinate system to form a spatial-temporal prior distribution map of the reflection area.

3. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 1, characterized in that, Step S2 includes: Regions with probability values ​​higher than a threshold in the prior spectrum of high-brightness pseudo-peaks are selected as candidate regions for image matching. Pixel-level sliding window matching is performed in the original image frame to extract the spatial response region. Extract the phase consistency spectrum features of the matching candidate region in the frequency domain, and determine its structural synergy in multiple directions and multiple frequency bands to identify the core region of brightness anomaly; By combining phase consistency spectrum features, curvature gradient feature analysis is performed on candidate regions to extract curvature continuity and gradient transition to screen out edge artifacts formed by optical reflection; The regions preserved after double verification are used to generate an edge evidence set, forming a continuous boundary set that expresses the structural contour of the lesion, and then undergoing post-processing.

4. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 3, characterized in that, Only when the phase consistency spectrum features of the image matching candidate region meet the conditions of stable phase response and high synergy in multiple directions are they included in the generation range of the edge evidence set.

5. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 1, characterized in that, Step S3 includes: Using edge evidence sets as input, and combining image acquisition parameters, a dynamic image sequence with respiratory rhythm and light source angle perturbation features is constructed to simulate the edge response trajectory in dynamic acquisition scenarios. The response changes of edge evidence points in dynamic image sequences are tracked along the time axis to evaluate their brightness gradient, texture continuity and spatial stability, and low-confidence edge points are screened out. The set of highly stable edge points is spatially reorganized to construct a coherent edge graphic structure, which serves as the basis for representing the boundary of the lesion region. Based on stable edge structures, texture features, dynamic response and spatial shape information are extracted, a structure-time response triple mapping model is constructed and a lesion fingerprint database is generated.

6. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 5, characterized in that, When constructing the lesion fingerprint database, the dynamic response of the edge structure is judged by the combined criteria of brightness change amplitude, texture direction consistency and boundary continuity. Only the edge regions that maintain stable response in all image frames are retained as the basis for fingerprint generation.

7. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 1, characterized in that, Step S4 includes: Based on the edge structure and brightness response information in the lesion fingerprint database, the spatial distribution trajectory of the optical interference source is deduced in reverse, a displacement model of the reflection source is constructed, and the interference behavior is calibrated. A causal verification graph is constructed based on the displacement model of the reflection source to quantify the causal relationship between the behavior nodes of the reflection source and the perturbation of image features, and the interference path and intensity are recorded. Based on the causal verification map, a three-domain coupled control command containing exposure parameters, polarization angle and image acquisition posture is generated and used to optimize image acquisition behavior; The control commands are sent back to the image acquisition chain to build a dynamic closed-loop control between the recognition chain and the acquisition chain, enabling real-time adjustment and feedback updates of image acquisition parameters.

8. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 7, characterized in that, In the three-domain coupled control command, the exposure parameters are adjusted synchronously by controlling the light source illuminance and exposure time, the polarization angle is dynamically adjusted according to the polarization sensitive range of the reflection source, and the image acquisition posture controls the lens incident angle according to the reflection path direction.

9. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 1, characterized in that, Step S5 includes: A phase conjugate modulation strategy is set up to adjust the emission phase of the light wave based on the reflection source trajectory data in order to achieve interference reduction and reduce high brightness interference response; The polarization rotation scanning operation is performed based on the reflection path dictionary, and the polarization direction of the light source is continuously changed within the image acquisition cycle, thereby enhancing the feature separation capability between the lesion edge and the reflection area. A time grid migration and energy reflection suppression mechanism is introduced. By synchronously adjusting the light source start-up delay and acquisition time, the incident cone-shaped light field is reconstructed to weaken the reflected echo energy. The dynamic response evaluation mechanism is activated to analyze the texture clarity and brightness balance of the lesion edge area in real time under the current light field configuration, and to complete the adaptive optimization and adjustment of the light field.

10. The method for early laryngeal cancer grading based on deep learning of laryngoscopy images according to claim 9, characterized in that, The phase conjugate modulation strategy introduces a phase adjustment unit at the light source emission end to set the phase of the emitted light wave to the conjugate state of the reflected wave of the previous cycle, so that the interference effect forms a local energy cancellation region on the tissue surface, which is used to weaken the specular reflection high brightness response.