An adaptive multi-photon endoscopy assisted diagnosis method and system based on multi-modal image processing
By using multiphoton endoscopy to acquire and screen consistent regions of interest in real time, image quality is optimized, solving the problem of unstable imaging quality in real-time endoscopic diagnosis of colorectal tumors, and achieving efficient and stable tumor invasion identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIMEI UNIV
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-24
AI Technical Summary
In the real-time endoscopic diagnosis of colorectal tumors, the existing technology suffers from unstable imaging quality, which leads to inconsistencies in image feature extraction and multimodal fusion, affecting the real-time performance and accuracy of the diagnosis.
Multimodal photon imaging signals are acquired in real time using a multiphoton endoscope. Consistent regions of interest are selected, an image quality evaluation function is constructed, parameters are optimized to enhance image quality, and the results are input into an auxiliary diagnostic model for detection.
It improves the stability and efficiency of colorectal tumor invasion identification, meets the needs of real-time diagnosis, reduces the impact of noise and artifacts, and enhances the effectiveness of multimodal fusion.
Smart Images

Figure CN122115459B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multiphoton endoscopy-assisted diagnostic technology, and in particular to an adaptive multiphoton endoscopy-assisted diagnostic method and system based on multimodal image processing. Background Technology
[0002] Colorectal tumors are among the most common malignant tumors of the digestive system. Clinically, the diagnosis and assessment of the extent of tumor invasion in colorectal tumors usually rely on pathological examination, especially during endoscopic examinations. Pathological diagnosis still considers the preparation and staining of biopsy tissue (e.g., H&E, Masson, etc.) followed by microscopic interpretation by a pathologist as the "gold standard." However, traditional pathological sample preparation typically involves multiple steps such as fixation, dehydration, paraffin embedding, sectioning, and staining, making the entire process lengthy and difficult to meet the needs of rapid intraoperative diagnosis and real-time decision-making.
[0003] To improve diagnostic efficiency, various image-based computational aid diagnostic systems have emerged in recent years. For example, feature extraction from pathological slide images or images obtained through endoscopy / microscopy can enable lesion identification, boundary localization, or pathological grading prediction. Meanwhile, multimodal fusion methods, by simultaneously utilizing information from different imaging channels or different viewpoints / modalities, can improve diagnostic performance and robustness to some extent. However, existing image-based aid diagnostic methods still face significant challenges in transitioning from offline acquisition to real-time in vivo applications.
[0004] In in vivo real-time imaging scenarios, image quality is often no longer stable and controllable, but rather coupled with real-time constraints and fluctuates continuously. Flexible endoscopes operating within the cavity are inevitably affected by factors such as intracavitary movement, respiration / peristalsis, changes in the relative position of instruments, field-of-view obstruction, and focal plane stability. Simultaneously, to meet the requirements of "real-time acquisition and rapid output," the system needs to make trade-offs between exposure, scanning speed, and sampling frequency. These factors collectively lead to phenomena such as motion artifacts, local blurring, focus / depth-of-field fluctuations, and incomplete multi-channel information in the imaging results. Furthermore, when quality deteriorates and is accompanied by missing local information or multimodal spatial misalignment, diagnostic algorithms based on image feature extraction and cross-modal alignment / fusion become unstable: on the one hand, noise and artifacts may interfere with texture and structural representation, causing feature distribution drift; on the other hand, inconsistencies in multimodal information prevent the registration and fusion processes from reliably converging under the consistency assumption, thus causing significant fluctuations in diagnostic output with the input quality state, and even leading to misinterpretations.
[0005] Existing enhancement, denoising, registration, or fusion methods for improving image visibility typically assume that the input image is "basically usable" within a certain quality range, or primarily serve offline / post-processing scenarios. Under real-time imaging conditions, because image quality can change rapidly over time, and local areas may experience more severe quality degradation, these methods often struggle to guarantee that the enhancement / fusion results retain reliable structural information relevant to the diagnostic task. In other words, existing enhancement / fusion strategies lack task-oriented constraints on changes in "quality availability," making it difficult to achieve a stable improvement consistent with diagnostic reliability under real-time requirements.
[0006] The purpose of this invention is to design an adaptive multiphoton endoscopic auxiliary diagnostic method and system based on multimodal image processing to address the problems existing in the prior art. Summary of the Invention
[0007] In view of this, the purpose of this invention is to propose an adaptive multiphoton endoscopy-assisted diagnostic method and system based on multimodal image processing, which can solve the above-mentioned problems.
[0008] This invention provides an adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing, comprising: S1 acquires multimodal photon imaging signals in real time through a multiphoton endoscope, preprocesses each modal signal, and filters candidate regions of interest based on image brightness. A consistent region of interest is obtained by filtering the multimodal photon imaging signals of the candidate regions of interest according to the consistency index of the multimodal photon imaging signals of the candidate regions of interest. S2 extracts image features for each consistent region of interest, constructs an image quality evaluation function based on real-time analyzable metrics of the image features, and calculates the image quality score for each consistent region of interest. S3 determines image optimization parameters based on the image features and image quality score of each consistent region of interest, under the conditions of satisfying the constraints of image real-time performance and information integrity, so as to maximize the quality gain of the fused image and obtain the optimization strategy for each consistent region of interest. S4 uses a consistent region of interest optimization strategy to select regions of interest to be optimized, applies the optimization strategy to the regions of interest to be optimized to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.
[0009] The present invention further provides an adaptive multiphoton endoscopy-assisted diagnostic system based on multimodal image processing, comprising: The filtering module is used to acquire multimodal photon imaging signals in real time through a multiphoton endoscope, preprocess each modal signal, filter candidate regions of interest based on image brightness based on the multimodal photon imaging signals, and obtain consistent regions of interest based on the consistency index of the multimodal photon imaging signals of the candidate regions of interest. The scoring module is used to extract image features for each consistent region of interest, construct an image quality evaluation function based on real-time analyzable indicators of image features, and calculate the image quality score for each consistent region of interest. The strategy module is used to determine image optimization parameters based on the image features and image quality score of each consistent region of interest, under the constraints of image real-time performance and information integrity, so as to maximize the quality gain of the fused image and obtain the optimization strategy for each consistent region of interest. The execution module filters regions of interest to be optimized based on a consistent region of interest optimization strategy, applies the optimization strategy to the regions of interest to be optimized to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.
[0010] The beneficial effects of this invention are: First, by reducing the entire field of view data to a consistent region of interest (ROI), the amount of data that needs to be processed for subsequent enhancement, quality assessment, and model inference is reduced, thereby improving the system frame rate and meeting real-time requirements. Only candidate regions related to tumor invasion characteristics are retained, suppressing the influence of background tissue, out-of-field regions, and imaging artifacts on the recognition results, thus improving the stability of the invasion probability output. The definition of a consistent ROI allows different modalities (fluorescence lifetime / second harmonic / third harmonic) to be analyzed and optimized at the same spatial location, enhancing the effectiveness of multimodal fusion.
[0011] Secondly, a quality evaluation index is constructed using signal-to-noise ratio and gradient statistics, enabling the system to quickly determine the reliability and structural clarity of images in real-time. Image quality scores and subsequent quality gains are used to quantify the net benefit of enhancement strategies, avoiding performance instability caused by relying solely on subjective or single-modal strength parameter selection.
[0012] Third, by focusing on image quality gain and selecting optimal parameters while ensuring real-time execution, the enhanced image is made more effective for invasiveness identification, rather than simply becoming clearer. Real-time constraints prevent recognition delays caused by timeouts; information integrity constraints ensure that the statistics of key fluorescence lifetime parameters are not corrupted, thereby reducing the risk of visual enhancement but loss of diagnostic information. Joint optimization of sampling distribution ratio and fusion weights enables the system to adapt to imaging conditions and lesion characteristics of different ROIs, improving the contribution of multimodal fusion to the determination of invasiveness severity.
[0013] Fourth, by combining quality score threshold preservation and Top-K image gain ranking for joint screening, enhancement is performed only on ROIs most likely to bring diagnostic gains, significantly improving real-time efficiency. The screening logic ensures that ROIs entering enhancement and inference have sufficient basic quality and that there is a clear net improvement after enhancement, reducing the impact of erroneous inputs on the output of the auxiliary diagnostic model from the source. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart of the method in this embodiment. Detailed Implementation
[0016] To facilitate understanding by those skilled in the art, the structure of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that, unless otherwise specified, the order of the steps mentioned in this embodiment can be adjusted according to actual needs, and they can even be executed simultaneously or partially simultaneously.
[0017] like Figure 1 As shown, this embodiment of the invention provides an adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing, comprising: S1 acquires multimodal photon imaging signals in real time through a multiphoton endoscope, preprocesses each modal signal, and filters candidate regions of interest based on image brightness. A consistent region of interest is obtained by filtering the multimodal photon imaging signals of the candidate regions of interest according to the consistency index of the multimodal photon imaging signals of the candidate regions of interest. In this step, multimodal acquisition yields large-scale image data pixel-by-pixel / frame-by-frame. Directly performing enhancement, feature extraction, and diagnostic prediction on the entire image would result in excessive computation and increased inference latency, making it difficult to meet real-time requirements. This approach quickly identifies consistent regions of interest, reducing the interference of noise and false positives on prediction, while simultaneously improving the accuracy and stability of intrusion level assessment while ensuring the system's real-time acquisition and recognition capabilities.
[0018] S101 acquires two-photon induced fluorescence lifetime imaging signals and their time-domain parameters, second harmonic imaging signals, and third harmonic imaging signals of colorectal tissue in real time. The signals are then filtered, background corrected, and intensity normalized to obtain preprocessed fluorescence lifetime images, second harmonic images, and third harmonic images. In this step, fluorescence lifetime images reflect changes in the fluorescent molecular environment and biochemical state within the tissue, providing temporal information related to the pathological stage. Second harmonic imaging is highly sensitive to non-centrosymmetric structures in tissues (such as fibrous structures and collagen-related microstructures), characterizing structural arrangement and morphological differences. Third harmonic imaging is more sensitive to microstructural changes such as discontinuities or gradients in the refractive index of the medium, supplementing spatial texture information related to lesions. By simultaneously acquiring and utilizing the complementary information of these three modalities, cross-modal characterization can be formed at both the molecular level (lifetime temporal characteristics) and the microstructural level (SHG and THG texture / structural characteristics), thereby improving the accuracy and robustness of candidate region screening and subsequent diagnostic prediction.
[0019] S102 performs real-time synchronization and spatial matching processing on the preprocessed fluorescence lifetime image, second harmonic image, and third harmonic image, so that the same pixel coordinate corresponds to the same local tissue region in the three modes, resulting in a spatially aligned multimodal image set. In this step, within the same frame acquisition, slight delays and registration errors exist between modalities due to the system's scanning speed. Real-time synchronization and spatial matching in S102 ensure that the ROI maintains the same geometric position on both the lifetime image and the THG image.
[0020] S103 acquires the signal intensity of each pixel in the second harmonic image as its brightness. When the brightness of a pixel is not less than a preset brightness threshold, the pixel is marked as a bright pixel, and a brightness mask is generated. S104 performs connected component analysis on the luminance mask to obtain multiple connected bright regions. Bright regions with a pixel count not less than a preset pixel count threshold are identified as candidate regions of interest.
[0021] In this step, the sensitivity of SHG to tissue structures (such as fibrous structures / non-centrosymmetric structural contrast) is utilized to quickly and in real-time extract candidate regions. A certain lesion area in the colorectal mucosa exhibits stronger structural contrast on SHG. After setting a brightness threshold, the pixels corresponding to these regions are marked as "bright pixels" and thus included in subsequent candidate region extraction.
[0022] Occasional individual bright speckles in the image will form minimal connected components after thresholding. If their pixel count is below a preset threshold, they are excluded to avoid treating "noise points" as candidate ROIs for consistency calculation.
[0023] S105 extracts the temporal features of the fluorescence lifetime image of each candidate region of interest, extracts the texture features of the third harmonic image of the candidate region of interest, analyzes the imaging consistency of the temporal features and texture features, and identifies the candidate regions of interest whose imaging consistency meets the imaging consistency threshold as consistent regions of interest.
[0024] S1041 extracts the first lifetime component and the second lifetime component of the fluorescence lifetime image of each candidate region of interest as its temporal features. S1042 performs a texture operator on the third harmonic image of each candidate region of interest to extract its texture features; S1043 inputs the temporal and texture features into the feature encoding module of the auxiliary diagnostic model to obtain temporal embedding features and texture embedding features respectively. The imaging consistency between the two is calculated by cosine similarity. If the consistency exceeds the consistency threshold, the current candidate region of interest is taken as the consistent region of interest.
[0025] In this step, changes in the fluorescent molecular environment within the lesion tissue lead to different component proportions or values in lifetime decay. Representing these as first / second lifetime components can form a stability criterion in cross-modal consistency assessment.
[0026] Some lesion areas exhibit richer texture details or more pronounced spatial structure changes in THG images. Statistical measures extracted by texture operators can better distinguish these areas.
[0027] When a candidate ROI belongs to the same pathological state in the lifespan image and the THG image, its encoded temporal embedding and texture embedding have high directional consistency in the unified space, and the cosine similarity will exceed the threshold, thus being judged as a consistent region of interest; otherwise, it will be eliminated.
[0028] S2 extracts image features for each consistent region of interest, constructs an image quality evaluation function based on real-time analyzable metrics of the image features, and calculates the image quality score for each consistent region of interest. For each region of interest, S201 extracts pixel values from the fluorescence lifetime image and calculates the signal-to-noise ratio based on the mean of the region signal and the standard deviation of the noise estimated from the preset background region. S202 calculates the gradient magnitude on the second harmonic image for each region of interest and performs weighted statistics within the region of interest to obtain the second harmonic gradient statistics. S203 calculates the gradient magnitude on the third harmonic image for each region of interest and performs weighted statistics within the region of interest to obtain the third harmonic gradient statistics. For each region of interest, S204 uses preset weights to weight and fuse the signal-to-noise ratio, second harmonic gradient statistics, and third harmonic gradient statistics to obtain an image quality score.
[0029] In this step, due to the occurrence of local occlusion, intensity fluctuations, noise changes, and inconsistent imaging conditions of different modalities during real-time acquisition, directly applying the same strategy to all consistent regions of interest may result in inconsistent image quality after enhancement, thereby reducing the stability and consistency of the invasion probability output.
[0030] Signal-to-noise ratio (SNR) estimation is performed on the fluorescence lifetime modes of each region of consistent interest (ROI) to reflect whether lifetime-related temporal information is dominated by noise and whether it possesses stable lifetime characteristic representation capabilities. For example, when air bubbles or instruments obstruct the colorectal mucosa, the fluorescence lifetime signal in that region may experience increased noise and a shift in lifetime component estimation. The calculated SNR will decrease, resulting in a corresponding reduction in the weight of that ROI in subsequent scoring, thus avoiding the use of regions with "unreliable lifetime information" for enhancement and invasion prediction.
[0031] Gradient statistics for SHG (Self-Enhancing Area) are used to characterize the clarity of structural boundaries and the richness of texture detail in the region, thus representing the amount of available information in the SHG modality. When there is slight defocus or change in the angle of the tissue surface in the field of view, the brightness of the SHG may still be high, but the edge transitions become smoother (lower gradient magnitude). Gradient statistics can capture situations of "unclear texture / structure," thus reflecting in the quality score that the ROI is not an ideal enhancement object.
[0032] The THG gradient statistic is used to characterize the clarity of structural information related to texture / refractive index discontinuities in the THG mode for a region of interest (ROI). THG is sensitive to microstructural changes, but texture details deteriorate under noise or poor acquisition conditions. For example, in the same lesion's adjacent region, changes in tissue thickness can lead to THG detail attenuation. In such cases, even if the ROI performs well on SHG, the THG gradient statistic may still be low.
[0033] S3 determines image optimization parameters based on the image features and image quality score of each consistent region of interest, under the conditions of satisfying the constraints of image real-time performance and information integrity, so as to maximize the quality gain of the fused image and obtain the optimization strategy for each consistent region of interest. In this step, during real-time acquisition using multiphoton endoscopy, the "sampling allocation ratio" and "fusion weight" of different modalities within the ROI are jointly optimized. This ensures that the fused image achieves maximum quality gain within a short time budget (meeting real-time requirements) (beneficial for subsequent auxiliary diagnosis / invasiveness identification), while guaranteeing that key biological information (such as statistical fidelity of lifespan key parameters) is not lost due to over-optimization. The specific steps are as follows: S301 performs joint optimization modeling on regions of consistent interest, defining a set of variables for joint optimization, including: the sampling distribution ratio of different modalities and the multimodal fusion weights; In this step, during real-time scanning, lifetime modalities are highly sensitive to sampling duration; a fixed sampling ratio may lead to excessively large lifetime estimation variance. Furthermore, SHG / THG texture information is more sensitive to strong noise. By simultaneously incorporating both the "sampling distribution ratio (determining the allocation of sampling resources for each modality)" and the "fusion weight (determining the contribution of each modality's information in fusion)" into the variable set, a combined strategy that balances lifetime fidelity and overall quality improvement can be found at the ROI level, thereby supporting real-time intrusion detection.
[0034] In the joint optimization, S302 performs sampling and multimodal fusion on the consistent interest region based on the sampling distribution ratio of different modalities and the multimodal fusion weight to obtain a multimodal fused image, and calculates the image quality gain of the multimodal fused image; S3021 obtains the sampling results corresponding to the fluorescence lifetime image, second harmonic image and third harmonic image by sampling distribution ratio of multiple different modes, and fuses the sampling results based on multimodal fusion weight to obtain a multimodal fused image; S3022 extracts pixel values from the multimodal fused image and calculates the fused signal-to-noise ratio based on the mean of the regional signal and the noise standard deviation estimated from the preset background region; S3023 calculates the gradient magnitude on the multimodal fused image and performs weighted statistics to obtain the fused gradient statistics; S3024 calculates the fused image quality score by weighting the fused signal-to-noise ratio and the fused gradient statistics, and uses the difference between the fused image quality score and the original image quality score as the image quality gain.
[0035] In this step, during real-time acquisition and identification of tumor invasion, it is necessary to quickly determine whether the "current candidate sampling distribution ratio and fusion weight" truly improve the effective information content of the input image for the identification task. Therefore, by comparing the image quality before and after fusion, the difference between the fused image quality score and the original image quality score is calculated as the image quality gain. This gain is used to characterize the "net benefit" brought by the selected parameters, thus providing a directly comparable and real-time evaluable target quantity for subsequent constraint optimization.
[0036] In the joint optimization of regions of consistent interest, S303 applies real-time constraints and information integrity constraints, and uses image quality gain as the objective function to solve for the optimal sampling distribution ratio of different modalities and the multimodal fusion weights.
[0037] S3031 uses the duration of the fused image being lower than the duration threshold and the number of samples being lower than the sampling threshold as real-time constraints; S3032 uses the requirement that the statistical fidelity of key lifespan parameters be no less than their threshold as an information integrity constraint. S3033 aims to maximize image quality gain and solves for the optimal multimodal fusion weights and sampling distribution ratio.
[0038] In this step, in the application scenario of real-time acquisition and real-time identification of tumor invasion degree, the image optimization strategy needs to simultaneously meet the following constraints: real-time constraint, that is, the generation time of fused images and the number of samplings must be limited to ensure that the system outputs the invasion probability results within a limited time window; and information integrity constraint, that is, the statistical fidelity proxy of key life parameters needs to be kept above a preset threshold to avoid the weakening or distortion of key life features during sampling and fusion.
[0039] However, increasing the number of samples or extending the fusion calculation during optimization to improve image quality may result in a higher signal-to-noise ratio or clearer structural details, but it can cause the fused image generation time to exceed the real-time budget. Conversely, compressing the number of samples or shortening the fusion process to meet real-time requirements may reduce the statistical stability of fluorescence lifetime, causing the statistical fidelity surrogate of key lifetime parameters to fall below the threshold, thereby compromising information integrity. Therefore, there is an inherent conflict and trade-off between real-time constraints and information integrity constraints as optimization variables change.
[0040] Among all candidate parameter combinations that satisfy the constraints, the combination that maximizes the image quality gain is selected as the optimal solution. This constraint-optimal solution method avoids selecting parameter combinations that are not real-time or contain incomplete information, while ensuring that the quality improvement is as large as possible while maintaining real-time performance and lifetime information fidelity. This improves the accuracy and stability of subsequent real-time intrusion detection.
[0041] S4 uses a consistent region of interest (ROI) optimization strategy to select ROIs to be optimized, applies the optimization strategy to the ROIs to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.
[0042] S401 performs joint filtering of consistent regions of interest based on the image quality score and image quality gain of each consistent region of interest to obtain the regions of interest to be optimized; S4011 retains the consistent regions of interest whose image quality scores are higher than the image quality threshold, thus obtaining a set of consistent regions of interest for high-quality images. S4012 sorts each consistent region of interest in the set of consistent regions of interest for high-quality images from high to low according to image gain. S4013 selects the top K regions of interest from the sorted set as regions of interest to be optimized.
[0043] In this step, under real-time acquisition conditions, some Regions of Interest (ROIs) may suffer from occlusion, scattering artifacts, or excessively large variance in lifetime estimation. Resampling, redistributing, and fusing these ROIs may result in a "smoother but informationally flawed" outcome, causing the invasion probability output by the auxiliary diagnostic model to be based on incorrect representations. By setting a threshold "above the image quality threshold," these ROIs can be directly removed, ensuring the stability of subsequent real-time invasion probability output.
[0044] It is impossible to perform optimal sampling and fusion simultaneously on all high-quality ROIs. The Top-K strategy ensures diagnostic effectiveness while meeting real-time requirements. Some ROIs may have high quality scores but low gain, and further optimization may have limited contribution; sorting by gain can further exclude this situation.
[0045] For each region of interest to be optimized, S402 executes the corresponding optimal sampling distribution ratio of different modalities and multimodal fusion weights to obtain the enhanced image; S403 will enhance the image input to the pre-trained auxiliary diagnostic model to obtain the corresponding probability of colorectal tumor invasion.
[0046] In this step, the pre-trained auxiliary diagnostic model is trained offline using labeled data (e.g., colorectal tumor invasion level / probability related labels annotated by pathological results). This model learns to extract features related to the degree of invasion (e.g., lifespan key parameters representing boundary / texture differences in structural modalities) from the input multimodal enhanced image, thereby outputting the invasion probability or invasion level result for the corresponding ROI or frame. The auxiliary diagnostic model can employ an image classification network or a multimodal fusion network to predict the degree of tumor invasion. The auxiliary diagnostic model includes at least a feature extraction module and a prediction module. The feature extraction module can employ residual networks such as ResNet or visual Transformer; the prediction module outputs the probability or level of tumor invasion.
[0047] This invention provides an adaptive multiphoton endoscopy-assisted diagnostic system based on multimodal image processing, comprising: The filtering module is used to acquire multimodal photon imaging signals in real time through a multiphoton endoscope, preprocess each modal signal, filter candidate regions of interest based on image brightness, and obtain consistent regions of interest based on the consistency index of the multimodal photon imaging signals of the candidate regions of interest. The scoring module is used to extract image features for each consistent region of interest, construct an image quality evaluation function based on real-time analyzable indicators of image features, and calculate the image quality score for each consistent region of interest. The strategy module is used to determine image optimization parameters based on the image features and image quality score of each consistent region of interest, under the constraints of image real-time performance and information integrity, so as to maximize the quality gain of the fused image and obtain the optimization strategy for each consistent region of interest. The execution module filters regions of interest to be optimized based on a consistent region of interest optimization strategy, applies the optimization strategy to the regions of interest to be optimized to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.
[0048] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0049] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0050] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0051] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.
[0052] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The words first, second, and third, etc., do not indicate any order. These words can be interpreted as names.
[0053] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0054] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0055] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0056] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
Claims
1. An adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing, characterized in that, include: S1 acquires multimodal photon imaging signals in real time through a multiphoton endoscope, preprocesses each modal signal, and filters candidate regions of interest based on image brightness. A consistent region of interest is obtained by filtering the multimodal photon imaging signals of the candidate regions of interest according to the consistency index of the multimodal photon imaging signals of the candidate regions of interest. S2 extracts image features for each consistent region of interest (ROI), constructs an image quality evaluation function based on real-time analyzable metrics of these features, and calculates an image quality score for each ROI, including: For each region of interest, S201 extracts pixel values from the fluorescence lifetime image and calculates the signal-to-noise ratio based on the mean of the region signal and the standard deviation of the noise estimated from the preset background region. S202 calculates the gradient magnitude on the second harmonic image for each region of interest and performs weighted statistics within the region of interest to obtain the second harmonic gradient statistics. S203 calculates the gradient magnitude on the third harmonic image for each region of interest and performs weighted statistics within the region of interest to obtain the third harmonic gradient statistics. S204 uses preset weights to weight and fuse the signal-to-noise ratio, second harmonic gradient statistics, and third harmonic gradient statistics for each consistent region of interest to obtain an image quality score. Based on the image features and image quality scores of each consistent region of interest (ROI), S3 determines image optimization parameters to maximize the quality gain of the fused image while satisfying constraints on image real-time performance and information integrity. This yields the optimization strategy for each ROI, including: S301 performs joint optimization modeling on regions of consistent interest, defining a set of variables for joint optimization, including: the sampling distribution ratio of different modalities and the multimodal fusion weights; In the joint optimization, S302 performs sampling and multimodal fusion on the consistent interest region based on the sampling distribution ratio of different modalities and the multimodal fusion weight to obtain a multimodal fused image, and calculates the image quality gain of the multimodal fused image; In the joint optimization of regions of consistent interest, S303 applies real-time constraints and information integrity constraints, and uses image quality gain as the objective function to solve for the optimal sampling distribution ratio of different modalities and multimodal fusion weights. S4 uses a consistent region of interest (ROI) optimization strategy to select ROIs to be optimized, applies the optimization strategy to the ROIs to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.
2. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 1, characterized in that, The process involves real-time acquisition of multimodal photon imaging signals via a multiphoton endoscope, preprocessing each modal signal, and selecting candidate regions of interest (ROIs) based on image brightness. Consistent ROIs are then selected based on the consistency index of the multimodal photon imaging signals from the candidate ROIs, including: S101 acquires two-photon induced fluorescence lifetime imaging signals and their time-domain parameters, second harmonic imaging signals, and third harmonic imaging signals of colorectal tissue in real time. The signals are then filtered, background corrected, and intensity normalized to obtain preprocessed fluorescence lifetime images, second harmonic images, and third harmonic images. S102 performs real-time synchronization and spatial matching processing on the preprocessed fluorescence lifetime image, second harmonic image, and third harmonic image, so that the same pixel coordinate corresponds to the same local tissue region in the three modes, resulting in a spatially aligned multimodal image set. S103 acquires the signal intensity of each pixel in the second harmonic image as its brightness. When the brightness of a pixel is not less than a preset brightness threshold, the pixel is marked as a bright pixel, and a brightness mask is generated. S104 Performs connected component analysis on the brightness mask to obtain multiple connected bright regions. Bright regions with a pixel count not less than a preset pixel count threshold are identified as candidate regions of interest. S105 Extracts the temporal features of the fluorescence lifetime image of each candidate region of interest, extracts the texture features of the third harmonic image of the candidate region of interest, analyzes the imaging consistency of the temporal features and texture features, and identifies candidate regions of interest whose imaging consistency meets the imaging consistency threshold as consistent regions of interest.
3. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 2, characterized in that, The process of extracting the temporal features of the fluorescence lifetime image of each candidate region of interest, extracting the texture features of the third harmonic image of the candidate region of interest, analyzing the imaging consistency of the temporal features and texture features, and identifying candidate regions of interest whose imaging consistency meets the imaging consistency threshold as consistent regions of interest includes: S1041 extracts the first lifetime component and the second lifetime component of the fluorescence lifetime image of each candidate region of interest as its temporal features. S1042 performs a texture operator on the third harmonic image of each candidate region of interest to extract its texture features; S1043 inputs the temporal and texture features into the feature encoding module of the auxiliary diagnostic model to obtain temporal embedding features and texture embedding features respectively. The imaging consistency between the two is calculated by cosine similarity. If the consistency exceeds the consistency threshold, the current candidate region of interest is taken as the consistent region of interest.
4. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 1, characterized in that, In the joint optimization, sampling and multimodal fusion are performed on the consistent interest region based on the sampling distribution ratio of different modalities and the multimodal fusion weight to obtain a multimodal fused image. The image quality gain of the multimodal fused image is calculated as follows: S3021 obtains the sampling results corresponding to the fluorescence lifetime image, second harmonic image and third harmonic image by sampling distribution ratio of multiple different modes, and fuses the sampling results based on multimodal fusion weight to obtain a multimodal fused image; S3022 extracts pixel values from the multimodal fused image and calculates the fused signal-to-noise ratio based on the mean of the regional signal and the noise standard deviation estimated from the preset background region; S3023 calculates the gradient magnitude on the multimodal fused image and performs weighted statistics to obtain the fused gradient statistics; S3024 calculates the fused image quality score by weighting the fused signal-to-noise ratio and the fused gradient statistics, and uses the difference between the fused image quality score and the original image quality score as the image quality gain.
5. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 1, characterized in that, In the joint optimization within the region of consistent interest, real-time constraints and information integrity constraints are applied, and image quality gain is used as the objective function to solve for the optimal sampling distribution ratio of different modalities and the multimodal fusion weights, including: S3031 uses the duration of the fused image being lower than the duration threshold and the number of samples being lower than the sampling threshold as real-time constraints; S3032 uses the requirement that the statistical fidelity of key lifespan parameters be no less than their threshold as an information integrity constraint. S3033 aims to maximize image quality gain and solves for the optimal multimodal fusion weights and sampling distribution ratio.
6. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 1, characterized in that, The optimization strategy based on consistent regions of interest (ROIs) filters ROIs to be optimized, applies the optimization strategy to the ROIs to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results, including: S401 performs joint filtering of consistent regions of interest based on the image quality score and image quality gain of each consistent region of interest to obtain the regions of interest to be optimized; For each region of interest to be optimized, S402 executes the corresponding optimal sampling distribution ratio of different modalities and multimodal fusion weights to obtain the enhanced image; S403 will enhance the image input to the pre-trained auxiliary diagnostic model to obtain the corresponding probability of colorectal tumor invasion.
7. The adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing according to claim 6, characterized in that, The joint screening of consistent regions of interest (ROIs) based on the image quality score and image quality gain of each ROI yields the following ROIs to be optimized: S4011 retains the consistent regions of interest whose image quality scores are higher than the image quality threshold, thus obtaining a set of consistent regions of interest for high-quality images. S4012 sorts each consistent region of interest in the set of consistent regions of interest for high-quality images from high to low according to image gain. S4013 selects the top K regions of interest from the sorted set as regions of interest to be optimized.
8. An adaptive multiphoton endoscopic-assisted diagnostic system based on multimodal image processing, characterized in that, An adaptive multiphoton endoscopy-assisted diagnostic method based on multimodal image processing, according to any one of claims 1-7, includes: The filtering module is used to acquire multimodal photon imaging signals in real time through a multiphoton endoscope, preprocess each modal signal, filter candidate regions of interest based on image brightness, and obtain consistent regions of interest based on the consistency index of the multimodal photon imaging signals of the candidate regions of interest. The scoring module is used to extract image features for each consistent region of interest, construct an image quality evaluation function based on real-time analyzable indicators of image features, and calculate the image quality score for each consistent region of interest. The strategy module is used to determine image optimization parameters based on the image features and image quality score of each consistent region of interest, under the constraints of image real-time performance and information integrity, so as to maximize the quality gain of the fused image and obtain the optimization strategy for each consistent region of interest. The execution module filters regions of interest to be optimized based on a consistent region of interest optimization strategy, applies the optimization strategy to the regions of interest to be optimized to obtain an enhanced image, inputs the enhanced image into the auxiliary diagnostic model for detection, and outputs the detection results.