Digestive endoscopy decontamination quality evaluation method based on image recognition
By combining a dual-stream backbone network and a causal graph structure, the robustness and interpretability of the digestive endoscopy image quality assessment model under low-quality image conditions are solved. The model achieves dynamic causal relationship modeling and misjudgment mitigation between image quality and washout status, thereby improving the model's performance in real clinical scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN THE FIRST HOSPITAL
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing models for assessing the quality of digestive endoscopy images lack robustness and interpretability under low-quality image conditions. In particular, under real degradation conditions such as atomization, reflection, and motion blur, the false positive rate for residue identification is high, and they cannot effectively distinguish and model the complex implicit dependency between image quality and washing status.
Feature extraction is performed using a dual-stream backbone network based on shared convolutional parameters. An implicit causal graph structure is constructed by combining the LINGAM causal discovery algorithm and gradient masking mechanism. A causal intervention feature reweighting mechanism is used to perform virtual sharpness enhancement in low-quality image regions and monitor the sensitivity of output changes. A region confidence decay factor is generated, and a closed-loop optimized multi-task loss function is constructed to achieve end-to-end quality assessment.
It significantly enhances the model's ability to identify and suppress interference factors under low-quality image conditions, improves its robustness and reliability in real clinical scenarios, and effectively alleviates the problem of misjudgment caused by poor imaging conditions.
Smart Images

Figure CN122049477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing and intelligent assessment technology of digestive endoscopy, and in particular to a method for assessing the quality of digestive endoscopy cleaning and disinfection based on image recognition. Background Technology
[0002] With the rapid development of digital healthcare and intelligent assisted diagnosis and treatment, the automated assessment of the quality of endoscopy cleaning and disinfection is receiving increasing attention from medical institutions. Currently, in the field of endoscopy cleaning and disinfection quality assessment, mainstream technical solutions mostly rely on deep learning models based on image recognition. These models typically implement sub-tasks such as endoscopy image quality judgment and cleaning status (mainly residue identification) through single-task or simple multi-task learning frameworks. Mainstream algorithms generally employ convolutional neural networks (CNNs), image classification and segmentation techniques, attention mechanisms, and a small number of feature fusion strategies in model design, and obtain the discriminative model through supervised learning on the acquired clinical endoscopy image set.
[0003] Currently, models applied to image quality assessment and residual detection in digestive endoscopy often employ simple shared feature extraction structures or introduce attention-weighted fusion modules between task branches. While these models perform well in high-definition, low-noise testing scenarios, they often fail in real-world clinical situations with fluctuating image quality and real-world degradation such as fogging, reflection, and motion blur. This lack of effective differentiation and modeling of the complex, implicit dependencies between image quality and disinfection status leads to increased false-positive rates in residual identification under low-quality images, resulting in insufficient robustness and interpretability. For example, while some researchers have proposed multi-task convolutional neural networks that share low-level features to improve parameter utilization, simple task head separation is insufficient to address the systematic misleading effect of low-quality regions on disinfection status assessment. At the basic research level, some works have attempted to dynamically adjust the weights of different task branches using attention mechanisms, but these neglect modeling of the causal influence mechanism between tasks, only alleviating some feature confusion issues and failing to fundamentally address the root cause of image quality degradation misleading disinfection judgments. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, the present invention provides a method for evaluating the quality of digestive endoscope washing and disinfection based on image recognition.
[0005] The technical solution of this invention is implemented as follows: a method for assessing the quality of washing and disinfection of digestive endoscopes based on image recognition, comprising: S1: Perform dual-stream input processing on the original digestive endoscopy image and the degraded image enhanced by physical imaging simulation, wherein the physical imaging simulation enhancement includes simulating the degradation process of fogging, reflection and motion blur to generate a dual-path input dataset; S2: A dual-stream backbone network based on shared convolutional parameters is used to extract features from the dual-path input dataset. The bottom layer shares some convolutional layer parameters to maintain the consistency of basic features, while the high layer retains the ability to express differences and outputs multi-scale feature maps in the shared feature space. S3: Based on the LINGAM causal discovery algorithm and gradient masking mechanism, a trainable implicit causal graph structure is constructed, in which the potential causal direction between the image quality factor and the washout state judgment result is inferred through variational graph learning, and a dynamically optimized causal adjacency matrix is generated. S4: Based on the implicit causal graph structure, perform causal intervention feature reweighting operation on the shared feature space. When a conflict is detected between a low-quality image region and a residual judgment, the clarity is virtually improved and the sensitivity of output change is monitored to generate a region confidence decay factor to correct the feature weights. S5: Construct a differentiable causal attribution head and generate a residual distribution difference map under counterfactual generation conditions. Use the difference map as an auxiliary supervision signal and perform weighted fusion with the original task loss function to form a closed-loop optimized multi-task loss function. S6: The multi-task deep learning architecture is trained in stages using a progressive causal freezing strategy: the first stage fixes the causal graph structure and optimizes the parameters of each task branch; the second stage jointly fine-tunes the causal graph and all network parameters until the model converges. S7: Based on the trained model, perform end-to-end quality assessment on newly acquired digestive endoscopy images, and output a three-dimensional assessment report including image quality factors, washout status judgment results and causal attribution difference map, to achieve dynamic quality closed-loop management in clinical scenarios.
[0006] The present invention provides a method for assessing the washing and disinfection quality of digestive endoscopes based on image recognition, which has the following beneficial effects: This invention introduces an "implicit causal graph decoupling" mechanism to construct a dual-stream parallel deep network architecture with causal reasoning capabilities, significantly enhancing the model's ability to identify and suppress interfering factors under low-quality image conditions. After sharing a high-level feature space, this invention integrates a trainable causal reasoning module, utilizing variational graph learning combined with gradient masking strategies to dynamically mine the potential causal structure between quality factors such as image sharpness, contrast, and noise level and the washout state judgment, identifying key influencing paths such as "image blurring → residual false detection." This achieves a fundamental leap from correlation modeling to causal modeling. Based on this, a causal intervention feature reweighting mechanism is designed to proactively execute counterfactual intervention operations when a high-confidence residual prediction is detected in a low-quality region—that is, virtually improving the image quality of that region in the feature space and observing changes in output sensitivity. If the response is drastic, the original judgment weights are automatically reduced, effectively mitigating the problem of false judgments caused by poor imaging conditions and significantly improving the robustness and reliability of the model in real clinical scenarios. Attached Figure Description
[0007] Figure 1 This is a flowchart of a digestive endoscope washing and disinfection quality assessment method based on image recognition according to the present invention; Figure 2 This is a sub-flowchart of a digestive endoscope washing and disinfection quality assessment method based on image recognition according to the present invention; Figure 3 This is another sub-flowchart of the image recognition-based digestive endoscope washing and disinfection quality assessment method of the present invention. Detailed Implementation
[0008] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0009] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0010] like Figure 1 As shown, this invention provides a method for assessing the quality of washing and disinfection of digestive endoscopes based on image recognition, specifically including: S1: Perform dual-stream input processing on the original digestive endoscopy image and the degraded image enhanced by physical imaging simulation, wherein the physical imaging simulation enhancement includes simulating the degradation process of fogging, reflection and motion blur to generate a dual-path input dataset; S2: A dual-stream backbone network based on shared convolutional parameters is used to extract features from the dual-path input dataset. The bottom layer shares some convolutional layer parameters to maintain the consistency of basic features, while the high layer retains the ability to express differences and outputs multi-scale feature maps in the shared feature space. S3: Based on the LINGAM causal discovery algorithm and gradient masking mechanism, a trainable implicit causal graph structure is constructed, in which the potential causal direction between the image quality factor and the washout state judgment result is inferred through variational graph learning, and a dynamically optimized causal adjacency matrix is generated. S4: Based on the implicit causal graph structure, perform causal intervention feature reweighting operation on the shared feature space. When a conflict is detected between a low-quality image region and a residual judgment, the clarity is virtually improved and the sensitivity of output change is monitored to generate a region confidence decay factor to correct the feature weights. S5: Construct a differentiable causal attribution head and generate a residual distribution difference map under counterfactual generation conditions. Use the difference map as an auxiliary supervision signal and perform weighted fusion with the original task loss function to form a closed-loop optimized multi-task loss function. S6: The multi-task deep learning architecture is trained in stages using a progressive causal freezing strategy: the first stage fixes the causal graph structure and optimizes the parameters of each task branch; the second stage jointly fine-tunes the causal graph and all network parameters until the model converges. S7: Based on the trained model, perform end-to-end quality assessment on newly acquired digestive endoscopy images, and output a three-dimensional assessment report including image quality factors, washout status judgment results and causal attribution difference map, to achieve dynamic quality closed-loop management in clinical scenarios.
[0011] Step S1: Perform dual-stream input processing on the original digestive endoscopy image and the degraded image enhanced by physical imaging simulation, wherein the physical imaging simulation enhancement includes simulating the degradation processes of fogging, reflection, and motion blur to generate a dual-path input dataset. Specifically, this includes: S1.1: Obtain the raw digestive endoscopy image dataset, which is derived from high-resolution endoscopic images acquired clinically. The images contain typical regions of the digestive tract mucosa surface, serving as the input basis for subsequent image quality degradation simulation. Raw image data from clinical high-definition digestive endoscopy examinations are acquired. Color image samples from different locations on the surface of the digestive tract mucosa are collected using a medical high-definition imaging system and saved in a lossless image format to ensure the integrity of information in subsequent processing. The original endoscopic images are retrieved from the hospital's image storage server using the Picture Archiving and Communication System (PACS) interface protocol (parameter: DICOM standard version 3.0) to ensure that the images contain complete acquisition metadata, including information such as equipment model, lighting intensity, and shooting distance. Furthermore, by using the content-based image retrieval (CBIR) method (parameters: feature types are texture features and color histogram features), samples containing typical digestive tract mucosal feature regions are selected from the original image set, and samples with background noise ratios exceeding a preset threshold are removed. Furthermore, an edge detection algorithm (parameters: the high threshold of the Canny operator is set to 0.15 and the low threshold is set to 0.05) is used to perform structured contour analysis on the filtered images to determine the integrity of the tissue region and the clarity of the edges, and to form a region bounding box index. Furthermore, the initial sharpness index of the image is calculated using spatial frequency analysis, employing the formula:
[0012] in For clarity index, The amplitude spectrum of the image in the frequency domain. For frequency domain energy integral, The zero-frequency energy integral is used to determine the spatial information richness of an image; Using the above method, the raw high-definition endoscopic image data obtained in the previous step is transformed into a valid dataset that has been quality-screened and contains regions with diagnostic value, thereby providing high-quality basic input for subsequent image degradation simulation. For example, during a gastroscopy, a high-resolution electronic endoscope with a resolution of 1920×1080 pixels was used to acquire image data. The light source color temperature was 5500K, and the frame rate was 30fps. The image sequence with examination number 20230415007 was retrieved from the hospital's PACS system, and three clear images from the anterior wall of the gastric body, the pylorus, and the duodenal bulb were extracted. Based on the CBIR method, image samples with uniform color histogram distribution and texture complexity above 0.6 were selected, while samples with a boundary noise ratio exceeding 0.1 were removed. The Canny operator (high threshold 0.15, low threshold 0.05) was used to detect complete and unbroken edge contours of gastric mucosal folds. The sharpness index was calculated through spatial frequency analysis, where the frequency domain energy integral was... The zero-frequency energy integral is ,get = The sharpness index indicates that the image has high spatial information quality, and ultimately three valid samples are formed as the input basis for subsequent degradation simulation; S1.2: Based on the physical imaging simulation model, the original digestive endoscopy image is simulated to perform an image degradation process simulation, which includes simulating fogging effect, reflection interference and motion blur, so as to generate low-quality image samples with real clinical degradation characteristics; S1.3: The original digestive endoscopy image and the degraded image enhanced by physical imaging simulation are used as dual-path input data to construct a dual-stream input structure, in which one input is the original image and the other input is the corresponding degraded image, forming a one-to-one dual-stream input pair; The processing results of S1.1 and S1.2 are used as input conditions, including high-resolution original endoscopic images and their degraded image samples enhanced by physical imaging simulation; A path pairing index generation method (parameter: based on the original image set index ID and the degraded image set index ID) is adopted to establish a one-to-one correspondence between the two types of images in the data dimension and generate a dual-path input mapping table; Furthermore, a data synchronization loading algorithm (parameters: batch size of 32, I / O prefetch buffer depth of 4) is used to achieve parallel loading of the corresponding original image and degraded image, and maintain an index-aligned buffer structure to ensure the consistency of the input order of the dual-channel convolutional network. Furthermore, a dual-stream input pipeline construction method is used (parameter: input tensor dimension is...). This enables the placement and data allocation of input tensors for the original image path and the degraded image path, so that subsequent shallow shared convolutions can receive dual-channel data with the same format. Furthermore, a feature alignment preprocessing strategy is adopted (parameter: channel mean alignment threshold is...). Variance normalization range This achieves basic consistency in brightness, contrast, and color space between dual-path inputs, thereby reducing cross-path feature bias in shared convolutional kernels; By constructing the dual-path input structure as described above, the original high-definition image and the degraded simulation image are transformed into structured, index-aligned dual-stream input data with completely consistent tensor dimensions, thereby achieving efficient parallelism and semantic consistency in the subsequent dual-stream convolution feature extraction stage. For example, in a clinical scenario, the original digestive endoscopy image dataset contains 5000 images with a resolution of [missing information]. The high-definition images, after being degraded by the physical imaging simulation model, generate 5000 corresponding images with fogging intensity. Reflective interference ratio Motion fuzzy kernel scale Low-quality images. A dual-path indexing generation method uses a hash function to generate unique mapping IDs, ensuring no index drift during the pairing and loading of 5000 image pairs. The data synchronization loading algorithm sets the batch size to 64 to improve GPU memory utilization. During feature alignment preprocessing, the mean difference between the two paths in the RGB channels is calculated. variance After meeting the preset threshold, the input layer of the two-stream convolutional neural network is entered. After execution, the index consistency of the two-stream input structure is verified to be 100% (no mismatches), and the brightness and color distributions are completely aligned in the normalized space, ensuring the convergence stability of the model in the early stage of training and the consistency of cross-task feature representation. S1.4: Perform image normalization processing on the original image path in the dual-stream input structure, and standardize the brightness, contrast and color distribution of the image based on the image histogram equalization algorithm to obtain a unified image input format; S1.5: Perform image quality parameter annotation on the degraded image path in the dual-stream input structure, extract image quality factors using the image sharpness evaluation model, noise estimation module and fuzzy quantification algorithm, and generate image quality labels corresponding to the degraded image for training supervision of the subsequent causal inference module; For the degraded image path in the dual-stream input structure, the input object is a low-quality digestive endoscopy image generated by physical imaging simulation enhancement, which includes degradation features such as simulated fogging, reflection interference and motion blur. An image sharpness evaluation model (parameter: energy threshold setting based on Laplacian operator) is used to quantify the overall sharpness of degraded images and output a sharpness score matrix as a preliminary quality factor. Furthermore, through the noise estimation module (parameters: based on spatial noise standard deviation measure and frequency domain noise power spectrum analysis), the random noise intensity of the degraded image is quantitatively measured, and a noise level index vector is obtained; Furthermore, a unified scale measurement of the blur degree of degraded images is achieved through a fuzzy quantization algorithm (parameters: gradient magnitude decay rate and frequency domain high-frequency energy loss ratio), and a scalar of blur coefficients is generated. By employing a feature fusion processing method, the sharpness score matrix, noise level index vector, and blur coefficient scalar are weighted and combined in a normalized space to obtain a comprehensive image quality feature vector. The overall quality score is calculated using the following formula, based on a quality label generation algorithm. :
[0013] in, The average of the sharpness scores. This represents the average noise level index. For fuzzy coefficients, , , These are the weight parameters for the corresponding features; Through the above algorithm or processing method, the image quality features of the previous step are transformed into structured image quality labels corresponding to the degraded image, so as to realize the input data required for training supervision of the subsequent causal inference module; For example, in a clinical application scenario, the resolution of the degraded image generated by physical imaging simulation is 1920×1080 pixels. The sharpness assessment sets the Laplacian energy threshold to 150, and the average output sharpness score is 320. In the noise estimation module, the spatial noise standard deviation measure is 2.8, the average frequency domain noise power spectrum is 0.015, and the weighted average noise level index is 0.021. In the fuzzy quantization algorithm, the gradient amplitude decay rate is 0.45, the high-frequency energy loss ratio is 0.37, and the overall fuzzy coefficient is 0.41. Let the weight parameters be... =0.5, =0.3, =0.2, substituting into the comprehensive quality score formula, we get: The overall quality score output is approximately 160.067. This quality score is encoded as a quality label for degraded images and used to train the causal inference module. In model validation, this significantly improves the stability and consistency of residue identification under low-quality image conditions.
[0014] Step S2: Feature extraction is performed on the dual-path input dataset based on a dual-stream backbone network with shared convolutional parameters. The lower layers share some convolutional layer parameters to maintain consistency of basic features, while the higher layers retain differentiated expressive power, outputting multi-scale feature maps in the shared feature space. Specifically, this includes: S2.1: Based on a dual-path input dataset, one path is the original digestive endoscopy image and the other path is the degraded image enhanced by physical imaging simulation, a dual-stream convolutional neural network architecture is constructed to extract the low-level visual semantic features of the two path input data respectively. Based on a dual-path input dataset, one path is the original digestive endoscopy image and the other path is the degraded image enhanced by physical imaging simulation. A dual-stream convolutional neural network initialization method is adopted (parameters: number of input channels = 3, kernel size = 3×3, stride = 1) to realize parallel feature access of image data from the two paths. Furthermore, a spatial domain convolution preprocessing algorithm (parameters: Padding mode = Same, activation function = ReLU) is used to encode the local receptive field of the input image and obtain a feature mapping matrix with basic spatial resolution. Furthermore, a batch normalization method (parameters: batch size = 32, momentum = 0.99) is adopted to standardize the feature distribution across batches, thereby reducing internal covariate bias and generating normalized feature tensors for subsequent shallow feature extraction. Furthermore, the Local Response Normalization (LRN) algorithm (parameters: normalization radius = 5, α = 0.0001, β = 0.75) is introduced to suppress and enhance the response of local feature channels and generate channel feature maps after local optimization. Furthermore, in the shallow structure of each path of the dual-stream network, the parallel extraction of features of different receptive fields is achieved through multi-scale convolution groups (parameters: convolution kernel size combination = {3×3, 5×5}, dilation = {1, 2}), and the features are spliced and fused in the feature channel layer to obtain a multi-scale shallow feature set that adapts to the visual feature conditions of each path. Through the above convolution and normalization processing, the dual-stream input data from the previous step is transformed into a low-level semantic consistency feature map, realizing parallel shallow feature extraction and structured initial feature space construction of the dual-stream network. For example, in the case of inputting high-resolution clinical endoscopic images and degraded images, the original image size is set to 512×512 pixels, in three-channel RGB format; the degraded image maintains the same size and channel configuration after fogging simulation processing. The dual-stream convolutional network input layer receives two types of image data respectively. The 3×3 convolutional kernels of the convolutional group are used to capture minute texture features, and the 5×5 convolutional kernels are used to encode color and structural changes in larger regions; when the dilatation rate is set to 1 and 2, it corresponds to the standard convolution and expanded receptive field modes, respectively. After batch normalization, the variance of the feature value distribution is significantly reduced, and LRN processing improves the response intensity of mucosal texture edges. After parallel computation of the multi-scale convolutional groups, the number of feature channels of the two paths is merged into 64 dimensions. Fusion is achieved through channel concatenation, and the output feature mapping matrix size is 128×128×64, which is used in the subsequent shallow processing stage with shared convolutional parameters. The technical effect is that even under conditions of large differences in image quality, the consistency and comparability of the underlying semantic features can still be maintained. S2.2: In the shallow convolutional layers of the dual-stream backbone network, the convolutional kernel parameters of the two input paths are shared to ensure the semantic consistency of the image content during the basic feature extraction stage and output a low-dimensional feature map in the shared feature space. S2.3: Configure the parameters of the mid-to-high-level convolutional layers of the dual-stream backbone network independently to allow the two input paths to retain their differentiated feature expression capabilities under their respective image quality conditions during the high-level semantic feature extraction stage, thereby generating a high-level semantic feature map with path specificity. S2.4: Based on the multi-scale feature maps in the shared feature space, the feature pyramid network (FPN) structure is used to perform cross-layer fusion processing on the multi-scale feature maps to enhance the spatial resolution and semantic expressive power of the feature maps and generate high-resolution fused feature maps of the same scale. S2.5: Input the fused feature map into the feature normalization module, adjust the channel weights of the fused feature map based on the channel attention mechanism to enhance the feature channel response intensity related to image quality assessment and descaling status recognition, and output the normalized enhanced feature map. S2.6: The normalized enhanced feature map is used as the input to the subsequent causal inference module to support the modeling of the potential causal relationship between the image quality factor and the washout state, and to complete the mapping transformation of the two-stream features to the shared causal feature space.
[0015] like Figure 2 As shown, step S3 involves constructing a trainable implicit causal graph structure based on the LINGAM causal discovery algorithm and gradient masking mechanism. This structure infers the potential causal direction between the image quality factor and the washout state judgment result through variational graph learning, generating a dynamically optimized causal adjacency matrix. Specifically, this includes: S3.1: Based on the multi-scale feature map in the shared feature space, extract the high-level semantic feature vector related to the image quality factor and the washout state, and use it as the input condition for implicit causal graph modeling; S3.2: The LINGAM causal discovery algorithm is used to initially model the potential causal relationship between the image quality factor and the washout state feature vector, and the causal direction is inferred based on the non-Gaussianity assumption to generate an initial causal adjacency matrix; Based on the image quality factor feature vector and the washout state feature vector output in step S3.1, the LINGAM causal discovery algorithm (parameters: non-Gaussian assumption, data covariance matrix and sample size have been obtained in step S2) is used to achieve a preliminary modeling of the potential causal relationship between the two. Furthermore, by using the Independent Component Analysis (ICA) method (parameters: maximum likelihood estimation objective function, unmixing matrix dimension set to eigenvector dimension), the independent factors in the feature space are separated, and the source signal representation that meets the non-Gaussian independence condition is obtained; Furthermore, by evaluating the temporal arrangement of the source signals and the fitting residuals of the structural equation model (SEM), the directed weight matrix W required for causal orientation determination is calculated, where the components... Represents the path coefficient from quality factor i to state factor j; Furthermore, based on the linear non-Gaussian model framework of LINGAM, zero-constraint optimization is performed on the coefficient matrix W to eliminate path connections that are insignificant under the Gaussianity test, thereby obtaining a sparsified initial causal adjacency matrix A, where =1 indicates that a causal path from i to j exists. =0 indicates that there is no direct causal path; By using the above-mentioned causal direction inference processing method, the high-level semantic feature vector of the previous step is transformed into an explicit initial causal adjacency matrix, thus realizing the initial construction of the implicit causal graph structure. For example, after processing high-resolution endoscopic images acquired clinically in step S2, the image quality factor feature vector has a dimension of 6 (sharpness, contrast, noise level, brightness uniformity, artifact intensity, color fidelity), and the washout state feature vector has a dimension of 4 (percentage of residual pixels, uniformity of distribution, number of bright spots, percentage of low-brightness shadow areas). During ICA separation, maximum likelihood estimation is used, and the unmixing matrix is 6×6. Under the non-Gaussianity test, the absolute value of the skewness of the source signals is greater than the critical value of 2.5, satisfying the LINGAM hypothesis. A structural equation model is used for fitting, calculating each element of the coefficient matrix W, and testing the Gaussianity p-value. Insignificant paths are set to zero, ultimately obtaining the initial causal adjacency matrix A, where, for example… =1 indicates that the noise level has a direct causal impact on the proportion of residual pixels. In the subsequent S3.3 gradient mask optimization, this matrix will serve as the basic structure for trainable optimization, significantly improving the interpretability of the causal graph and the performance of joint model learning; S3.3: Introduce a gradient masking mechanism to trainably optimize the initial causal adjacency matrix. Based on the gradient response of the task loss to the feature channel during backpropagation, dynamically adjust the connection strength coefficient of each causal path to enhance the learnability of the causal graph. Based on the node connection relationship of the initial causal adjacency matrix, a gradient mask generation method (parameters: backpropagation gradient of task loss, feature channel index) is used to calculate the channel-by-channel gradient response of the task branch output error signal, thereby realizing the extraction of the sensitivity mapping relationship between feature channels and task loss. Furthermore, by using a gradient normalization method (parameter: L2 norm normalization coefficient), the gradient values of each channel are scaled proportionally to the L2 norm of the overall gradient vector to achieve uniformity of gradient magnitudes across different channels, thus preventing an imbalance in causal weight updates caused by an excessively large gradient in a single channel. Furthermore, the normalized gradient value is compared with the preset threshold by the channel mask matrix construction method (parameters: number of channels, mask threshold) to generate a 0 / 1 binary mask matrix, where the channel with a value of 1 corresponds to a high-sensitivity path with a gradient greater than the threshold, and the channel with a value of 0 corresponds to a low-sensitivity path, which is used to selectively participate in the causal weight update calculation. Furthermore, the initial causal adjacency matrix is dynamically optimized using a weighted connection strength update formula, as follows:
[0016] in, Let be the connection strength coefficient from node i to node j. The learning rate parameter, The corresponding element values of the mask matrix. Let be the gradient response values from node i to node j; Furthermore, the mask generation and connection strength update process described above is repeated in multiple training batches using an iterative optimization method (parameters: maximum number of iterations, convergence tolerance) to achieve gradual convergence optimization of the implicit causal graph structure. By using a gradient masking mechanism, the initial causal adjacency matrix of the previous step is transformed into a dynamically updatable connection strength matrix, thereby achieving the learningability improvement effect of the causal structure under the task loss drive. For example, in the scenario of joint modeling of image quality and decontamination status in gastrointestinal endoscopy, the number of feature channels is set to 256. Backpropagation is used to calculate the gradient vector of the task branch, with the absolute value of the gradient ranging from 0.002 to 0.45. L2 norm normalization is used, with the normalization coefficient set as the L2 norm value of the overall gradient vector. After normalization, the gradient range is between 0.0005 and 0.12. The mask threshold is set to 0.05, and approximately 102 channels in the generated binary mask matrix have a value of 1, corresponding to high-sensitivity paths. The learning rate η is set to 0.001, and the initial causal adjacency matrix (element range 0.1 to 0.8) is updated using the above update formula with weighted updates. The average increase in the connection strength of high-sensitivity paths is significantly improved, while the connection strength of low-sensitivity paths remains basically stable. After 50 iterations of optimization, the overall change of the matrix tends to stabilize. The final optimized causal adjacency matrix significantly enhances the robustness of causal inference and the ability to identify paths affected by image quality degradation in subsequent variational graph learning steps. S3.4: Perform variational graph learning based on the optimized causal adjacency matrix. By introducing a differentiable graph sampling strategy, the probability distribution of the causal structure graph is dynamically updated during training to improve the robustness and generalization ability of causal inference. Based on receiving the optimized causal adjacency matrix from S3.3 as the input condition for this sub-step, a variational graph learning method (parameters: implicit causal graph node set V, edge set E, causal direction weight matrix W) is adopted to realize the dynamic distribution model of the causal structure in the probability space. Furthermore, by employing a differentiable graph sampling strategy (parameters: node sampling temperature τ, edge existence Bernoulli parameter p), a continuous approximate representation of the existence and direction probability of each potential causal edge is achieved, resulting in a probabilistic adjacency matrix sample set. This is used to support the causal structure update in gradient backpropagation; Furthermore, by optimizing the objective function and adding a KL divergence regularization term (parameters: current distribution q(G) and prior distribution p(G)), the causal structure sampling distribution is constrained to control the prior structure, and a stability improvement index is generated to prevent training instability caused by unconstrained structural drift. Furthermore, the Monte Carlo sampling method (parameters: number of samplings K, sampling batch B) is adopted to realize the expected estimation of the model task loss under multiple graph structure sampling, and generate the expected loss gradient vector, which is used to update the continuous parameters of the causal relationship strength weight matrix W. Furthermore, the above sampling and expectation calculation process is used to calculate the loss gradient under a single sampling using the following formula:
[0017] in, For the number of samples, For the k-th sampling diagram structure, For the joint task loss function; Through the above variational graph learning and differentiable sampling processing, the optimization results of the previous step are transformed into a stable and robust causal relationship distribution in the conditional probability space, thereby improving the generalization and adaptability to random structure perturbations in multi-task modeling scenarios. For example, in a practical digestive endoscopy washing and disinfection quality assessment task, the node set V includes four factors: image sharpness, contrast, noise level, and residual detection score. The edge set E is determined in the initialization phase based on the adjacency matrix optimized in S3.3. The sampling temperature τ is set to 0.1, and the edge existence parameter p is obtained by applying the Sigmoid function to the elements of the continuous weight matrix W. The sampling batch B is 16, and the number of Monte Carlo samplings K per batch is 10. The joint task loss is calculated for each sampled image structure, where the image quality assessment branch uses the mean squared error loss, and the disinfection state branch uses the cross-entropy loss. The prior distribution p(G) in the KL regularization term is selected from the initial structure distribution obtained from the LINGAM causal inference in stage S3.2. The sampling gradient is cumulatively averaged, and the gradient update vector is obtained according to the above formula. This vector is backpropagated to the continuous representation of W, and the updated W is reflected in the change of structure connection probability in subsequent sampling. In this embodiment, the sampling distribution converges stably within 30 cycles during the training process, and the probability of the structural edge direction maintains a high weight among highly correlated factors and approaches zero among low correlated factors, which significantly improves the robustness and generalization ability of causal reasoning under complex and degenerate conditions. S3.5: Output the dynamically optimized causal adjacency matrix to the subsequent causal intervention module as structured prior knowledge for feature reweighting and causal compensation operations, so as to realize explicit modeling of causal relationships between tasks; The dynamically optimized causal adjacency matrix output by step S3.4 is used as input condition. The causal intervention module structure configuration interface is called to bind the adjacency matrix to the internal causal path storage unit of the causal intervention module, so as to ensure that the causal relationship on which the subsequent feature reweighting operation is based is the optimal structure prior of the current model training stage. A node mapping method (parameters: node number lookup table, feature channel index) is used to realize the explicit correspondence between the nodes of the causal adjacency matrix and the corresponding image quality factor vectors and washout state vectors in the shared feature space, and to generate a causal path mapping table to guide the path selection and intervention direction determination in the feature intervention process. Furthermore, by using the edge weight assignment method (parameters: non-zero element values of the adjacency matrix, edge weight normalization coefficient), the weighted loading of causal path connection strength within the causal intervention module is realized, and normalization processing is used to ensure that the contribution of each path is comparable under a unified scale during multi-path comprehensive intervention, thus obtaining the causal compensation weight vector. Furthermore, a structured prior embedding algorithm (parameters: adjacency matrix sparsity factor, prior preservation rate) is applied to embed the causal adjacency matrix as a graph structure prior into the feature graph processing link of the module, and the interpretability of explicit modeling of the critical path is maintained by sparsity constraints, while suppressing the interference effect brought by non-critical paths. By using the feature mapping transformation method (parameters: channel mapping matrix, scale matching coefficient), the causal path weights processed by prior embedding are applied to the multi-scale feature map of the shared feature space, ensuring that subsequent intervention operations and compensation mechanisms complete data mapping based on the explicit causal relationship of the current stage, and achieving consistency between feature reweighting and causal compensation. Using the above method, the results of the previous step are transformed into structured prior data that can directly drive the causal intervention module, thereby realizing explicit modeling of causal relationships between tasks and path controllability of subsequent low-quality area compensation operations. For example, in a clinical setting, a batch of digestive endoscopy images with different resolution levels are acquired. The causal adjacency matrix generated by step S3.4 has a dimension of 100×100 and a sparsity factor set to 0.85. The number of non-zero edges from image quality factor nodes to cleansing state nodes in the adjacency matrix is 12, with edge weights ranging from 0.15 to 0.62. The node number lookup table is configured as image quality factor node index [1-5] and cleansing state node index [6-10]. The feature channel index uses the same channel number as the shared feature space. The non-zero elements are normalized using an edge weight normalization coefficient of 0.8, resulting in normalized edge weights ranging from 0.12 to 0.50. When embedding the causal intervention module, the prior preservation rate is set to 0.9 to ensure that the main causal path information is not truncated. By applying a channel mapping matrix (1024 channels of size-matched shared feature map) and a scale matching coefficient of 1.0, prior data is embedded into a multi-scale feature map. When the intervention module detects an image region with a low resolution level of 1 and the causal path weight of that region is higher than 0.45, it initiates a compensation mechanism and performs exponential mapping on the activation value decay of that region. The output can be used as the feature reweighting input for step S4, thereby significantly improving the stability of the washout state determination under low-quality conditions.
[0018] like Figure 3 As shown, step S4 involves performing a causal intervention feature reweighting operation on the shared feature space based on the implicit causal graph structure. When a conflict between a low-quality image region and a residual judgment is detected, a region confidence decay factor is generated to correct the feature weights by virtually enhancing the clarity and monitoring the sensitivity of output changes. Specifically, this includes: S4.1: Based on the multi-scale feature map and implicit causal graph structure in the shared feature space, the potential causal path between the image quality factor and the washout state judgment result is identified. The image quality factor includes sharpness, contrast and noise level. The potential causal path is represented by non-zero edges in the causal adjacency matrix to determine the input object and intervention direction of causal intervention. The input conditions are multi-scale feature maps in the shared feature space and dynamic causal adjacency matrix optimized by step S3. The feature maps and causal adjacency matrix together serve as inputs for identifying potential causal paths between image quality factors and washout state determination. The causal path parsing method (parameter: the weight threshold of the causal adjacency matrix is 0.05) is used to filter the node pairs with non-zero edge connections in the causal adjacency matrix, and map the feature channel indices corresponding to these node pairs to the shared feature space to form a preliminary set of potential causal paths. Furthermore, by using an image quality factor decoding algorithm (parameters: sharpness is represented by gradient magnitude statistics, contrast by histogram standard deviation, and noise level by high-frequency domain energy ratio), the image quality variable nodes in the potential causal path set are quantitatively represented, and the source node quality factor vector of each path is obtained. Furthermore, a feature correlation calculation method (parameter: correlation threshold is 0.6) is used to calculate the linear correlation coefficient between the source node quality factor and the target node washout status feature, and generate a causal path correlation score; when the score is greater than the threshold, the path is marked as a high-confidence path; Furthermore, the causal direction consistency verification method (parameters: weighted Granger causality test, significance level 0.01) is used to re-verify the causal direction of the high-confidence path, and the verification results are mapped to the edge direction settings in the causal adjacency matrix to generate a list of verified causal paths. By using causal path analysis and verification, the shared features and initial causal structure from the previous step are transformed into highly reliable causal path data with a defined direction, enabling precise determination of the input object and intervention direction for causal intervention operations in low-quality image regions. For example, in a clinical digestive endoscopy scenario, the input multi-scale feature map size is 256×256×128. After filtering with a weight threshold of 0.05 using the causal adjacency matrix, 15 non-zero edge connection paths are obtained. For the image quality factor nodes of each path, the sharpness is calculated using gradient magnitude statistics, yielding an average sharpness of 120.4; the contrast is calculated using histogram standard deviation, yielding an average of 35.6; and the noise level is calculated using high-frequency domain energy ratio, yielding an average ratio of 0.18. Using the linear correlation coefficient calculation method, these image quality factors are matched with the corresponding washout state features, resulting in 9 paths with a correlation greater than 0.6. A weighted Granger causality test is performed on these 9 paths, with a significance level set at 0.01, and 7 of the paths pass the causal direction consistency verification. The node directions of these 7 paths are clearly marked in the causal adjacency matrix, providing specific intervention input objects and directions for the counterfactual intervention operation in step S4.2, improving the targeting and effectiveness of subsequent feature reweighting. S4.2: Perform counterfactual intervention on the detected low-quality image region, and virtually enhance the clarity representation of the region in the feature space based on the causal path. Specifically, local perturbation of the feature map is performed by introducing a differentiable image enhancement operator to generate an enhanced feature representation. S4.3: Perform forward inference of the washout state branch based on the enhanced feature representation, calculate the change sensitivity between the original judgment result and the enhanced judgment result, and use KL divergence to measure the difference in the output probability distribution to quantify the sensitivity of the image quality of this region to the residual judgment. S4.4: Generate a region confidence decay factor based on the change sensitivity, and use an exponential decay function to map the sensitivity to a weight value between 0 and 1, where high sensitivity corresponds to low weight, which is used to weight and decay the activation value of the corresponding region in the original feature map. Based on the change sensitivity matrix calculated in S4.3 as input conditions, an exponential decay function is used to generate a confidence weight mapping table for low-quality image regions, thereby achieving controllable adjustment of the weights of sensitive regions. Furthermore, by defining the basic form of the exponential decay function... (in The weighted values after decay. For sensitivity to change, (where is the attenuation coefficient) to map the change sensitivity to a continuous weight value distribution in the interval of 0 to 1, and obtain a monotonically decreasing relationship curve between sensitivity and weight value; Furthermore, a sensitivity normalization operator is introduced during the weight mapping process. This enables the unified processing of sensitivity values for different image regions, thus avoiding the distortion in weight generation caused by excessively large spans in sensitivity intervals. Furthermore, by performing segmented threshold adjustment on the normalized sensitivity, a high sensitivity threshold is set based on clinical experience. ,when Force the weight to be set to the minimum value. This is used to completely suppress the feature contributions of extremely low-quality regions; Furthermore, by performing element-wise matrix mapping operations, the weight values generated above are mapped to the spatial coordinates of the original feature map, thereby constructing a region-level weight decay matrix; By generating the region confidence decay factor matrix, a multiplicative weighting operation is performed on the activation values of each pixel in the original feature map of the previous step, and the change sensitivity quantification result of the previous step is transformed into an input control parameter that can be directly applied to the feature reweighting module, thereby achieving a precise suppression effect on low-quality image region features. For example, during the operation of a digestive endoscopy washing and disinfection quality assessment model, the resolution of the multi-scale feature map in the input shared feature space is 512×512, the change sensitivity matrix ranges from 0.02 to 0.94, and the attenuation coefficient is set. High sensitivity threshold The minimum weight value is 0. After normalization, a low-quality region with a change sensitivity of 0.88 is assigned a weight of 0, and the convolution activation value at the corresponding position in the original feature map is directly reduced from 2.35 to 0, effectively eliminating the risk of residual misclassification in this region. Another region with a change sensitivity of 0.45 has its weight calculated using an exponential decay function. The value was approximately 0.26, corresponding to the activation value of this region in the feature map decreasing from 1.80 to approximately 0.47, which significantly improved the robustness and consistency of the model to low-quality images during the multi-task inference stage. S4.5: Apply the region confidence decay factor to the multi-scale feature map in the shared feature space, and reweight the feature map through element-wise multiplication to generate a corrected feature representation, which serves as the input basis for subsequent task branch inference, thereby improving the model's robustness and judgment consistency for low-quality images.
[0019] Step S5: Construct a differentiable causal attribution head and generate a residual distribution difference map under counterfactual generation conditions. Use this difference map as an auxiliary supervision signal and weightedly fuse it with the original task loss function to form a closed-loop optimized multi-task loss function. Specifically, this includes: S5.1: Based on the multi-scale feature map output by the shared feature space, a differentiable causal attribution head network structure is constructed. The causal attribution head contains multiple fully connected layers and normalization operations to support gradient propagation for causal inference paths. S5.2: Perform interpretability processing based on gradient class activation mapping (Grad-CAM) on the feature map input to the causal attribution head to generate a causal attribution heatmap, which represents the causal contribution of image regions to the determination of washout residue under counterfactual intervention conditions. Based on the multi-scale feature map data input by the causal attribution head, the Gradient Class Activation Mapping (Grad-CAM) method (parameter: select the predicted class of the washout state branch as the target class) is used to calculate the gradient response of each channel in the convolutional feature map in the class determination. Furthermore, gradient information is transmitted between the fully connected layer of the causal attribution head and the shared feature extraction layer through backpropagation (parameter: keep the weights of the causal path nodes at the current training state value), thereby realizing the extraction of the global average pooling gradient vector of each convolutional channel and obtaining the channel importance weight coefficients. Furthermore, by using a weighted summation method (parameters: convolution channel weight coefficients, corresponding channel feature maps), feature response accumulation in the spatial dimension is achieved, and a preliminary spatial activation map is generated to characterize the degree of influence of each spatial location on the current judgment task output; Furthermore, by replacing counterfactual features under causal intervention conditions (parameter: replacing low-quality region features with ideal quality-enhancing features), the activation map is recalculated, and two sets of spatial activation values are generated, corresponding to the original conditions and counterfactual intervention conditions, respectively. Furthermore, by using the spatial difference mapping method (parameters: original spatial activation map and counterfactual spatial activation map), the difference between the two sets of activation maps at corresponding spatial locations is calculated, and a causal attribution heatmap is generated after pixel-level normalization. This heatmap quantitatively characterizes the causal contribution of each image region to the judgment output of the washout residue under the counterfactual intervention condition. By combining gradient-based activation mapping with causal intervention conditions, the feature map response results of the previous step are transformed into a spatial thermal distribution with causal interpretability, thereby realizing the model's visual causal interpretation effect in multi-task joint learning. For example, under the dual-flow path conditions of high-resolution endoscopic images acquired clinically and their degraded versions, the "residue presence" category for washing residue determination is selected as the Grad-CAM target category, and the gradient response of the 512-channel feature map input to the causal attribution head is calculated for this category. The gradient vector is then subjected to global average pooling to obtain 512 channel importance weights, whose values are normalized to the range of 0 to 1. These weights are used to weight and sum the feature maps under the original conditions to generate a 64×64 spatial activation map. Subsequently, under counterfactual conditions, the detected low-resolution region features are replaced with a feature matrix enhanced based on an adaptive sharpening operator, and the activation map is recalculated. By comparing the original and counterfactual activation maps, a causal attribution heatmap is generated through pixel-by-pixel difference and encoded in pseudo-color within the range of 0 to 255 using linear normalization. The high-value region of this heatmap is significantly concentrated at a mucosal fold in the center of the image, indicating that this region has a high causal influence on the residue determination result under image resolution degradation, verifying the ability of this step to visualize and explain the model's decision-making basis. S5.3: Based on the generated causal attribution heatmap and the original image input, perform counterfactual generation operation to simulate the residual distribution prediction map under ideal image quality conditions, so as to obtain the counterfactual residual distribution map; S5.4: Calculate the pixel-level difference between the original residual distribution prediction map and the counterfactual residual distribution prediction map to generate a residual distribution difference map, which is used to characterize the degree of causal influence of image quality degradation on the washout status judgment result; S5.5: The residual distribution difference map is used as an auxiliary supervision signal and weighted and fused with the image quality assessment task loss and the washout status judgment task loss to construct a multi-task joint loss function containing causal attribution constraints; The residual distribution difference map generated by S5.4, along with the image quality assessment task loss value and the decontamination status judgment task loss value acquired simultaneously, are used as input data to perform weighted fusion calculation within a multi-task joint optimization framework. The task weight allocation algorithm is adopted (parameters: Assign weights to the image quality assessment task. Assigning weights to the cleansing and disinfection status task. To assist in supervising the causal attribution constraints, the weights of the supervision signals are used to set the proportion of each task's loss in the initial fusion stage, and are dynamically adjusted based on the gradient contribution of the difference map. The value of ; Furthermore, an adaptive weight adjustment method based on gradient sensitivity (parameter: weight update rate) is used. This enables real-time weight adjustment of causal attribution constraints in the multi-task loss function, and the calculation formula is as follows:
[0020] in, This represents the current value of the causal constraint loss term. This is the current value of the total loss function; Furthermore, the multi-task loss is constructed using the total loss fusion formula, as follows:
[0021] in, For image quality assessment task loss, To assess the mission loss based on the decontamination status, The auxiliary loss is a causal constraint generated by the residual distribution difference map; Furthermore, a loss balancing mechanism based on normalization weighting (parameter: weight normalization threshold τ) is used to normalize each loss component to avoid excessive loss in a single task that could cause a shift in the global optimization direction. Through the aforementioned weighted fusion algorithm and dynamic weight adjustment mechanism, the causal structure constraints fed back by the S5.4 residual distribution difference map are effectively superimposed on the original task loss, achieving simultaneous optimization of causal interpretability and task accuracy. For example, in a specific implementation process, let the initial weights of the image quality assessment task be set. The initial weight for the disinfection status determination task is 0.4. The initial weight for the causal constraint task is 0.4. The weight update rate η is set to 0.05, and the weight normalization threshold τ is set to 1.0. During the training iteration, the system detects a batch of causal constraint losses. The total loss is 0.12. The value is 0.65. The causal weight is updated according to the formula. Received the updated Approximately 0.2092. The weight set is updated accordingly. =0.4、 =0.4、 ≈0.2092, substituting into the total loss formula, we can obtain... The final total loss value is approximately 0.2331. With this parameter configuration, the model significantly improves judgment consistency and causal explanatory power under low-quality image conditions while maintaining task accuracy. S5.6: Perform gradient backpropagation based on the multi-task joint loss function to update the parameters of the causal attribution head, task branch head, and shared feature extraction network, so as to achieve synergistic optimization of the model between causal interpretability and task accuracy.
[0022] Step S6: A progressive causal freezing strategy is used to train the multi-task deep learning architecture in stages: the first stage fixes the causal graph structure and optimizes the parameters of each task branch; the second stage jointly fine-tunes the causal graph and all network parameters until the model converges. Specifically, this includes: S6.1: Based on the constructed trainable implicit causal graph structure and multi-task loss function, initialize all parameters of the multi-task deep learning architecture, including shared convolutional layer parameters, task-specific branch parameters, and causal adjacency matrix parameters, to generate an initial training model. Based on the constructed trainable implicit causal graph structure and multi-task loss function, the input conditions include: initial values of parameters of the shared convolutional feature extraction layer generated by the dual-stream backbone network, initial network weights of the causal attribution head and task-specific branches, initial state of the causal adjacency matrix obtained by optimization according to step S5, and weight coefficient settings in the joint loss function. A parameter initialization method (parameters: based on He Normal distribution, mean 0, standard deviation calculated inversely proportional to the number of output channels of the previous layer) is adopted to initialize the weights of the shared convolutional layer in order to adapt to the stable propagation of gradients in deep convolutional networks; Furthermore, the Xavier Uniform method (parameter: the weight range depends on the square root of the number of neurons in the preceding and following layers) is used to initialize the network parameters of the task-specific branch and obtain an initial weight matrix that meets the gradient balancing requirements. Furthermore, a random orthogonal matrix initialization method is adopted (parameter: the matrix dimension is consistent with the number of nodes in the causal adjacency matrix) to achieve the initial setting of the parameters of the causal adjacency matrix and generate a structured sparse layout of the connection weights between nodes, so as to enhance the learnability and initial stability of the causal structure. Furthermore, by normalizing the weights of the joint loss function (parameter: the sum of the loss weights for each task is normalized to 1.0), the initial parameters of the multi-task loss function are set, and a weight vector balancing the gradient contributions between tasks is generated. ,in For image quality assessment task loss weights, The task loss weight is determined based on the disinfection status. Through the above algorithms and processing methods, the causal structure optimization results of the previous step are transformed into global initial training model parameters, realizing the unified initialization effect of the multi-task deep learning architecture, and laying a stable parameter foundation for the causal freeze training in the first stage of S6.2. For example, in a real-world clinical endoscopic disinfection quality assessment scenario, using a training dataset containing 5000 high-resolution original images and their degraded versions, the initial parameters of the shared convolutional layers are configured as He Normal distributions, where the standard deviation is... , Input the number of channels. The product of the convolution kernel spatial dimensions; the task-specific branch parameters are taken using the XavierUniform method, taking the upper limit of the range. , The number of input neurons, The number of output neurons is set; the causal adjacency matrix is initialized as a 10×10 dimensional random orthogonal matrix, and the proportion of non-zero elements is kept below 0.3 through sparsity constraints. In the joint loss function, the weights of the image quality assessment task and the deduplication state judgment task are set to 0.45 and 0.55, respectively, and the total weight is kept at 1.0 after normalization. After performing the above initialization, the gradient distribution of the network on the pre-training validation set tends to be balanced, avoiding the gradient explosion or vanishing problem caused by differences in parameter scales, and providing a stable convergence environment for the subsequent first-stage single-task optimization training; S6.2: In the first stage of training, the parameters of the causal adjacency matrix are frozen. Based on the multi-scale feature map in the shared feature space, the parameters of the image quality assessment branch and the washout state judgment branch are optimized separately to obtain a task-specific model that has initially converged. Based on the initialized multi-task deep learning architecture parameters, the input conditions include multi-scale feature maps in the shared feature space, image quality assessment task branch parameters, and descaling status judgment task branch parameters. A parameter freezing operation (object: causal adjacency matrix) is used to maintain a stable topology of the implicit causal structure in the early stage of training, so as to avoid early convergence deviation caused by the disordered fluctuation of causal path weights with the task gradient. Furthermore, the parameters of the image quality assessment task branch are iteratively optimized using the stochastic gradient descent algorithm (learning rate: 0.001, weight decay coefficient: 1e-4), and convergence trend curve data is obtained to monitor the improvement of the independent performance of the branch. Furthermore, the Adam optimization algorithm (parameters: β1=0.9, β2=0.999) is adopted to achieve efficient optimization of the branch parameters of the cleansing state judgment task, and to generate the accuracy and cross-entropy loss change results on the validation set; Furthermore, by using the multi-task branch gradient isolation method, the backpropagation gradients of the two tasks are applied to their respective branch parameter sets, and the gradient update of the frozen causal adjacency matrix is blocked, thereby achieving path-selective control of the gradient flow. Through the above optimization process, the task-specific branch parameters of the initial network are adjusted to weight configurations that are stable in their respective tasks, thus achieving a preliminary convergence of the task-specific model and providing a stable starting point for subsequent unfreezing of the causal graph structure and joint fine-tuning. For example, 5000 high-resolution endoscopic images and corresponding degraded images acquired clinically are selected. A multi-scale feature map of size (256×256×64) is generated using the feature extraction module as training input. When freezing the causal adjacency matrix parameters, the matrix dimension is fixed at (10×10), corresponding to the causal relationships of 10 potential factor nodes. For the image quality assessment branch, a stochastic gradient descent algorithm is used, with a batch size of 32 and 50 training iterations. The mean squared error after convergence is calculated using the following formula: ,in The true image quality factor. Here, n is the number of samples, and n represents the predicted value. For the cleansing state judgment branch, the Adam optimizer is used with a batch size of 16 and 40 iterations. The cross-entropy loss is calculated using the following formula: ,in For the true label distribution, To predict the probability distribution. After the two branches are trained independently, the prediction error of the image quality factor is significantly reduced, the accuracy of the washout state judgment is significantly improved, and the performance curve on the validation set shows a stable upward trend compared with the initial freezing stage, laying the training baseline for the second stage of joint fine-tuning of the causal structure; S6.3: Based on the convergence of the task branch, and based on the counterfactual residual distribution difference map generated by the differentiable causal attribution head, calculate the influence gradient of the causal intervention module on the feature space to update the node connection weights in the causal adjacency matrix, thereby optimizing the implicit causal graph structure. S6.4: In the second stage of training, the freezing of the causal adjacency matrix parameters is lifted, and the shared convolutional layer, task branch and causal graph structure are jointly fine-tuned end-to-end based on the multi-task loss function to obtain the global optimal model that coordinates the optimization of causal relationship and task prediction. The shared feature space multi-scale feature map is based on the task branch parameters that have been optimized and converged in the first stage. The input conditions include the task-specific branch parameters that have been optimized with a fixed causal graph structure and the weight coefficient settings of the multi-task joint loss function. A parameter unfreezing operation is used to release the node connection weights and their corresponding gradient update paths in the causal adjacency matrix, so as to allow them to participate in the global backpropagation optimization process. Furthermore, by using a multi-task joint loss function (parameters: image quality assessment loss weight α, washout state judgment loss weight β, causal attribution constraint loss weight γ), the synchronous iterative update of parameters of the shared convolutional feature extraction layer, task branch head, and causal inference module is achieved, and the collaborative gradient matrix of causal relationship and task prediction result is obtained. Furthermore, based on the adaptive learning rate scheduling algorithm (parameters: initial learning rate η=0.001, decay factor δ=0.75, decay period T=10), the learning rate of different parameter subsets is dynamically adjusted during the optimization process to enhance the convergence balance between causal structure optimization and task prediction accuracy. Furthermore, a gradient truncation mechanism (parameter: gradient norm threshold ρ=5.0) is adopted to prevent gradient explosion caused by the adjustment of the causal graph structure and to generate a stable parameter update step vector. Through a global joint fine-tuning process, the causal adjacency matrix parameters and shared convolutional feature parameters frozen in the previous stage are transformed into a globally optimal parameter set containing causal path optimization information under a unified optimization framework, thereby achieving the expected technical effects of enhanced causal relationship interpretability and improved task prediction accuracy. For example, in training clinical digestive endoscopy image data, the multi-task joint loss function is set to α=0.4, β=0.4, γ=0.2, the initial learning rate η=0.001, the decay factor δ=0.75, the decay period T=10 rounds, and the gradient norm threshold ρ=5.0. During the global joint fine-tuning phase, the causal adjacency matrix parameters participate in gradient updates starting from the 11th round. The Adam optimizer (parameters: β1=0.9, β2=0.999) is used to jointly optimize the shared convolutional layers and the causal inference module. For the update of the connection weights of the causal graph, the gradient vector of the multi-task joint loss is calculated, and its norm is kept less than 5 by gradient truncation. In the 15th round, the mean square error of image quality assessment on the validation set decreased from 0.028 in the previous stage to 0.019, the accuracy of the washout state judgment increased from 0.87 to 0.92, the interpretability score of the causal attribution difference map was significantly improved, and the global optimal parameter set was output for subsequent deployment, realizing the joint optimization of robust recognition and high interpretability under image quality fluctuation conditions. S6.5: Based on the image quality factor prediction error, the accuracy of the washout status judgment, and the interpretability score of the causal attribution difference map on the validation set, evaluate the multi-task performance of the globally optimal model, and select the best model snapshot for subsequent deployment according to the evaluation results, so as to realize dynamic quality closed-loop management in clinical scenarios.
[0023] Step S7: Based on the trained model, perform end-to-end quality assessment on newly acquired digestive endoscopy images, and output a three-dimensional assessment report including image quality factors, washout status judgment results, and causal attribution difference maps, realizing dynamic quality closed-loop management in clinical scenarios. Specifically, this includes: S7.1: Acquire newly acquired digestive endoscopy images as input for model inference, and extract features from the images based on the trained dual-stream backbone network to generate multi-scale feature maps in the shared feature space, providing basic feature representations for subsequent causal inference and task prediction; S7.2: Based on the generated multi-scale feature map, input it into the trained implicit causal graph structure to perform causal path inference between the image quality factor and the washout state judgment result, so as to obtain the dynamic causal weight of each node in the causal adjacency matrix, which is used to guide the subsequent feature reweighting operation. S7.3: Based on the output causal adjacency matrix, perform causal intervention feature reweighting operation on local image regions in the shared feature space. When image quality degradation and residual judgment conflict are detected, the region clarity is virtually improved and the sensitivity of the washout branch output change is monitored to generate a region confidence decay factor to correct the feature weights, thereby improving the robustness of the judgment result. S7.4: Input the corrected features into the differentiable causal attribution head to generate a residual distribution difference map under the counterfactual generation condition, and use this difference map as an auxiliary supervision signal to perform weighted fusion with the original task prediction results to form a multi-dimensional comprehensive quality assessment output. S7.5: Based on the fusion of multi-task outputs, a three-dimensional evaluation report is generated, which includes image quality factors, disinfection status judgment results and causal attribution difference diagrams. The impact of endoscopic image quality degradation path on disinfection residue identification is displayed in a visual form, realizing dynamic quality closed-loop management and decision support in clinical scenarios. Based on the multi-task output fusion in S7.4, a data structure assembly method (parameters: image quality factor vector, washout state classification result matrix, causal attribution difference image matrix) is adopted to realize the initial construction of the three-dimensional evaluation report; Furthermore, by using a metadata association and binding method (parameters: task output identifier, timestamp, patient information index), a two-way association between report data and the clinical management system is achieved, and traceable evaluation data records are obtained. Furthermore, through a multi-channel visualization rendering algorithm (parameters: color map table, transparency grading rules, spatial resolution setting), the image quality factor is presented in the form of a numerical scale, the washout status judgment result is overlaid on the surface of the endoscopic image with a regional classification label, and the causal attribution difference map is displayed in a semi-transparent overlay manner, generating an interactive evaluation screen in three-dimensional space. Furthermore, through the indicator normalization processing algorithm (parameter: minimum-maximum normalization range is set to 0 to 1), a unified quantitative scale for indicators of different dimensions in the report is achieved, and histograms, line charts and other statistical representations are generated as data appendices to the report. The automatic report template formatting module (parameters: font style, layout rules, and image-to-text ratio) transforms structured task output data into a complete report containing a cover summary, image distribution map, statistical analysis chart, and causal path explanation paragraph, enabling documented output that can be directly used for clinical decision support. For example, in a gastroscopy scenario, newly acquired high-definition endoscopic images are input into a trained model to obtain an image quality factor vector {sharpness 0.82, contrast 0.76, noise level 0.21}. The washout status judgment matrix shows that one of the three regions is identified as containing residue, and the sensitivity coefficient for this region in the causal attribution difference image matrix is 0.93. Based on a multi-channel visualization rendering algorithm, the sharpness map uses a blue gradient, the washout residue area is covered with red semi-transparent overlay, and the causal attribution difference map uses a green overlay, with transparency controlled by the sensitivity coefficient. Mini-maximum normalization is performed on the quality factors.
[0024] The numerator is the original sharpness minus the minimum noise level of 0.21, and the denominator is the maximum possible value minus the minimum noise level, achieving a unified quality score calculation. The rendered results are formatted into a report template, presenting the above values as a bar chart with a causal path explanation: "Image blurring → High-sensitivity residual misjudgment path," providing an intuitive visual analysis of the conflict between sharpness and the removal status. In clinical validation, this report significantly improved the accuracy of physicians' judgment on residual identification in low-quality images and effectively reduced repeated cleaning operations caused by misjudgments.
[0025] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0026] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and rules of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for assessing the quality of washing and disinfection of digestive endoscopes based on image recognition, characterized in that, Includes the following steps: S1: Perform dual-stream input processing on the original digestive endoscopy images and the degraded images enhanced by physical imaging simulation to generate a dual-path input dataset; S2: Based on a dual-stream backbone network with shared convolutional parameters, feature extraction is performed on the dual-path input dataset, and multi-scale feature maps in the shared feature space are output. S3: Based on the LINGAM causal discovery algorithm and gradient masking mechanism, a trainable implicit causal graph structure is constructed, in which the potential causal direction between the image quality factor and the washout state judgment result is inferred through variational graph learning, and a dynamically optimized causal adjacency matrix is generated. S4: Based on the implicit causal graph structure, perform causal intervention feature reweighting operation on the shared feature space. When a conflict is detected between a low-quality image region and a residual judgment, generate a region confidence decay factor by virtually improving the clarity and monitoring the sensitivity of output changes. S5: Construct a differentiable causal attribution head and generate a residual distribution difference map under counterfactual generation conditions. Use the difference map as an auxiliary supervision signal and perform weighted fusion with the original task loss function to form a closed-loop optimized multi-task loss function. S6: A progressive causal freezing strategy is adopted to train the multi-task deep learning architecture in stages. In the first stage, the causal graph structure is fixed and the parameters of each task branch are optimized. In the second stage, the causal graph and all network parameters are fine-tuned together until the model converges. S7: Perform end-to-end quality assessment on newly acquired digestive endoscopy images based on the trained model and output a three-dimensional assessment report.
2. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 1, characterized in that, Step S1 specifically includes: Acquire raw digestive endoscopy image data, which are derived from high-resolution endoscopic images acquired clinically, and which contain typical areas of the digestive tract mucosa surface; The original digestive endoscopy image data are simulated using a physical imaging simulation model to simulate the image degradation process, which includes simulating fogging effects, reflection interference, and motion blur, generating low-quality image samples with real clinical degradation characteristics. The original digestive endoscopy image data and the degraded image data enhanced by physical imaging simulation are used as dual-path input data to construct a dual-stream input structure, in which one input is the original image and the other input is the corresponding degraded image, forming a one-to-one dual-stream input pair; Image normalization processing is performed on the original image path in the dual-stream input structure. The brightness, contrast and color distribution of the image are standardized based on the image histogram equalization algorithm to obtain a unified image input format. Image quality parameters are labeled for the degraded image path in the dual-stream input structure. Image quality factors are extracted using an image sharpness assessment model, a noise estimation module, and a fuzzy quantification algorithm to generate image quality labels corresponding to the degraded image.
3. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 2, characterized in that, The original digestive endoscopy images are automatically retrieved using the Medical Image Archiving and Communication System Interface Protocol. The data preprocessing, which includes complete mucosal structures and removes abnormal noise regions, is performed in combination with the CBIR screening algorithm and Canny edge detection. The sharpness index is calculated using spatial frequency analysis, which quantifies spatial information by the ratio of frequency domain energy integral to zero frequency energy integral.
4. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 1, characterized in that, Step S2 specifically includes: Based on a dual-path input dataset, one path is the original digestive endoscopy image and the other path is the degraded image enhanced by physical imaging simulation. A dual-stream convolutional neural network architecture is constructed to extract the low-level visual semantic features of the two path input data respectively. In the shallow convolutional layers of the dual-stream backbone network, the convolutional kernel parameters of the two input paths are shared, and the output is a low-dimensional feature map in the shared feature space. The parameters of the mid-to-high-level convolutional layers of the dual-stream backbone network are configured independently to allow the two input paths to retain their differentiated feature expression capabilities under their respective image quality conditions during the high-level semantic feature extraction stage, thereby generating a high-level semantic feature map with path specificity. Based on multi-scale feature maps in a shared feature space, a feature pyramid network structure is used to perform cross-layer fusion processing on the multi-scale feature maps to generate a high-resolution fused feature map of a unified scale. The high-resolution fused feature map is input into the feature normalization module, and the channel weights of the high-resolution fused feature map are adjusted based on the channel attention mechanism to output the normalized enhanced feature map.
5. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 1, characterized in that, Step S3 specifically includes: Based on multi-scale feature maps in the shared feature space, high-level semantic feature vectors related to image quality factors and de-washing states are extracted. The LINGAM causal discovery algorithm is used to initially model the potential causal relationship between the image quality factor and the washout state feature vector. The causal direction is inferred based on the non-Gaussianity assumption, and an initial causal adjacency matrix is generated. A gradient masking mechanism is introduced to trainably optimize the initial causal adjacency matrix. Based on the gradient response of the task loss to the feature channel during backpropagation, the connection strength coefficient of each causal path is dynamically adjusted. Variational graph learning is performed based on the optimized causal adjacency matrix. By introducing a differentiable graph sampling strategy, the probability distribution of the causal structure graph is dynamically updated during training. The dynamically optimized causal adjacency matrix is output to the subsequent causal intervention module.
6. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 5, characterized in that, Step S3 further includes combining independent component analysis and evaluating the temporal arrangement and structural equation model between source signals to automatically generate an initial directional weighted adjacency matrix between the image quality factor and the washout state factor, and using gradient masking to filter highly sensitive paths for targeted trainable optimization.
7. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 1, characterized in that, Step S4 specifically includes: Based on the multi-scale feature map and implicit causal graph structure in the shared feature space, the potential causal path between the image quality factor and the washout status judgment result is identified. Counterfactual intervention is performed on the detected low-quality image regions. Based on the causal path, the sharpness representation of the region is virtually enhanced in the feature space. A differentiable image enhancement operator is introduced to locally perturb the feature map and generate an enhanced feature representation. Based on the enhanced feature representation, perform forward inference of the cleansing state branch, calculate the change sensitivity between the original judgment result and the enhanced judgment result, and use KL divergence to measure the difference in the output probability distribution. Based on the change sensitivity, a region confidence decay factor is generated, and an exponential decay function is used to map the sensitivity to a weight value between 0 and 1, where high sensitivity corresponds to low weight. The region confidence decay factor is applied to the multi-scale feature map in the shared feature space, and the feature map is reweighted through element-wise multiplication to generate the corrected feature representation.
8. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 7, characterized in that, The image quality factors include sharpness, contrast, and noise level, and the potential causal path is represented by non-zero edges in the causal adjacency matrix.
9. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 1, characterized in that, Step S5 specifically includes: Based on the multi-scale feature map output from the shared feature space, a differentiable causal attribution head network structure is constructed. The causal attribution head includes multiple fully connected layers and normalization operations. Perform interpretability processing based on gradient class activation mapping on the feature map input to the causal attribution head to generate a causal attribution heatmap; Based on the causal attribution heatmap and the original image input, a counterfactual generation operation is performed to simulate the residual distribution prediction map under ideal image quality conditions, and obtain the counterfactual residual distribution map. The original residual distribution prediction map and the counterfactual residual distribution prediction map are compared at the pixel level to generate a residual distribution difference map. The residual distribution difference map is used as an auxiliary supervision signal and weighted and fused with the image quality assessment task loss and the washout status judgment task loss to construct a multi-task joint loss function containing causal attribution constraints. Gradient backpropagation is performed based on the multi-task joint loss function to update the parameters of the causal attribution head, task branch heads, and shared feature extraction network.
10. The method for assessing the quality of digestive endoscope washing and disinfection based on image recognition according to claim 9, characterized in that, The causal attribution heatmap represents the causal contribution of image regions to the determination of washing residue under counterfactual intervention conditions.