Endoscope global background motion preserving method and electronic device
Patent Information
- Application Number
- CN202611029296.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-29
AI Technical Summary
[0002]内窥镜镜头常因器械操作、组织牵拉或生理活动而产生复杂且难以预测的全局运动,这种全局背景运动为手术视觉辅助系统带来了严峻挑战,而传统方法在面对该挑战时往往存在一定的局限性
[0046]获取原始双路内窥镜图像与环境状态数据,结合注意力机制与自适应滤波进行数据预处理,生成多模态内窥镜数据集,本步骤通过对异构数据进行精准时空对齐,为后续操作提供了严格一致的数据基础。
Smart Images

Figure CN122841184A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to an endoscope global background motion preservation method and electronic device. Background Technology
[0002] Endoscopic lenses often exhibit complex and unpredictable global motion due to instrument manipulation, tissue traction, or physiological activities. This global background motion poses a serious challenge to surgical visual aid systems, and traditional methods often have certain limitations in addressing this challenge.
[0003] On the one hand, traditional methods mostly rely on electronic image stabilization or rigid motion estimation and compensation based on feature points. However, in the acquired endoscopic images, the anatomical structural information representing the geometric shape of the tissue and the physiological functional information representing specific biological processes are often highly coupled and mutually contaminated. Traditional methods often find it difficult to achieve high-fidelity and physically interpretable separation of the two in complex dynamic scenes, which makes it difficult to accurately model and compensate for the complex non-rigid deformation of biological tissues, thus damaging the spatial consistency and anatomical realism of the scene corresponding to the endoscopic image.
[0004] On the other hand, traditional methods often lack an understanding of surgical semantics, making it difficult to dynamically and adaptively enhance key local areas while stabilizing the overall background. Summary of the Invention
[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for maintaining global background motion in an endoscope, the method comprising:
[0006] The raw dual-channel endoscope images and environmental state data are acquired, and data preprocessing is performed by combining attention mechanism and adaptive filtering to generate a multimodal endoscope dataset.
[0007] A pre-trained physical perception hierarchical decoupling network is used to decouple the features of the multimodal endoscope dataset to obtain a multidimensional decoupled dataset;
[0008] The neural implicit surface technique and meta-learning algorithm are used to simultaneously combine the multidimensional decoupled dataset to reconstruct multidimensional scenes and generate a reconstructed scene multidimensional dataset.
[0009] Based on the adaptive dual-path fusion mechanism, dual-path endoscopic images are fused simultaneously with the reconstructed scene multidimensional dataset to generate enhanced endoscopic images.
[0010] As a further aspect of the present invention, original dual-channel endoscopic images and environmental state data are acquired, and data preprocessing is performed using an attention mechanism and adaptive filtering to generate a multimodal endoscopic dataset, including:
[0011] The original dual-channel endoscope images are represented as two original Bayer images of the endoscope in two different wavelength bands, and the environmental status data includes at least the operation stage label and instrument status data.
[0012] A spatiotemporal convolutional network is used to extract features from the original dual-channel endoscope images and environmental state data to obtain endoscope image feature vectors and endoscope state feature vectors.
[0013] Based on the endoscope image feature vector and the endoscope state feature vector, temporal offset inference is performed to obtain the temporal offset;
[0014] Adaptive Kalman filtering is used, combined with the time offset, to perform time offset compensation and time alignment, and to obtain the corrected dual-channel endoscope images and the corresponding alignment environment state data.
[0015] The corrected dual-channel endoscopic images and the corresponding aligned environmental state data are output as a multimodal endoscopic dataset.
[0016] As a further aspect of the present invention, a pre-trained physical perception hierarchical decoupling network is used to decouple the features of the multimodal endoscope dataset to obtain a multidimensional decoupled dataset, including:
[0017] The physical perception hierarchical decoupling network performs forward propagation and feature decoupling on the multimodal endoscope dataset to generate structural feature maps and functional feature maps;
[0018] Monte Carlo Dropout is used to perform multiple forward propagations to obtain the forward propagation results; uncertainty quantification is performed based on the forward propagation results to generate a confidence graph of the decoupling results;
[0019] The region within the confidence map of the decoupling result that is lower than the preset confidence threshold is defined as the feature fuzzy region, and a multi-hypothesis decoupling set is generated based on the forward propagation result corresponding to the feature fuzzy region.
[0020] The quality of the multi-hypothesis decoupling set is evaluated based on a preset hypothesis quality evaluation function to obtain the hypothesis quality evaluation result.
[0021] Based on the hypothetical quality assessment results, the structural feature map and functional feature map are optimized to generate an enhanced structural feature map and an enhanced functional feature map.
[0022] The enhanced structural feature map, the enhanced functional feature map, and the decoupling result confidence map are encapsulated into a multidimensional decoupling dataset for output.
[0023] As a further aspect of the present invention, the physical perception hierarchical decoupling network includes:
[0024] The physical perception hierarchical decoupling network includes a hierarchical feature decoupling submodule and a differentiable physical rendering submodule;
[0025] The hierarchical feature decoupling submodule is used to extract shared features from the input multimodal endoscopy dataset, and combine two parallel branch networks to decode structural feature maps and functional feature maps from the shared features;
[0026] The differentiable physical rendering submodule is represented as a multilayer perceptron based on a bidirectional reflection distribution function, used to apply physical realism constraints to structural features.
[0027] The physical perception hierarchical decoupling network is trained using a phased learning strategy and a preset joint loss function until the preset training convergence condition is met.
[0028] As a further aspect of the present invention, neural implicit surface technology and meta-learning algorithms are used to simultaneously combine the aforementioned multidimensional decoupled dataset for multidimensional scene reconstruction, generating a reconstructed scene multidimensional dataset, including:
[0029] An image scene reconstruction model is constructed based on neural implicit surface technology and meta-learning algorithm. The image scene reconstruction model includes a scene representation network and a context information encoder.
[0030] The scene representation network is used to represent the geometric and appearance features of endoscopic images using neural implicit surface techniques, generating a scene representation set.
[0031] The context information encoder is used to encode the multidimensional decoupled dataset into a context information vector;
[0032] A phased reconstruction training strategy is adopted to train the image scene reconstruction model, and the parameters are fine-tuned by combining a meta-learning algorithm.
[0033] As a further aspect of the present invention, the method further includes:
[0034] Input the multidimensional decoupled dataset into the image scene reconstruction model to perform multidimensional scene reconstruction;
[0035] After the multidimensional decoupled dataset is input into the image scene reconstruction model, the context information encoder encodes the multidimensional decoupled dataset to generate a context information vector.
[0036] An inner loop update mechanism based on model-independent meta-learning is used to fine-tune the parameters of the scene representation network;
[0037] The fine-tuned scene representation network, based on the context information vector, simultaneously combines a depth estimation algorithm to reconstruct the scene and generate a set of reconstructed scene representations.
[0038] The context information vector and the reconstructed scene representation set are encapsulated into a reconstructed scene multidimensional dataset for output.
[0039] As a further aspect of the present invention, based on an adaptive dual-path fusion mechanism, dual-path endoscopic images are fused simultaneously with the reconstructed scene multidimensional dataset to generate enhanced endoscopic images, including:
[0040] Inference is performed based on the reconstructed scene multidimensional dataset and multidimensional decoupled dataset to generate a spectral mixing weight map and a local enhancement parameter map;
[0041] Based on the spectral mixing weight map and the set of reconstructed scene representations contained in the reconstructed scene multidimensional dataset, a basic fused endoscopic image is generated;
[0042] The base fused endoscope image is locally optimized based on the local enhancement parameter map to generate an enhanced endoscope image.
[0043] In another aspect, embodiments of the present invention also provide an electronic device, comprising:
[0044] At least one processor; and at least one memory communicatively connected to the processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform an endoscopic global background motion holding method provided in an embodiment of the present invention.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The raw dual-channel endoscopic images and environmental state data are acquired, and data preprocessing is performed using attention mechanisms and adaptive filtering to generate a multimodal endoscopic dataset. This step provides a strictly consistent data foundation for subsequent operations by accurately aligning heterogeneous data in time and space.
[0047] A pre-trained physical perception hierarchical decoupling network is used to decouple the features of the multimodal endoscope dataset to obtain a multidimensional decoupled dataset. This step separates the anatomical structure information and physiological function information in the endoscope image with high fidelity. This separation can independently perform background motion compensation and functional information enhancement to avoid mutual interference between the two, thereby ensuring the spatial consistency and anatomical authenticity of the endoscope image.
[0048] By employing neural implicit surface technology and meta-learning algorithms, multidimensional scene reconstruction is performed simultaneously with the aforementioned multidimensional decoupled dataset, generating a reconstructed scene multidimensional dataset. This step involves constructing a scene reconstruction model that includes semantic information and supports arbitrary viewpoint queries, thereby encoding visual, geometric, and semantic information contained in endoscopic images. This provides rich contextual information and spatial relationships for subsequent image fusion.
[0049] Based on the adaptive dual-path fusion mechanism, dual-path endoscopic images are fused simultaneously with the reconstructed scene multidimensional dataset to generate enhanced endoscopic images. This step achieves the fusion of multidimensional information provided in the previous steps, thereby achieving the preservation of global background motion and adaptive enhancement of local details. Attached Figure Description
[0050] Figure 1 This is a flowchart of the steps of an endoscope global background motion preservation method according to the present invention;
[0051] Figure 2 This is a schematic diagram of the physical perception layered decoupling network in the global background motion preservation method for endoscopes of the present invention. Detailed Implementation
[0052] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart of the steps of an endoscope global background motion preservation method according to the present invention. Figure 2 This is a schematic diagram of the physical perception layered decoupling network in the global background motion preservation method for endoscopes according to the present invention. The following is a detailed description of the global background motion preservation method for endoscopes.
[0053] Step S1: Obtain the original dual-channel endoscope images and environmental state data, and perform data preprocessing by combining attention mechanism and adaptive filtering to generate a multimodal endoscope dataset.
[0054] Understandably, the original dual-channel endoscopic images represent two original Bayer images from different spectral bands, i.e., image streams from two image sensors with complementary spectral responses. For example, the first sensor is configured with an RGGB Bayer filter array suitable for the visible light band, and the second sensor is configured with a BGGR Bayer filter array suitable for the near-infrared band, and both have consistent resolution, bit depth, and frame rate. The environmental status data includes at least operation stage labels and instrument status data. The operation stage labels are labels corresponding to actual operation signals, such as foot switches, voice commands, etc., which originate from the central control system of the operating room or human-computer interaction signals of the surgeon. These signals can be represented as discrete enumerated state variables. The instrument status data is a multi-channel digital signal. For example, channel 1: instrument type {0: no instrument, 1: gripper, 2: electrosurgical unit, 3: scissors, 4: needle holder}; channel 2: {instrument opening degree: (0, 100%)}; channel 3: {electrosurgical unit power level: [0, 10]}.
[0055] Specifically, a spatiotemporal convolutional network is used to extract features from the original dual-channel endoscopic images and environmental state data to obtain endoscopic image feature vectors and endoscopic state feature vectors. Based on the endoscopic image feature vectors and endoscopic state feature vectors, temporal offset inference is performed to obtain the temporal offset.
[0056] Furthermore, an adaptive Kalman filter is used, combined with the time offset, to perform time offset compensation and time alignment, to obtain the corrected dual-channel endoscope images and the corresponding alignment environment state data, and the corrected dual-channel endoscope images and the corresponding alignment environment state data are output as a multimodal endoscope dataset.
[0057] In one possible embodiment, for the original dual-channel endoscopic images, a lightweight spatiotemporal convolutional neural network with shared weights, such as the MobileNetV3-Small architecture, is used to extract spatial features. A shallow temporal convolutional layer is then used to analyze the optical flow motion information between consecutive frames, ultimately outputting a spatiotemporal feature vector for each frame of each channel image. For example, assuming the received image is the first... The visible light path image of the frame shows that the frame image is the abdominal adipose tissue region. At this time, the spatiotemporal convolutional network will extract the features representing the yellow tone and porous sponge texture of the region. At the same time, combined with the (t-1)th frame, it will identify information such as slight periodic displacement caused by breathing, and encode the extracted spatial and temporal information together into a 512-dimensional vector.
[0058] For the operation stage label, its corresponding discrete signal is first converted into a one-hot code, and then mapped into a multi-dimensional continuous vector through an embedding layer that has pre-learned the distributed representation of different surgical stages in the semantic space. This vector is the operation stage vector. For example, when the label of the "separation" stage is received from the central control system of the operating room, the vector generated after the embedded layer is trained may be closer to other stage labels involving tissue cutting or stripping in the vector space.
[0059] For instrument status data, since it is a signal that is a mixture of continuous and discrete signals, it is first necessary to perform a normalization operation, and then input it into a two-layer perceptron to encode it into a multi-dimensional instrument status vector. For example, when the electrosurgical unit is activated and the power is set to level 5, assuming that the structure of the perceptron in this case is an input layer, a hidden layer of 64 dimensions, and an output layer of 64 dimensions, the vector output by the perceptron will simultaneously encode the mixed status information of "the instrument type is an energy tool" and "currently in a medium energy output state".
[0060] This enables the acquisition, at any given moment, of an endoscope image feature vector composed of spatiotemporal feature vectors corresponding to the original dual-channel endoscope images, such as a 1024-dimensional vector formed by stitching together 512-dimensional image feature vectors from both channels; and an endoscope state feature vector composed of operation stage vectors and instrument state vectors.
[0061] Next, a multi-head attention mechanism is used for multimodal data alignment. For the current processing time, a sliding window with a time span of [-a, +a] for a total of 2a+1 frames is constructed. The sequence of endoscopic image feature vectors of all frames within the window is used as the query, and the sequence of endoscopic state feature vectors within the same window is used as the key and value. By calculating the dot product of the query and the key and scaling it, an attention weight matrix is generated. This matrix is used to characterize the correlation strength between the image features at each time step and the state features at all times. For example, assuming that at time t-2, the instrument state feature shows that the electrosurgical unit is activated, in an ideal attention weight matrix, the image features of frame t, t-1, or t-3 will generate a high attention score with the instrument state feature of frame t-2. By analyzing the pattern of the attention weight matrix along the time dimension, the corresponding time offset is obtained. For example, since the image change lags behind the instrument trigger, the peak offset of the weight distribution is analyzed to preliminarily estimate the possible time offset of the image stream relative to the instrument state stream.
[0062] An adaptive Kalman filter is used to smooth, predict, and optimize the time offset. Specifically, the Kalman filter defines the system state as the delay of each mode relative to a reference clock and its rate of change. For example, a state vector can be represented as... ,in and These represent the instantaneous delays of the instrument state flow and the operation phase flow relative to the reference image flow, respectively. and This is represented by the rate of change corresponding to the delay; in each processing cycle The adaptive Kalman filter sequentially performs the following two computational processes: state prediction, which is based on the state estimate from the previous time step. Calculate the prior state estimate using the preset state transition matrix F. Covariance of prior error The state transition matrix F models the assumption that the delay changes at a constant rate within a short time window; that is, it is assumed that the delay changes at a constant rate within a short time. State update involves introducing the observation vector at the current moment. That is, the data estimated from the multimodal data alignment stage. and The estimated time offset is calculated by taking the Kalman gain from the observations. With prior prediction The residuals between the two are weighted and fused to output the posterior state estimate. With posterior error covariance This completes one iteration; the two calculation processes are continuously executed recursively, and finally the optimal delay estimate of each mode at the current time is output by combining the optimal estimate.
[0063] It should be noted that, in order to achieve optimal tracking of the dynamic environment, a recent observation residual window is maintained, and statistical properties such as covariance of the residuals within this window are calculated. The estimated value of the observation noise covariance matrix is dynamically updated. That is, when the scene is stable and the observation quality is high, the estimated value of the observation noise covariance matrix decreases, indicating that the filter increases its confidence in the current observation; when the scene changes drastically and the observation noise increases, the estimated value of the observation noise covariance matrix increases accordingly, indicating that the filter relies more on the internal state prediction model, thereby smoothing abnormal disturbances and maintaining estimation stability. At the same time, the process noise covariance matrix is fine-tuned according to the historical dynamics of state changes.
[0064] After obtaining the optimal delay estimate of the Kalman filter output, a weighted fusion algorithm is used to perform weighted fusion of multimodal data, combining circular buffer management and a weighted fusion algorithm. Specifically, two circular buffers are maintained for alternating read and write operations. The original data of the original dual-channel endoscopic images and environmental state data are uniformly remapped to a virtual global time axis based on the image stream, according to their inherent timestamps and the corresponding delay compensation values in the optimal delay estimate. For a target output time point, multiple data samples that are closest to the target output time point on the remapped time axis are retrieved from the circular buffers of each modality. A Gaussian kernel function is used to calculate the absolute difference between the timestamp of the data sample and the target output time point to obtain the corresponding fusion weight. The smaller the absolute difference, the higher the fusion weight. Based on the fusion weight, the data of each modality is independently weighted and averaged to generate the fused data of that modality at the current time point.
[0065] Finally, the time-corrected dual-channel endoscope images and the corresponding alignment environment state data are obtained, and the corrected dual-channel endoscope images and the corresponding alignment environment state data are output as a multimodal endoscope dataset.
[0066] Step S2: Use a pre-trained physical perception hierarchical decoupling network to decouple the features of the multimodal endoscope dataset to obtain a multidimensional decoupled dataset.
[0067] Specifically, the physical perception hierarchical decoupling network performs forward propagation and feature decoupling on the multimodal endoscope dataset to generate structural feature maps and functional feature maps. Monte Carlo Dropout is used to perform multiple forward propagations to obtain the forward propagation results. Uncertainty quantification is performed based on the forward propagation results to generate a confidence map of the decoupling results.
[0068] It should be noted that the physical perception hierarchical decoupling network is represented as a deep convolutional neural network that follows an encoding, decoupling, and decoding architecture, and has an embedded differentiable physical model as an intrinsic constraint. It also includes at least a hierarchical feature decoupling submodule and a differentiable physical rendering submodule.
[0069] The hierarchical feature decoupling submodule is used to extract shared features from the input multimodal endoscopic dataset and decode structural and functional feature maps from the shared features using two parallel branch networks. Specifically, the hierarchical feature decoupling submodule includes a shared feature encoder and a hierarchical decoupling head. The shared feature encoder is composed of a lightweight U-Net deformable and is used to receive calibrated dual-channel endoscopic images. Through a series of convolutional and pooling layers, it extracts a multi-scale shared feature pyramid from the calibrated dual-channel endoscopic images. This multi-scale shared feature pyramid contains at least potential anatomical clues such as edges and textures, as well as physiological functional clues such as vascular patterns and fluorescence signal distribution. The hierarchical decoupling head consists of two parallel, structurally symmetrical but functionally different branch networks: a structural decoding branch and a functional decoding branch. Each branch consists of several residual convolutional blocks and progressive upsampling layers. The structural decoding branch is responsible for reconstructing structural feature maps representing tissue geometry and surface texture from the shared features, while the functional decoding branch is responsible for reconstructing functional feature maps representing physiological activities or the distribution of specific molecular markers from the shared features.
[0070] It should be noted that after the structural decoding branch and the functional decoding branch included in the hierarchical decoupling head, there is an independent small fully connected network. This fully connected network is used to map high-dimensional features to a low-dimensional common embedding space. In this space, the contrastive loss function is optimized to maximize the mutual information of structural and functional embedding vectors from the same spatial location. That is, it encourages the structural decoding branch and the functional decoding branch to share necessary contextual information, while minimizing the similarity of unrelated position vectors, thereby improving the accuracy of feature decoupling.
[0071] The differentiable physical rendering submodule is represented as a multilayer perceptron based on a bidirectional reflectance distribution function, used to impose physical realism constraints on structural features. Specifically, this module is represented as a differentiable bidirectional reflectance distribution function parameterized by the multilayer perceptron, whose input includes at least the following parameters: spatial position, i.e., the normalized image coordinates corresponding to the corrected dual-channel endoscopic images; surface normal vector, used to describe the orientation of the microscopic surface, for example, the corresponding surface normal vector is obtained by calculating the structural feature map based on the predicted depth map or by using the Sobel operator and differentiable integration; and illumination direction. λ is a three-dimensional unit vector used to characterize the main direction of the endoscope light source; spectral bands are scalar or one-hot vectors used to distinguish the physical characteristics of different imaging channels. For example, λ=0 represents the visible light channel, used to reflect surface reflection, and λ=1 represents the near-infrared fluorescence channel, used to reflect subsurface scattering and fluorescence absorption and emission; tissue optical parameter set is a vector retrieved from a preset tissue optical property knowledge base based on the current scene, such as gastric mucosa or adipose tissue. This vector contains key parameters of the tissue type in the relevant bands, such as basic reflectivity, scattering coefficient, and absorption coefficient.
[0072] The differentiable physical rendering submodule, based on the above parameters, renders each pixel in the image... With the corresponding channel Generate the theoretical reflectance under given illumination and viewing direction. The pixel synthesis function is used to synthesize pixels and generate corresponding rendered dual-channel endoscope images. The pixel synthesis function can be expressed as follows: ,in, Represented as the composite pixel value; The intensity of the structural feature map at this spatial location is used to characterize the reflection contribution of the geometric surface; The intensity of the functional feature map at this spatial location is used to characterize the emission contribution of internal functional signals such as fluorescence; and These are channel-dependent scaling factors and mixing coefficients used to model camera gain, spectral response crosstalk, etc.
[0073] A phased learning strategy and a preset joint loss function are employed to train the physical perception hierarchical decoupling network until a preset training convergence condition is met. Specifically, firstly, on a large-scale dataset of natural and synthetic medical images, a data-driven reconstruction loss is used to pre-train the shared feature encoder and the hierarchical decoupling head. This stage employs a joint loss function combining L1 reconstruction loss and structural similarity loss. Next, physical model constraints and contrastive learning constraints are introduced, and secondary training is performed on a real historical endoscopic image dataset. The loss used in this stage can be expressed as... ,in, To correct the reconstruction loss of dual-channel endoscopic images and ensure the fidelity of the decoupling process; For physical consistency loss, a differentiable physical rendering submodule is used to generate corresponding rendered dual-channel endoscope images for the calibrated dual-channel endoscope images, and loss calculations are performed between these rendered dual-channel endoscope images and the calibrated dual-channel endoscope images. For example, according to... Calculate the corresponding loss, where and 2. To correct the dual-channel endoscopic images, and To render dual-channel endoscopic images; To compare the mutually exclusive loss, for a batch of training samples, spatial locations are randomly sampled, positive sample pairs are constructed based on the embedding vectors of the structural feature map and the functional feature map at the same location, and negative sample pairs are constructed based on the embedding vectors of the structural feature map and the functional feature map at different locations or different samples. The InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, thereby keeping the structural feature map and the functional feature map independent. For structural feature maps The applied total variational regularization term is used to encourage its piecewise smoothness; This is for the functional feature map An L1 sparse regularization term is applied to encourage non-zero responses only in specific regions such as blood vessels and lesions, thereby ensuring that the functional feature map conforms to the locality of the functional signal. This is the preset weighting coefficient for the corresponding loss term.
[0074] An adaptive moment estimation optimizer is used, and it is trained using a cosine annealing strategy until the total loss is reached. Training is considered complete when the decrease in the value of the feature map is less than a preset threshold within 20 consecutive training cycles, and the edge preservation index of the structural feature map, the overlap between the functional feature map and the region annotated by the expert, and the mutual information estimation of the decoupling fluctuate less than a preset threshold within 10 consecutive cycles. Alternatively, training is terminated when the preset maximum training cycle is reached.
[0075] Furthermore, the region within the confidence map of the decoupling result that is lower than the preset confidence threshold is defined as the feature fuzzy region. A multi-hypothesis decoupling set is generated based on the forward propagation result corresponding to the feature fuzzy region. The quality of the multi-hypothesis decoupling set is evaluated based on the preset hypothesis quality evaluation function to obtain the hypothesis quality evaluation result.
[0076] Furthermore, based on the hypothetical quality assessment results, the structural feature map and functional feature map are optimized to generate enhanced structural feature map and enhanced functional feature map. The enhanced structural feature map and enhanced functional feature map, along with the decoupling result confidence map, are then encapsulated into a multidimensional decoupling dataset for output.
[0077] In one possible embodiment, the calibrated dual-channel endoscopic image of the current frame is input into a physical perception hierarchical decoupling network. For the same frame input, the physical perception hierarchical decoupling network performs T Monte Carlo Dropout-based forward propagation iterations. Each forward propagation generates a set of structural feature maps and functional feature maps, thereby obtaining a forward propagation set. The standard deviation of each pixel in this set on the structural feature map and functional feature map is calculated. and And combined with preset weighting coefficients for the same pixel position and Weighted fusion is performed, and Gaussian normalization is applied to normalize the weighted values to generate a comprehensive decoupling confidence value. A decoupling result confidence map is generated based on the comprehensive decoupling confidence value at each pixel location. All regions in the decoupling result confidence map that are below a preset confidence threshold are defined as feature fuzzy regions. For the feature fuzzy regions, the corresponding structural and functional feature values in the T forward propagations are extracted to construct a local multi-hypothesis decoupling set. Each fuzzy pixel corresponds to a hypothesis set, which contains T possible structure-function value pairs.
[0078] The quality of the multi-hypothesis decoupling set is evaluated based on a preset hypothesis quality evaluation function, which can be expressed as follows: ,in, Spatial consistency scores can be calculated by assessing the gradient continuity between each structure-function value pair within the multi-hypothesis decoupling set and the decoupling result with the surrounding high-confidence regions, such as the smoothness difference between the current pixel and the adjacent high-confidence pixel values. The smaller the difference, the higher the score. The time consistency score can be obtained by calculating the negative exponent of the difference between the structure-function value pair in the current frame's multi-hypothesis decoupling set and the decoupling result at the corresponding position in the historical frame, so as to achieve the effect that the smoother the change, the higher the score. To obtain the prior conformity score, the function value of the structure-function value pair can be evaluated based on the current tissue type to determine whether it conforms to the typical distribution pattern of the tissue. For example, a pre-trained vascular network detector can be used to determine whether the assumed vascular morphology is reasonable. The weight coefficients for the corresponding terms are used to obtain the hypothesis quality assessment score for each hypothesis in the local multi-hypothesis decoupling set corresponding to each fuzzy region.
[0079] A weighted fusion strategy is employed to weight and fuse the forward propagation results corresponding to each pixel. Specifically, the final value of each pixel is obtained by weighted averaging of its T hypotheses based on their hypothesis quality assessment scores. The hypothesis quality assessment scores of each hypothesis within the multi-hypothesis decoupling set are mapped to corresponding weighting coefficients using the Softmax function. The structural and functional values at corresponding positions in the structural and functional feature maps obtained from the T forward propagations are then weighted and fused according to these weighting coefficients to generate the enhanced value for that pixel. For pixels in non-feature-ambiguous regions, the mean of the T forward propagations is directly used as the final value, generating corresponding enhanced structural and functional feature maps. The enhanced structural feature map represents the tissue anatomical structure characterization optimized by multiple hypotheses, achieving suppression of functional signal interference and transient artifacts. The enhanced functional feature map represents the physiological functional signal distribution optimized by multiple hypotheses, achieving elimination of the obfuscation effect of structural texture.
[0080] For example, when processing an intestinal image with motion blur, multiple decoupling results are first obtained through 10 Dropout forward propagations. Calculations show that the confidence level of the intestinal villus edge region is low. After extracting 10 hypotheses for this region, the hypothesis quality assessment function is used to calculate that 3 of the hypotheses score highly in terms of spatial continuity, temporal consistency, and anatomical rationality. That is, they satisfy the requirements of being consistent with the clear villus morphology, matching the villus motion trajectory of the previous frame, and conforming to villus morphological features, respectively. These 3 high-quality hypotheses are then fused using softmax weighting (the reason for weighting 3 high-quality hypotheses here is that the high-quality hypotheses have high hypothesis quality assessment scores, so their weight coefficients after mapping are also high; the other low-quality hypotheses have low scores and their corresponding weight coefficients are also low and can be ignored). The final output enhanced feature map maintains structural continuity at the villus edge and accurately separates the functional signals between villi.
[0081] Finally, the enhanced structural feature map, the enhanced functional feature map, and the decoupling result confidence map are encapsulated into a multidimensional decoupling dataset for output.
[0082] Step S3: Using neural implicit surface technology and meta-learning algorithm, multidimensional scene reconstruction is performed simultaneously with the multidimensional decoupled dataset to generate a reconstructed scene multidimensional dataset.
[0083] Specifically, an image scene reconstruction model is constructed based on neural implicit surface technology and meta-learning algorithm. The image scene reconstruction model includes a scene representation network and a context information encoder.
[0084] Understandably, the scene representation network is used to represent the geometric and appearance features of endoscopic images using neural implicit surface techniques to generate a scene representation set. Specifically, the scene representation network includes a symbolic distance function field and a conditional radiation field constructed with a multilayer perceptron. The scene representation network takes the three-dimensional spatial coordinates of the query point, the encoding vector of the observation direction, the corresponding context information vector, and a rendering mode label as input.
[0085] The encoded vectors of three-dimensional spatial coordinates and viewing direction are used to capture the basic spatial and viewpoint dependencies. The encoded vectors can be encoded using hash encoding techniques. For example, multi-resolution hash encoding techniques can be used to divide the three-dimensional space into multiple grids. Each grid vertex stores a learnable feature vector. For any query point, its multi-scale features are quickly obtained through trilinear interpolation, and the corresponding encoded vector is output.
[0086] The context information vector serves as a conditional input, controlling the network to output different scene attributes based on different operational states.
[0087] The rendering mode label is represented as a scalar with a value range of [0,1]. This label is only enabled in step S4 when the final fused image is generated. It serves as an independent input dimension to control the modality of the final rendering output. When the value of the rendering mode label is 0, the scene representation network should render a pure structural scene, i.e., an image containing only tissue anatomical textures and no functional markers. When the value of the rendering mode label is 1, the scene representation network should render a pure functional scene, i.e., an image that highlights functional signals and suppresses natural textures. When the value of the rendering mode label is within (0, 1), the scene representation network should render an intermediate transition state with the corresponding ratio. For example, if the value of the rendering mode label is 0.4, the final rendered image should be presented with a ratio of 0.4 * pure structural scene + 0.6 * pure functional scene.
[0088] The main framework of the scene representation network is a multilayer perceptron. Its processing flow is divided into two logical branches: a symbolic distance function branch and a conditional radiation field branch. This branch concatenates the encoded vector of the three-dimensional spatial coordinates with the context information vector and inputs it into a shared backbone multilayer perceptron to map from conditional coordinates to the global geometry of the scene. Then, it combines the symbolic distance function to output a symbolic distance value, which represents the signed distance from the query point to the nearest tissue surface. Positive values represent the outside, negative values represent the inside, and zero values are defined as being located on the surface, thus realizing an implicit geometric representation of the scene. The conditional radiation field branch takes the intermediate geometric features extracted by the shared backbone multilayer perceptron, the encoded vector of the viewing direction, and the context information vector and inputs them into a dedicated color and feature output head. This head outputs the RGB color code and high-dimensional semantic feature vector of the corresponding point. Finally, the scene representation network outputs a scene representation set containing the symbolic distance value, RGB color code, and high-dimensional semantic feature vector.
[0089] For example, when querying a point on the surface of the liver, the scene representation network takes its spatial coordinates and contextual information in the separation phase as input. The symbolic distance function field branch outputs a value close to zero to indicate that it is currently located on the tissue surface, while the conditional radiation field branch outputs the normal liver tissue color and feature vectors representing soft tissue at that point. If the contextual information changes to "after electrocoagulation", the output of the conditional radiation field branch for that spatial point may change to the dark brown of charred tissue and features representing "thermal damage", while its output symbolic distance value remains unchanged, thus achieving dynamic and conditional representation of appearance and function.
[0090] It should be noted that, in order to obtain the corresponding two-dimensional image from the scene representation set, the scene representation network adopts a differentiable volume rendering method. For each pixel on the image, a ray is emitted from the center of the camera, and multiple points are sampled along the ray. The scene representation set corresponding to these points is obtained. Through the differentiable volume rendering formula, the symbolic distance value is converted into density, and the color and semantic features corresponding to the RGB color code and the high-dimensional semantic feature vector are integrated along the ray. Finally, the predicted color and feature map of the pixel are synthesized. During training, by comparing the difference between the synthesized image and the real image, the backpropagation algorithm is used to optimize the parameters corresponding to the two branches.
[0091] The context information encoder is used to encode the multidimensional decoupled dataset into a context information vector. Specifically, the context information encoder is represented as a lightweight Transformer encoder. Its input includes aligned environmental state data within the multimodal endoscopy dataset and enhanced structural feature maps and enhanced functional feature maps of the multidimensional decoupled dataset. The encoder captures the complex relationships between the input data through self-attention and cross-attention mechanisms and outputs a context information vector. This vector is used to represent contextual semantic information such as "what operation is currently in progress, what tools are being used, how the tissue moves, and where functional signals are significant".
[0092] A phased reconstruction training strategy is adopted to train the image scene reconstruction model, and a meta-learning algorithm is used for parameter fine-tuning. Specifically, a phased course learning strategy is adopted, dividing the training process into the following two stages: a basic scene reconstruction pre-training stage, in which a dataset containing multi-view endoscopic image sequences and registered depth map ground truth values is used as training data, trains the image scene reconstruction model to reconstruct physically plausible geometric surfaces and appearances from monocular or multi-view images. The loss function used in this stage can be expressed as: ,in, The photometric reconstruction loss term is obtained by calculating the L1 loss and structural similarity loss between the rendered image and the real image derived from the model. The geometric consistency loss term is obtained by calculating the L1 loss between the rendered depth map corresponding to the symbolic distance value output by the field branch of the symbolic distance function and the corresponding depth map of the sensor. For example, the Eikonal regularization loss and the segmentation mask loss; , and These are the weighting coefficients for the corresponding loss terms.
[0093] It should be noted that, because the input to the scene representation network includes a rendering mode label, during the training phase, The value is the sum of the reconstruction losses of the two rendering modes, i.e. ,in The reconstruction loss is calculated in RGB space based on the corresponding enhanced structure feature map of the training data. Similarly, The reconstruction loss is calculated in RGB space based on the corresponding augmented feature map of the training data; through this training method, the scene representation network is forced to learn to switch rendering modes according to the value of the rendering mode label.
[0094] In the meta-learning fine-tuning stage, the training data used in this stage consists of endoscopic video sequences with rich annotations, such as stage labels, instrument status, and tissue motion annotations. A model-independent meta-learning algorithm is employed. Within the meta-training loop, each independent anatomical site is treated as a meta-task. For each meta-task, a small number of frames, such as 5 frames, are sampled as a support set. The image, depth, and corresponding contextual information vectors of the support set are input into the image scene reconstruction model, and the loss is calculated. The model performs several gradient descent steps on the corresponding parameters to obtain parameters that are quickly adapted for the task. In the outer loop of the meta-training, the initial parameter set of the image scene reconstruction model is optimized based on a large number of meta-tasks so that the sum of the losses on their respective query sets is minimized after all tasks are quickly adapted in the inner loop. The loss used in this stage is the loss function used in the pre-training stage of the basic scene reconstruction.
[0095] The training ends when the model's rendering quality improves to a preset threshold after a fixed number of iterations on a reserved, unseen validation task set, or when the total loss of the validation set does not decrease within a preset maximum number of consecutive training epochs. If the total loss of the validation set does not decrease within 15 consecutive epochs, early stopping is triggered to save the best-performing model parameters.
[0096] Furthermore, the multidimensional decoupled dataset is input into the image scene reconstruction model for multidimensional scene reconstruction. After the multidimensional decoupled dataset is input into the image scene reconstruction model, the context information encoder encodes the multidimensional decoupled dataset to generate a context information vector.
[0097] Furthermore, an inner loop update mechanism based on model-independent meta-learning is adopted to fine-tune the parameters of the scene representation network. The fine-tuned scene representation network, based on the context information vector, simultaneously combines with a depth estimation algorithm to reconstruct the scene and generate a reconstructed scene representation set. The context information vector and the reconstructed scene representation set are then encapsulated into a reconstructed scene multidimensional dataset for output.
[0098] In one possible embodiment, the aligned environment state data in the multimodal endoscopy dataset and the enhanced structural feature map and enhanced functional feature map of the multidimensional decoupled dataset are input into the context information encoder to generate the corresponding context information vector. In the initial stage of endoscopic image acquisition, such as the first 30 seconds or the first 100 frames after the start of acquisition, the image frames acquired in the initial stage, the corresponding depth estimates, enhanced structural feature maps and enhanced functional feature maps, and the context information vector are cached and constructed into a support set.
[0099] Starting with the initial parameters of the pre-trained image scene reconstruction model, the support set data is input into the model, the reconstruction loss of the corresponding parameters is calculated, and a small number of gradient descent update operations are performed in combination with the reconstruction loss to obtain adaptive model parameters for the current patient. For example, for a patient who is obese and has thick peritoneal fat, the general model may reconstruct a tissue surface that is too thin. By using the initial few frames of endoscopic images and related data of the patient, rapid learning and parameter fine-tuning are performed to more accurately adapt to the individual characteristics of the current patient.
[0100] The scene representation network of the adjusted image scene reconstruction model outputs the corresponding scene representation set based on the corresponding input data. At this time, the output scene representation set is the corresponding reconstructed scene representation set.
[0101] Step S4: Based on the adaptive dual-path fusion mechanism, the dual-path endoscopic images are fused simultaneously with the reconstructed scene multidimensional dataset to generate enhanced endoscopic images.
[0102] Specifically, inference is performed based on the reconstructed scene multidimensional dataset and the multidimensional decoupled dataset to generate a spectral mixing weight map and a local enhancement parameter map. Based on the spectral mixing weight map and the reconstructed scene representation set contained in the reconstructed scene multidimensional dataset, a basic fused endoscope image is generated. The basic fused endoscope image is then locally optimized according to the local enhancement parameter map to generate an enhanced endoscope image.
[0103] In one possible embodiment, a pre-trained semantic fusion strategy network is employed. This network uses an encoder-decoder architecture, where the encoder fuses the reconstructed scene multidimensional dataset and the multidimensional decoupled dataset, encoding them into a comprehensive scene state vector. The decoder then generates a spectral mixing weight map and a local enhancement parameter map based on this comprehensive scene state vector. Specifically, the data is first preprocessed. For the context information vector within the reconstructed scene multidimensional dataset, a fully connected layer is used for mapping, forming a global semantic conditional tensor that matches the image space size. For the reconstructed scene representation set within the reconstructed scene multidimensional dataset, a multi-channel feature map with the same resolution as the output image is rendered using differentiable rendering technology. The image is defined as a scene semantic feature map. For the enhanced structural feature map and enhanced functional feature map of the multidimensional decoupled dataset, the two are concatenated into a two-channel image, and a shallow convolutional network such as two 3×3 convolutional layers is used for feature extraction to generate a spatial detail feature map. For the confidence map of the decoupling result of the multidimensional decoupled dataset, after activation and normalization by a sigmoid function, it is directly used as a spatial reliability mask to dynamically weight the information flow from the spatial detail feature map and the scene semantic feature map during the feature fusion stage. For example, in low confidence regions, the information contribution of the enhanced structural feature map and the enhanced functional feature map will be suppressed, and the semantic fusion strategy network will rely more on the information from the reconstructed scene multidimensional dataset.
[0104] Next, the four types of data processed above are input into a lightweight visual Transformer encoder. This encoder performs multi-dimensional feature fusion through self-attention and cross-modal attention, outputting a comprehensive scene state vector. For example, a multi-scale feature pyramid architecture is used for feature fusion. In the high-resolution feature layer of the encoder, spatial detail feature maps and scene semantic feature maps are aligned and fused through element-wise addition and weighting operations based on spatial reliability masks, realizing the integration of detail information and geometric priors. In the medium-resolution feature layer, spatial adaptive normalization is used to inject the global semantic condition tensor into the feature map output by the high-resolution feature layer to achieve global semantic alignment. In the low-resolution feature layer, a multi-head cross-attention mechanism is used to perform global inference on the feature map output by the medium-resolution feature layer to generate a comprehensive scene state vector.
[0105] The decoder, through a series of upsampling and skip connections to the corresponding layers of the encoder, gradually restores the spatial resolution and decodes the comprehensive scene state vector output by the encoder into a spectral blending weight map and a local enhancement parameter map. The value of each pixel in the spectral blending weight map is the value of the rendering mode label of the input scene representation network. The local enhancement parameter map is a multi-channel map, and each channel defines the intensity of different post-processing operations such as edge sharpening intensity, contrast stretching range, and specific color gain applied at the corresponding pixel position.
[0106] Based on the rendering mode label value of each pixel in the spectral blending weight map Rerun the scene representation network to obtain the corresponding pure structural properties. Scenarios with pure functional attributes , That is, the set of reconstructed scene representations without considering functional signals; similarly... This refers to the set of reconstructed scene representations without considering structural signals, where the base fused pixel value corresponding to that pixel is... This generates a basic fused endoscopic image, for example, the image corresponding to a certain pixel. and The RGB color codes in the image represent pale yellow, the natural color of the tissue, and bright cyan, representing blood flow function. Furthermore, this point... If the value is 0.3, then the final color of this point is 0.3 * pale yellow + 0.7 * bright cyan, thus presenting a color dominated by cyan with a slightly textured background.
[0107] Each channel of the local enhancement parameter map is analyzed, and the base fused endoscope image is locally optimized based on the analyzed control parameters to finally generate the corresponding enhanced endoscope image. For example, assuming that the local enhancement parameter map contains two operations: adaptive edge sharpening and local contrast stretching, for the adaptive edge sharpening operation, Laplacian filtering or unsharpening masks of appropriate intensity are applied to the corresponding pixel areas according to the values of the corresponding channels in the local enhancement parameter map to enhance tissue edges and instrument contours; for the local contrast stretching operation, linear or nonlinear contrast transformations with different slopes are performed on different regions of the image according to the values of the corresponding channels, so that dark details and bright details are visible at the same time.
[0108] An electronic device, comprising:
[0109] At least one processor; and at least one memory communicatively connected to the processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method proposed in the embodiments of the present invention.
[0110] The following is a detailed introduction to the various components of the electronic device:
[0111] In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement Embodiment 1 of this invention, such as one or more digital signal processors (DSPs) or one or more field-programmable gate arrays (FPGAs).
[0112] The processor can perform various functions of an electronic device by running or executing software programs stored in memory and by calling data stored in memory.
[0113] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.
[0114] The memory can be a real-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through an interface circuit of an electronic device; this embodiment of the invention does not specifically limit this.
[0115] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via limited means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0116] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0117] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0118] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for maintaining global background motion in an endoscope, characterized in that, It includes the following steps: The raw dual-channel endoscope images and environmental state data are acquired, and data preprocessing is performed by combining attention mechanism and adaptive filtering to generate a multimodal endoscope dataset. A pre-trained physical perception hierarchical decoupling network is used to decouple the features of the multimodal endoscope dataset to obtain a multidimensional decoupled dataset; The neural implicit surface technique and meta-learning algorithm are used to simultaneously combine the multidimensional decoupled dataset to reconstruct multidimensional scenes and generate a reconstructed scene multidimensional dataset. Based on the adaptive dual-path fusion mechanism, dual-path endoscopic images are fused simultaneously with the reconstructed scene multidimensional dataset to generate enhanced endoscopic images.
2. The method for maintaining global background motion in an endoscope according to claim 1, characterized in that, Raw dual-channel endoscopic images and environmental state data are acquired, and data preprocessing is performed using an attention mechanism and adaptive filtering to generate a multimodal endoscopic dataset, including: The original dual-channel endoscope images are represented as two original Bayer images of the endoscope in two different wavelength bands, and the environmental status data includes at least the operation stage label and instrument status data. A spatiotemporal convolutional network is used to extract features from the original dual-channel endoscope images and environmental state data to obtain endoscope image feature vectors and endoscope state feature vectors. Based on the endoscope image feature vector and the endoscope state feature vector, temporal offset inference is performed to obtain the temporal offset; Adaptive Kalman filtering is used, combined with the time offset, to perform time offset compensation and time alignment, and to obtain the corrected dual-channel endoscope images and the corresponding alignment environment state data. The corrected dual-channel endoscopic images and the corresponding aligned environmental state data are output as a multimodal endoscopic dataset.
3. The method for maintaining global background motion in an endoscope according to claim 1, characterized in that, A pre-trained physical perception hierarchical decoupling network is used to decouple features from the multimodal endoscope dataset to obtain a multidimensional decoupled dataset, including: The physical perception hierarchical decoupling network performs forward propagation and feature decoupling on the multimodal endoscope dataset to generate structural feature maps and functional feature maps; Monte Carlo Dropout is used to perform multiple forward propagations to obtain the forward propagation results; uncertainty quantification is performed based on the forward propagation results to generate a confidence graph of the decoupling results; The region within the confidence map of the decoupling result that is lower than the preset confidence threshold is defined as the feature fuzzy region, and a multi-hypothesis decoupling set is generated based on the forward propagation result corresponding to the feature fuzzy region. The quality of the multi-hypothesis decoupling set is evaluated based on a preset hypothesis quality evaluation function to obtain the hypothesis quality evaluation result. Based on the hypothetical quality assessment results, the structural feature map and functional feature map are optimized to generate an enhanced structural feature map and an enhanced functional feature map. The enhanced structural feature map, the enhanced functional feature map, and the decoupling result confidence map are encapsulated into a multidimensional decoupling dataset for output.
4. The method for maintaining global background motion in an endoscope according to claim 3, characterized in that, A hierarchical decoupling network for physical perception includes: The physical perception hierarchical decoupling network includes a hierarchical feature decoupling submodule and a differentiable physical rendering submodule; The hierarchical feature decoupling submodule is used to extract shared features from the input multimodal endoscopy dataset, and combine two parallel branch networks to decode structural feature maps and functional feature maps from the shared features; The differentiable physical rendering submodule is represented as a multilayer perceptron based on a bidirectional reflection distribution function, used to apply physical realism constraints to structural features. The physical perception hierarchical decoupling network is trained using a phased learning strategy and a preset joint loss function until the preset training convergence condition is met.
5. The method for maintaining global background motion in an endoscope according to claim 1, characterized in that, Using neural implicit surface techniques and meta-learning algorithms, multidimensional scene reconstruction is performed simultaneously with the aforementioned multidimensional decoupled dataset, generating a reconstructed scene multidimensional dataset, including: An image scene reconstruction model is constructed based on neural implicit surface technology and meta-learning algorithm. The image scene reconstruction model includes a scene representation network and a context information encoder. The scene representation network is used to represent the geometric and appearance features of endoscopic images using neural implicit surface techniques, generating a scene representation set. The context information encoder is used to encode the multidimensional decoupled dataset into a context information vector; A phased reconstruction training strategy is adopted to train the image scene reconstruction model, and the parameters are fine-tuned by combining a meta-learning algorithm.
6. The method for maintaining global background motion in an endoscope according to claim 5, characterized in that, The method further includes: Input the multidimensional decoupled dataset into the image scene reconstruction model to perform multidimensional scene reconstruction; After the multidimensional decoupled dataset is input into the image scene reconstruction model, the context information encoder encodes the multidimensional decoupled dataset to generate a context information vector. An inner loop update mechanism based on model-independent meta-learning is used to fine-tune the parameters of the scene representation network; The fine-tuned scene representation network, based on the context information vector, simultaneously combines a depth estimation algorithm to reconstruct the scene and generate a set of reconstructed scene representations. The context information vector and the reconstructed scene representation set are encapsulated into a reconstructed scene multidimensional dataset for output.
7. The method for maintaining global background motion in an endoscope according to claim 1, characterized in that, Based on an adaptive dual-path fusion mechanism, dual-path endoscopic images are fused simultaneously using the reconstructed scene multidimensional dataset to generate enhanced endoscopic images, including: Inference is performed based on the reconstructed scene multidimensional dataset and multidimensional decoupled dataset to generate a spectral mixing weight map and a local enhancement parameter map; Based on the spectral mixing weight map and the set of reconstructed scene representations contained in the reconstructed scene multidimensional dataset, a basic fused endoscopic image is generated; The base fused endoscope image is locally optimized based on the local enhancement parameter map to generate an enhanced endoscope image.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and at least one memory communicatively connected to the processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform an endoscopic global background motion holding method as described in any one of claims 1 to 7.