ERCP image intelligent identification system based on multi-modal fusion

The multi-modal ERCP image recognition system addresses subjective analysis in ERCP by automating parameter measurement and reducing radiation exposure, enhancing procedural accuracy and efficiency.

CN120318207AInactive Publication Date: 2025-07-15THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Patent Information

Application Number
CN202510489791.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional ERCP image analysis relies on physician experience, lacks quantitative standards, cannot automatically quantify key parameters, and X-ray fluoroscopy extends the accumulated dose of radiation, making it difficult to achieve real-time early warning and auxiliary decision-making.

Method used

The ERCP image intelligent recognition system with multimodal fusion is adopted, including image acquisition, preprocessing, multimodal fusion model, parameter measurement and real-time interaction module. Through private information desensitization, image distortion correction, dynamic range optimization, multi-scale feature extraction and augmented reality display, precise segmentation of anatomical structures and parameter quantization are achieved.

Benefits of technology

It improves the objectivity and accuracy of ERCP image analysis, reduces the accumulated radiation dose, improves the efficiency of intraoperative decision-making, and ensures automatic quantification and real-time early warning of key parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318207A_ABST
    Figure CN120318207A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of medical image processing and intelligent diagnosis, discloses an ERCP image intelligent identification system based on multi-modal fusion, comprising an image acquisition module, an image preprocessing module, a multi-modal fusion model module, a parameter measurement module, an early warning judgment module and a real-time interaction module. The image acquisition module is used for acquiring an ERCP radiography original image in a DICOM format, performing privacy information desensitization processing on the original image, and reserving pixel pitch and catheter size parameters in DICOM metadata; the system has the beneficial effects that privacy protection and parameter retention are realized through the image acquisition module, precise structure segmentation is realized in combination with a multi-modal fusion model, key indexes are automatically quantized by means of the parameter measurement module, intraoperative visualization is enhanced in cooperation with the real-time interaction module, and the accuracy of the system is improved. The problems that traditional ERCP analysis depends on experience and lacks quantitative standards and decision lags are effectively solved, and the method has the remarkable advantages that diagnosis precision is improved, operation time is shortened, and the medical resource utilization rate is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing and intelligent diagnosis, and particularly relates to an intelligent recognition system for ERCP images based on multimodal fusion. Background Art

[0002] ERCP, i.e., endoscopic retrograde cholangiopancreatography, is a general term for inserting an endoscope through the mouth into the descending part of the duodenum, introducing special instruments through the duodenal papilla into the bile duct or pancreatic duct, injecting contrast agent for imaging under X-ray fluoroscopy, introducing a sub-endoscope / ultrasound probe for observation, collecting exfoliated cells / tissues, etc., completing the diagnosis of biliary and pancreatic diseases, and performing corresponding interventional treatments based on the diagnosis. It is an important technique for diagnosing and treating biliary and pancreatic system diseases and is widely used in the clinical treatment of diseases such as common bile duct stones, biliary strictures, and pancreatic duct dilation.

[0003] Traditional ERCP image analysis has significant professionalism and subjectivity. Intraoperative judgment mainly relies on the intuitive interpretation of X-ray angiography images by physicians, and there are the following defects: due to the complex anatomical structure and variable morphology of the bile and pancreatic ducts, and the low image contrast, clinicians need to have rich experience to accurately identify the common bile duct, main pancreatic duct, and lesion area; ERCP examinations only provide imaging information and cannot automatically quantify key parameters such as bile duct length, angle, and diameter, which restricts the classification of diseases and the selection of treatment plans; in the face of difficult intubation, complex stones, etc., it is difficult for doctors to timely obtain the risk level and cannot achieve active early warning and auxiliary decision-making; due to relying on real-time X-ray fluoroscopy to guide the operation, the extended intraoperative time increases the cumulative radiation dose of patients and operators.

[0004] Therefore, an intelligent recognition system for ERCP images based on multimodal fusion is proposed. Summary of the Invention

[0005] The purpose of the present application is to provide an intelligent recognition system for ERCP images based on multimodal fusion, which has the advantages of improving the objectivity of ERCP image analysis, realizing automatic quantification of key parameters, reducing the cumulative radiation dose, and improving the intraoperative decision-making efficiency.

[0006] According to one aspect of the present application, there is provided an intelligent recognition system for ERCP images based on multimodal fusion, including: an image acquisition module, configured to acquire the original ERCP contrast images in DICOM format, perform desensitization processing on the original images, and retain the pixel pitch and catheter size parameters in the DICOM metadata; an image preprocessing module, configured to perform noise reduction, distortion correction based on the catheter size parameters, and dynamic range optimization on the desensitized original images to generate standardized image data; a multimodal fusion model module, configured to extract the segmentation masks and sub-pixel level key anatomical point coordinates of the bile duct and pancreatic duct from the standardized image data; a parameter measurement module, configured to calculate the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct based on the segmentation masks, the sub-pixel key anatomical point coordinates, and the pixel pitch parameters; a warning judgment module, configured to judge the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct according to preset rules to generate a warning signal; and a real-time interaction module, configured to superimpose the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size of filling defects in the bile duct and pancreatic duct on the original images in an augmented reality manner, and perform risk marking according to the warning signal.

[0007] As a preferred solution of an intelligent recognition system for ERCP images based on multimodal fusion of the present application, the image preprocessing module includes: a noise reduction unit, configured to reduce image noise of the desensitized original images by using a non-local means filtering algorithm to generate a noise-reduced image; a gradient calculation unit, configured to extract gray gradient features of the noise-reduced image through an edge detection algorithm; a marker point extraction unit, configured to locate catheter marker points based on the catheter size parameters and the gray gradient features; a distortion correction unit, configured to perform an affine transformation on the noise-reduced image based on the catheter marker points to generate a geometrically corrected image; and a dynamic optimization unit, configured to perform adaptive histogram equalization on the geometrically corrected image to adjust the contrast, and verify whether the sharpness of the adjusted image meets a preset gradient threshold, and output standardized image data.

[0008] As a preferred solution of an ERCP image intelligent recognition system based on multimodal fusion in this application, the multimodal fusion model module includes: a feature extraction unit for performing multi-scale feature extraction on the standardized image data through a pre-trained convolutional neural network to generate feature maps containing different resolution levels; a local feature enhancement unit for enhancing the response of minute anatomical structure features in the feature maps through an attention mechanism; a global relationship modeling unit for capturing multi-scale anatomical context information of the feature maps through atrous spatial pyramid pooling operations; a feature fusion unit for fusing the enhanced minute anatomical structure features and multi-scale anatomical context features to generate fused features; and a segmentation and localization unit for outputting segmentation masks of the bile duct and pancreatic duct and sub-pixel level key anatomical point coordinates based on the fused features.

[0009] As a preferred solution of an ERCP image intelligent recognition system based on multimodal fusion in this application, the segmentation and localization unit includes: a segmentation sub-unit for decoding the fused features to generate segmentation masks of the bile duct and pancreatic duct; and a high-resolution localization sub-unit for extracting high-resolution spatial information of the fused features through an HRNet network and outputting the sub-pixel level key anatomical point coordinates.

[0010] As a preferred solution of an ERCP image intelligent recognition system based on multimodal fusion in this application, the system further includes a dynamic calibration module, and the dynamic calibration module performs: extracting catheter size and pixel pitch parameters from the DICOM metadata; calculating the theoretical pixel length of catheter marker points in the image according to the actual physical length of the catheter size and the pixel pitch parameters; measuring the actual pixel length of catheter marker points based on the standardized image data generated by the image preprocessing module; taking the ratio of the theoretical pixel length to the actual pixel length as a calibration factor; multiplying the calibration factor by the pixel pitch parameters to generate corrected pixel pitch parameters, and inputting the corrected pixel pitch parameters into the parameter measurement module.

[0011] As a preferred solution of an ERCP image intelligent recognition system based on multi-modal fusion in this application, the parameter measurement module includes: a centerline extraction unit for generating the bile duct centerline through a topological thinning algorithm based on the segmentation mask; a direction vector calculation unit for determining the main direction vector of the bile duct centerline through a principal component analysis algorithm based on the key anatomical point coordinates; an angle calculation unit for calculating the included angle between the main direction vectors to generate the distal bile duct angulation; a length calculation unit for performing curve integration along the bile duct centerline and calculating the distal bile duct length in combination with the pixel pitch parameter; a diameter analysis unit for generating a dynamic cross-section along the normal direction of the bile duct centerline and calculating the minimum diameter of the full circumference of the dynamic cross-section as the bile duct diameter in combination with the pixel pitch parameter; and a filling defect detection unit for detecting and calculating the size and quantity of filling defects.

[0012] As a preferred solution of an ERCP image intelligent recognition system based on multi-modal fusion in this application, the filling defect detection unit includes: a density difference comparison subunit for comparing the contrast agent density distribution differences between the segmentation mask and the standardized image data, and identifying the areas where the local density is lower than the preset threshold as candidate filling defect areas; an instance segmentation and verification subunit for performing connected region analysis on the candidate filling defect areas to determine independent candidate filling defect instances, and excluding artifact interference in combination with morphological rules to generate effective filling defect instances; a size annotation and statistics subunit for annotating the maximum transverse diameter of each effective filling defect instance and counting the number of instances as the filling defect quantity; and a size conversion subunit for converting the maximum transverse diameter of each effective filling defect instance into the filling defect size based on the pixel pitch parameter, and the filling defect size being the physical size of the maximum transverse diameter.

[0013] As a preferred solution of an ERCP image intelligent recognition system based on multi-modal fusion in this application, the warning signal generation conditions of the warning judgment module include: the distal bile duct angulation is lower than the preset angulation threshold; the diameter of the bile duct or pancreatic duct exceeds the preset diameter threshold; the filling defect size reaches the preset size threshold; and the filling defect quantity reaches the preset quantity threshold.

[0014] As a preferred solution of an ERCP image intelligent recognition system based on multi-modal fusion in the present application, where the steps of superimposing the bile duct length, angulation degree, diameter, and filling defect size onto the original image in an augmented reality manner and performing risk marking based on the warning signal are specifically as follows: enhancing the display of the bile duct and pancreatic duct in different colors respectively; displaying the bile duct length, angulation degree, diameter, and filling defect size in the form of dynamic floating labels in real time; when a warning signal is detected, based on the segmentation mask and the key anatomical point coordinates, determining the spatial position of the anatomical structure corresponding to the warning signal, and marking this spatial position in the original image with a visual marker as a risk marking.

[0015] As a preferred solution of an ERCP image intelligent recognition system based on multi-modal fusion in the present application, the system further includes: a model optimization module for compressing the number of parameters of the multi-modal fusion model module through a knowledge distillation algorithm.

[0016] Compared with the prior art, by using an ERCP image intelligent recognition system based on multi-modal fusion according to an embodiment of the present application, privacy protection and parameter retention can be achieved through the image acquisition module, accurate structure segmentation can be realized by combining with the multi-modal fusion model, key indicators can be automatically quantified relying on the parameter measurement module, intraoperative visualization can be enhanced with the cooperation of the real-time interaction module, effectively solving the problems of traditional ERCP analysis relying on experience, lacking a quantification standard, and decision lag, and having the significant advantages of improving the diagnostic accuracy, shortening the operation time, and optimizing the utilization rate of medical resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 It is a block diagram of an ERCP image intelligent recognition system based on multi-modal fusion of the present invention.

[0019] Figure 2 It is a block diagram of the image preprocessing module of an ERCP image intelligent recognition system based on multi-modal fusion of the present invention.

[0020] Figure 3 It is a block diagram of the multi-modal feature fusion module of an ERCP image intelligent recognition system based on multi-modal fusion of the present invention.

[0021] Figure 4Block diagram of the parameter measurement module of an intelligent ERCP image recognition system based on multi-modal fusion according to the present invention.

[0022] Figure 5 Block diagram of the filling defect detection unit of an intelligent ERCP image recognition system based on multi-modal fusion according to the present invention.

[0023] Figure 6 Operation result diagram of an intelligent ERCP image recognition system based on multi-modal fusion according to the present invention. Detailed implementation manners

[0024] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0025] Application overview

[0026] In traditional ERCP image analysis systems, there are contradictions between privacy information desensitization and key parameter retention in the DICOM format image processing flow. The geometric distortion correction of the original image relies on manual marking, resulting in spatial registration errors. Insufficient dynamic range optimization causes the loss of micro-anatomical structure features. The lack of a multi-modal data fusion mechanism limits the accuracy of segmentation masks and key point localization, and it is impossible to establish the mapping between the spatial geometric relationship of anatomical structures and physical parameters, thereby causing calculation deviations in quantization indicators such as bile duct length and angulation degree. Intraoperative real-time evaluation relies on manual experience judgment, lacking an early warning signal generation mechanism based on objective rules. The spatial registration error between anatomical parameters and the original image in augmented reality interaction leads to decision-making delays, and the extended operation time exacerbates the radiation accumulation risk.

[0027] For example, in the operation scenario of ERCP for common bile duct stone cases, when the traditional system processes DICOM original images, the privacy desensitization operation will clear the catheter size parameters, resulting in an excessive error in the extraction of the bile duct center line in the subsequent distortion correction link using a fixed scale factor for affine transformation. The high-frequency filtering algorithm in the noise reduction process destroys the gray gradient features of the catheter marking points, making it impossible for the multi-scale feature extraction network to accurately locate the bile-pancreatic duct bifurcation point. The global histogram equalization operation in the dynamic range optimization link reduces the contrast between the filling defect area and normal tissues, and the segmentation model misjudges small-diameter stones as artifacts. In the parameter calculation link, due to the lack of sub-pixel key point coordinates, the measurement error of the bile duct angulation degree is too large, forcing the operator to perform multiple fluoroscopies to confirm the anatomical structure, and the single operation time is extended.

[0028] If the above problems are not solved, the remaining geometric distortion in the image preprocessing stage will lead to the distortion of three-dimensional space reconstruction, affect the measurement accuracy of the length of the bile duct stricture segment, and cause errors in stent size selection. The deviation in the positioning of key anatomical points will miscalculate the angle between the bile and pancreatic ducts, resulting in errors in the endoscopic insertion path planning and increasing the risk of duodenal papilla injury. The insufficient sensitivity of filling defect detection will miss small stones, leading to complications such as postoperative cholangitis. The delay in calculating quantization parameters forces the operator to extend the fluoroscopy time, and the patient's surface radiation dose may exceed the safety threshold by several times. Without a real-time warning mechanism, bile duct dilation lesions with sudden diameter changes cannot be identified in time, delaying the early diagnosis window period of malignant tumors.

[0029] Facing the above problems, this application first considers how to retain the physical parameters crucial for subsequent analysis during the privacy de-sensitization process and attempts to establish a key metadata protection mechanism in the image acquisition stage. For the errors caused by the dependence on manual marking in geometric distortion correction, an automated spatial registration method based on the inherent size parameters of the catheter is explored. To solve the problem of the loss of small structural features caused by insufficient dynamic range optimization, a collaborative processing strategy of contrast enhancement and gradient threshold verification is studied. To achieve accurate quantitative analysis of anatomical structures, a multi-modal feature fusion network is designed to simultaneously improve the accuracy of segmentation masks and key point positioning. Facing the real-time requirements during the operation, an augmented reality interaction technology is developed to achieve the spatial synchronous mapping of image data and anatomical parameters.

[0030] Exemplary system

[0031] Such as Figures 1-5As shown in the figure, an intelligent recognition system for ERCP images based on multimodal fusion according to an embodiment of the present application includes: an image acquisition module, configured to acquire the original ERCP contrast images in DICOM format, perform privacy information desensitization processing on the original images, and retain the pixel spacing and catheter size parameters in the DICOM metadata; an image preprocessing module, configured to perform noise reduction, distortion correction based on the catheter size parameters, and dynamic range optimization on the desensitized original images to generate standardized image data; a multimodal fusion model module, configured to extract the segmentation masks and sub-pixel level key anatomical point coordinates of the bile duct and pancreatic duct from the standardized image data; a parameter measurement module, configured to calculate the length of the distal bile duct, the angle degree of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct based on the segmentation masks, sub-pixel key anatomical point coordinates, and pixel spacing parameters; an early warning judgment module, configured to judge the length of the distal bile duct, the angle degree of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct according to preset rules to generate an early warning signal; a real-time interaction module, configured to superimpose the length of the distal bile duct, the angle degree of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size of filling defects in the bile duct and pancreatic duct on the original image in an augmented reality manner, and perform risk marking according to the early warning signal.

[0032] Among them, the privacy information desensitization processing refers to irreversibly deleting or replacing sensitive information such as the patient's name and ID number in the DICOM image, which can be specifically implemented by using a data masking algorithm or an encrypted hash algorithm. On the premise of ensuring the privacy and security of the patient, the pixel spacing and catheter size parameters are retained to provide basic data support for subsequent physical size calculation.

[0033] Noise reduction refers to using a non-local means filtering algorithm to eliminate image noise interference, specifically by calculating the similarity between pixel blocks for weighted average processing to improve the signal-to-noise ratio of the image.

[0034] Distortion correction refers to geometric correction based on the correspondence between the actual physical size of the catheter marker points and the image coordinates, specifically using an affine transformation algorithm to adjust the image spatial position to eliminate the projection distortion generated by X-ray fluoroscopy.

[0035] Dynamic range optimization refers to adjusting the image gray distribution through adaptive histogram equalization, specifically dynamically expanding the gray range according to the local area contrast to enhance the distinguishability between the bile and pancreatic ducts and the surrounding tissues.

[0036] The segmentation mask refers to a binary image area output by a deep learning model, specifically using a convolutional neural network with an encoder-decoder architecture to extract the contours of the bile duct and pancreatic duct to achieve accurate segmentation of anatomical structures.

[0037] The coordinates of sub - pixel - level key anatomical points refer to the localization of anatomical landmarks through a high - resolution neural network. Specifically, the HRNet network is used for interpolation calculation within pixels to achieve spatial localization with an accuracy better than that of a single pixel.

[0038] The augmented reality - based overlay means spatially registering and displaying the quantization parameters with the original image. Specifically, it realizes the real - time overlay mapping of measurement data and anatomical structures through DICOM coordinate system conversion and 3D projection algorithms.

[0039] The core innovation of this application lies in constructing a full - process intelligent analysis system from image pre - processing to real - time decision - making. By multi - modal data fusion and spatial geometric calculation, subjective image interpretation is transformed into objective quantitative indicators, and intraoperative real - time warning is achieved through augmented reality interaction. Through the collaborative design of de - identification processing and physical parameter retention, data security is guaranteed while ensuring measurement accuracy; through distortion correction based on catheter size and dynamic range optimization, geometric distortion and insufficient contrast of X - ray images are specifically addressed; through the joint output of segmentation masks and sub - pixel localization, multi - dimensional feature extraction of anatomical structure morphology and spatial position is realized; through the linkage mechanism of parameter measurement and warning rules, a standardized diagnostic evaluation system is established; through augmented reality overlay and risk marking, intraoperative operation time and radiation exposure risk are reduced.

[0040] The working process and principle of this application are as follows:

[0041] In the first step, the image acquisition module obtains the original DICOM - format ERCP angiography images, performs privacy information de - identification processing, and simultaneously retains pixel spacing and catheter size parameters, which provide basic data support for subsequent quantitative analysis.

[0042] In the second step, the image pre - processing module performs noise reduction on the de - identified original images, uses the retained catheter size parameters for distortion correction to solve the geometric distortion problem of X - ray images, and then performs dynamic range optimization to improve image quality and generate standardized image data.

[0043] In the third step, the multi - modal fusion model module extracts the segmentation masks of bile ducts and pancreatic ducts and the coordinates of sub - pixel - level key anatomical points from the standardized image data. This step integrates anatomical structure segmentation and key point localization to achieve accurate morphological feature extraction.

[0044] In the fourth step, the parameter measurement module calculates the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and number of filling defects in the bile duct and pancreatic duct based on the segmentation masks, sub - pixel key anatomical point coordinates, and pixel spacing parameters. This process transforms image features into quantifiable clinical diagnostic indicators.

[0045] In the fifth step, the early warning judgment module judges the calculated parameters according to the preset rules to generate an early warning signal. Thus, the system establishes an objective evaluation standard and reduces the subjective judgment deviation of traditional technologies.

[0046] In the sixth step, the real-time interaction module superimposes the calculated parameters on the original image in an augmented reality manner and performs risk marking according to the early warning signal, realizing the spatial registration of image data and anatomical structures and improving the intraoperative decision-making efficiency.

[0047] The above-mentioned modules work together to form a closed-loop system, effectively solving the technical defects of traditional ERCP, such as relying on subjective experience, insufficient quantification, and lack of real-time feedback.

[0048] As a preferred embodiment, the solution of this application is specifically implemented as follows:

[0049] The image acquisition module obtains the original ERCP angiography images in DICOM format from the PACS system, performs privacy information desensitization processing on the original images, removes sensitive information such as patient names and IDs, and at the same time retains the pixel spacing and catheter size parameters in the DICOM metadata. These parameters are stored in an independent data structure.

[0050] The image preprocessing module first uses the non-local mean filtering algorithm to perform noise reduction processing on the desensitized original images, then based on the retained catheter size parameters, locates the catheter marker points through the edge detection algorithm, performs affine transformation using these marker points for distortion correction, and finally applies the adaptive histogram equalization method to adjust the image contrast and verify whether the sharpness of the adjusted image meets the preset gradient threshold, and outputs the standardized image data.

[0051] The multi-modal fusion model module uses a pre-trained convolutional neural network to perform multi-scale feature extraction on the standardized image data, enhances the feature response of tiny anatomical structures through the attention mechanism, and uses the atrous spatial pyramid pooling operation to capture multi-scale anatomical context information. It fuses the enhanced tiny anatomical structure features and multi-scale anatomical context features to generate fusion features. Based on the fusion features, it generates the segmentation masks of the bile duct and pancreatic duct through decoding processing, and uses the HRNet network to extract high-resolution spatial information, and outputs the coordinates of sub-pixel-level key anatomical points.

[0052] The parameter measurement module generates the centerlines of the bile duct and pancreatic duct through a topological refinement algorithm based on the segmentation mask, determines the main direction vector of the bile duct centerline using the principal component analysis algorithm, calculates the angle between the main direction vectors to obtain the angulation degree of the distal bile duct, performs a curve integral along the bile duct centerline, calculates the length of the distal bile duct in combination with the pixel pitch parameter, generates a first dynamic cross-section along the normal direction of the bile duct centerline, generates a second dynamic cross-section along the normal direction of the pancreatic duct centerline, calculates the minimum diameter of the full circumference of the first dynamic cross-section as the bile duct diameter, calculates the minimum diameter of the full circumference of the second dynamic cross-section as the pancreatic duct diameter, and identifies and calculates the size and quantity of filling defects by comparing the contrast agent density distribution differences between the segmentation mask and the standardized image data.

[0053] The early warning judgment module judges the calculated length of the distal bile duct, the angulation degree of the distal bile duct, the bile duct diameter, the pancreatic duct diameter, and the size and quantity of filling defects in the bile duct and pancreatic duct according to preset rules, and generates corresponding early warning signals when these parameters reach or exceed the preset thresholds.

[0054] The real-time interaction module superimposes the calculated length of the distal bile duct, the angulation degree of the distal bile duct, the bile duct diameter, the pancreatic duct diameter, and the size of filling defects in the bile duct and pancreatic duct onto the original image in the form of dynamic floating labels, and when an early warning signal is detected, based on the segmentation mask and the coordinates of key anatomical points, marks the corresponding anatomical structure spatial position in the original image as a risk marker.

[0055] Through the above solutions, the present application realizes the full-process optimization of ERCP images from data security processing to intelligent diagnosis. The image acquisition module retains key physical parameters on the premise of ensuring patient privacy security, providing basic data support for subsequent quantitative analysis. The image preprocessing module specifically solves the problem of geometric distortion of X-ray images and improves the image quality. The multi-modal fusion model module realizes accurate morphological feature extraction. The parameter measurement module converts imaging features into clinically quantifiable diagnostic indicators. The early warning judgment module establishes an objective evaluation standard and reduces subjective judgment deviation. The real-time interaction module improves the intraoperative decision-making efficiency. These improvements effectively solve the technical defects of traditional ERCP relying on subjective experience, insufficient quantification, and lack of real-time feedback, improve the accuracy and efficiency of ERCP surgery, and reduce the risk of cumulative radiation.

[0056] In some of the above solutions of the present application, there are problems such as inappropriate selection of noise reduction algorithms resulting in loss of details, inaccurate positioning of catheter marker points causing deviation in geometric distortion correction, and insufficient image sharpness after dynamic range optimization affecting subsequent anatomical structure recognition.

[0057] In response to this, the present application further proposes that the image preprocessing module includes a noise reduction unit, a gradient calculation unit, a marker point extraction unit, a distortion correction unit, and a dynamic optimization unit.

[0058] Among them, the noise reduction unit processes the image using a non-local means filtering algorithm with a search window of 21×21 pixels, and the noise suppression coefficient can be set between 0.05 and 0.15. The gradient calculation unit calculates the horizontal and vertical gradients through the Sobel operator, and the convolution kernel size can be selected as 3×3 or 5×5 pixels. The marker point extraction unit sets the gradient threshold according to the catheter diameter parameter. When the catheter diameter is 3 mm, the gradient threshold can be set to 120-150 gray levels. The distortion correction unit calculates the affine transformation matrix using the least squares method, and the number of control points is not less than 4 groups. The dynamic optimization unit divides the image into sub-blocks of 32×32 pixels to perform histogram equalization, and calculates the sharpness index using the Laplace operator. The gradient threshold range is set to 0.15-0.25.

[0059] Specifically, the noise reduction unit eliminates noise by weighted averaging of non-locally similar regions, and preserves the continuity of the catheter edge structure. The horizontal and vertical gradient maps generated by the gradient calculation unit form a two-dimensional feature space for subsequent marker point detection. The marker point extraction unit establishes a spatial coordinate system in combination with the catheter physical size parameters, and determines the region where the gradient amplitude exceeds the set threshold as valid marker points. The distortion correction unit constructs an affine transformation model based on the marker point coordinates, and eliminates the projection deformation error through linear transformation. During the process of enhancing the contrast, the dynamic optimization unit avoids local overexposure through block equalization, and at the same time ensures that the output image meets the input requirements of the subsequent segmentation model based on edge intensity verification. This processing flow controls the processing intensity of each stage through parameterization, realizes image standardization while preserving anatomical details, and provides high-precision input data for the multi-modal fusion model.

[0060] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0061] The image preprocessing module includes a noise reduction unit, a gradient calculation unit, a marker point extraction unit, a distortion correction unit, and a dynamic optimization unit.

[0062] The noise reduction unit uses a non-local means filtering algorithm to perform noise reduction processing on the desensitized original image. Specifically, the search window size can be selected as 21×21 pixels, the similar window size as 7×7 pixels, and the filtering intensity parameter h is set to 10. In this way, the image noise can be effectively reduced while retaining the edge detail information.

[0063] The gradient calculation unit uses the Canny edge detection algorithm based on the Sobel operator to extract the gray gradient features of the noise-reduced image. Among them, the high threshold can be set to the 90% quantile of the image gray value, and the low threshold is set to 1 / 3 of the high threshold, so that the edge contour information in the image can be accurately captured.

[0064] The marker point extraction unit locates the catheter marker points based on the catheter size parameters and the extracted gray gradient features in the DICOM metadata. For example, the template matching method can be used, taking the theoretical shape of the catheter marker points as the template, sliding and searching on the gradient image to find the best matching position as the marker point coordinates.

[0065] The distortion correction unit performs an affine transformation on the denoised image using the located catheter marker points. At least 3 non - collinear marker points can be selected to construct an affine transformation matrix to achieve geometric correction of the image.

[0066] The dynamic optimization unit performs adaptive histogram equalization on the geometrically corrected image. The image can be divided into 8×8 grids, and histogram equalization is performed separately for each grid. Then, the results of adjacent grids are fused using the bilinear interpolation method. Subsequently, the Sobel operator is used to calculate the gradient magnitude of the image to verify whether the preset gradient threshold is reached. If the threshold is not reached, the parameters of histogram equalization can be appropriately adjusted, and the optimization process is repeated.

[0067] Through the above technical solutions, the present application can effectively solve the problems existing in the image pre - processing process. The non - local mean filtering algorithm can retain image details while reducing noise, avoiding the loss of key anatomical information that may be caused by traditional filtering methods. The marker point extraction method based on catheter size parameters and gray gradient features improves the accuracy of catheter marker point positioning, providing a reliable benchmark for subsequent geometric distortion correction. The dynamic optimization method combining adaptive histogram equalization and gradient verification not only enhances the image contrast but also ensures that the output image has sufficient sharpness, providing high - quality input data for subsequent anatomical structure recognition. This staged image - processing flow realizes the precise optimization of image quality, provides standardized high - quality image data for the multi - modal fusion model module, and thus improves the performance and reliability of the entire ERCP image intelligent recognition system.

[0068] In some of the above - mentioned solutions of the present application, due to the low contrast of micro - anatomical structures and the significant multi - scale morphological differences in ERCP images, traditional single - scale feature extraction methods are difficult to simultaneously capture micro - features and global spatial relationships, resulting in blurred boundaries of the segmentation mask and insufficient accuracy in key point positioning.

[0069] In response to this, the present application further proposes a multi - modal fusion model module, including a feature extraction unit, a local feature enhancement unit, a global relationship modeling unit, a feature fusion unit, and a segmentation and positioning unit.

[0070] Among them, the feature extraction unit performs multi-scale feature extraction on the normalized image data through a pre-trained convolutional neural network. ResNet or VGG can be used as the basic network to output high-resolution detail features and low-resolution semantic features. For example, three-scale feature maps of 64×64, 32×32, and 16×16 are extracted through the first three convolutional layers respectively. The local feature enhancement unit enhances the tiny structure features through the channel attention mechanism. Specifically, the SE module or the CBAM module is used to dynamically adjust the feature weights in the channel dimension. The spatial compression ratio can be set to 16 in the calculation of the channel attention weights to balance the calculation efficiency and the feature selection ability. The global relationship modeling unit captures multi-scale context information through the dilated spatial pyramid pooling operation. For example, parallel dilated convolutional layers with dilation rates of 6, 12, and 18 are used to expand the receptive field while maintaining the resolution of the feature map. The feature fusion unit performs feature fusion through channel concatenation and 1×1 convolution. After channel concatenation, the dimension can reach 2048, and it is reduced to 512 through 1×1 convolution to reduce the calculation amount. The segmentation and localization unit outputs the segmentation result and the coordinate prediction synchronously through a two-branch structure. The segmentation branch adopts the U-Net decoding structure, and the localization branch adopts HRNet to maintain high-resolution features. The number of stages of HRNet can be set to 4, and the output resolutions of each stage are 64×64, 32×32, 16×16, and 8×8 in sequence.

[0071] Specifically, after the normalized image data is input into the feature extraction unit, multi-scale features are extracted through different convolutional layers. Among them, the high-resolution feature map retains the details of the bile duct wall folds, and the low-resolution feature map contains the topological information of the pancreatic duct branches. The local feature enhancement unit performs channel attention weighting on the high-resolution features. For example, in the 512-channel features, a 512-dimensional channel weight vector is generated through global average pooling. After passing through the fully connected layer and the Sigmoid activation, the feature response at the bile duct branch is increased by 1.2-1.5 times, and at the same time, the noise response in the background area is suppressed by 30-40%. The global relationship modeling unit performs multi-scale context modeling on the low-resolution features and processes them in parallel through dilated convolutional layers with different dilation rates. The convolutional kernel with a dilation rate of 6 captures the association between adjacent branches, and the convolutional kernel with a dilation rate of 18 establishes the spatial connection between the main road and the distal branches. The feature fusion unit concatenates the weighted high-resolution features and the multi-scale context features and realizes channel dimension alignment through 1×1 convolution to form a fused feature map of 512×32×32. In the segmentation and localization unit, the segmentation branch restores the resolution through upsampling and outputs the segmentation mask, and the localization branch generates sub-pixel coordinates through the multi-resolution feature fusion of HRNet. For example, the coordinate prediction accuracy is improved to the 0.1-pixel level by using the bicubic interpolation method. Thus, through hierarchical feature processing, cross-scale spatial associations are established while maintaining the details of tiny structures, increasing the intersection over union index of bile duct segmentation by 8-12% and reducing the average error of key point localization to less than 1.2 pixels.

[0072] As a preferred embodiment, the solution of this application is specifically implemented as follows:

[0073] The multi-modal fusion model module includes a feature extraction unit, a local feature enhancement unit, a global relationship modeling unit, a feature fusion unit, and a segmentation and localization unit.

[0074] The feature extraction unit uses a pre-trained ResNet-101 convolutional neural network to perform multi-scale feature extraction on the normalized image data. The ResNet-101 network contains 5 convolutional stages, and the resolution of the output feature map of each stage decreases in turn, forming a feature pyramid from high resolution to low resolution. The feature extraction unit selects the outputs of the conv2, conv3, conv4, and conv5 layers of ResNet-101 as the multi-scale feature maps.

[0075] The local feature enhancement unit applies a channel attention mechanism to the feature map. First, it obtains the global descriptor of each channel through global average pooling, then uses a two-layer fully connected network to learn the correlation between channels, generates a channel weight vector, and finally multiplies the weight vector by the original feature map to achieve selective enhancement of the features of tiny anatomical structures.

[0076] The global relationship modeling unit uses atrous spatial pyramid pooling operation. It applies 3x3 convolutions with 4 different dilation rates in parallel on the feature map, and the dilation rates are 1, 6, 12, and 18 respectively. The convolutional kernels with different receptive fields capture context information at multiple scales at the same time, effectively establishing the association between far and near features.

[0077] The feature fusion unit concatenates the enhanced local features and the multi-scale context features along the channel dimension, and then reduces the number of channels through 1x1 convolution to generate fused features.

[0078] The segmentation and localization unit outputs a segmentation mask and key point coordinates based on the fused features at the same time. The segmentation branch uses 4 layers of transposed convolution to gradually upsample and restore the feature resolution, and finally generates a 2-channel segmentation mask through 1x1 convolution, corresponding to the bile duct and pancreatic duct respectively. The localization branch uses the HRNet network to maintain high-resolution features, extracts multi-scale features through 4 parallel branches and repeatedly performs feature fusion, and finally regresses and outputs the sub-pixel coordinates of key anatomical points.

[0079] Through the above technical solutions, the present application achieves accurate segmentation of the bile duct and pancreatic duct in ERCP images and accurate positioning of key anatomical points. Multi-scale feature extraction and fusion effectively capture image information at different scales, the attention mechanism enhances the feature response of tiny anatomical structures, and the dilated convolution expands the receptive field and establishes the association of near and far features. This multi-modal fusion method significantly improves the boundary clarity of the segmentation mask and the accuracy of key point positioning, providing a reliable data basis for subsequent parameter measurement and early warning judgment. At the same time, this method has strong generalization ability, can adapt to ERCP images of different patients and different imaging conditions, and improves the robustness and clinical practicability of the system.

[0080] Through the above technical solutions, the present application can effectively solve the problems existing in the image preprocessing process. The non-local means filtering algorithm can retain image details while reducing noise, avoiding the loss of key anatomical information that may be caused by traditional filtering methods. The marker point extraction method based on the catheter size parameter and gray gradient feature improves the accuracy of catheter marker point positioning, providing a reliable benchmark for subsequent geometric distortion correction. The dynamic optimization method combining adaptive histogram equalization and gradient verification not only enhances the image contrast but also ensures that the output image has sufficient sharpness, providing high-quality input data for subsequent anatomical structure recognition. This staged image processing flow realizes the precise optimization of image quality, provides standardized high-quality image data for the multi-modal fusion model module, and thus improves the performance and reliability of the entire ERCP image intelligent recognition system.

[0081] In some of the above solutions of the present application, conventional decoding processing may lose high-resolution spatial details, resulting in insufficient positioning accuracy of key anatomical points and unable to meet the sub-pixel level positioning requirements, thereby affecting the accuracy of subsequent parameter measurement.

[0082] In response to this, the present application further proposes that the segmentation and positioning unit includes a segmentation sub-unit and a high-resolution positioning sub-unit: the segmentation sub-unit performs decoding processing on the fusion features to generate segmentation masks of the bile duct and pancreatic duct; the high-resolution positioning sub-unit extracts the high-resolution spatial information of the fusion features through the HRNet network and outputs the sub-pixel level key anatomical point coordinates.

[0083] Among them, the decoding process of the segmentation sub-unit uses a skip connection structure to fuse shallow features and deep features, so that the morphological features of the end of the bile duct branches are retained during the generation of the segmentation mask. The high-resolution localization sub-unit includes four parallel feature extraction branches. Each branch maintains feature maps with downsampling rates of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively. Multi-scale feature interaction is achieved through a cross-resolution feature exchange module. The feature map output by the highest-resolution branch is directly used for coordinate regression. In the feature fusion stage, the multi-branch features are unified to the original resolution through bilinear interpolation and then perform pixel-wise weighted summation. The weighting coefficients are dynamically adjusted through learnable parameters. For example, the weight coefficient of the highest-resolution branch is set in the range of 0.6 - 0.8 to ensure the dominant role of the original spatial details.

[0084] Specifically, during the generation of the segmentation mask, the fused features are restored to the size of the input image through three deconvolution operations. After each deconvolution, channel concatenation is performed with the feature map of the corresponding scale in the encoding stage, and feature fusion is achieved through a 3×3 convolutional layer. During the localization of key anatomical points, the parallel branches of the HRNet network continuously retain the feature maps of the original resolution, avoiding the loss of spatial information caused by multiple downsamplings. When processing the bile duct bifurcation site, the fine structural features at the 0.5mm level extracted by the high-resolution branch are combined with the deep semantic features, and the coordinates of the key points are output through a coordinate regressor. Experiments show that this structure can control the localization error of the key points within 1 / 32 pixels, improving the localization accuracy by about 67% compared with the conventional decoding structure. In the scenario of calculating the angle of the bile duct, by simultaneously using the topological structure information provided by the segmentation mask and the spatial relationship of the sub-pixel coordinates, the angular measurement error is reduced from ±3.2° to ±1.5°, effectively supporting the accuracy of subsequent clinical decisions.

[0085] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0086] The segmentation and localization unit includes a segmentation sub-unit and a high-resolution localization sub-unit. The segmentation sub-unit decodes the fused features to generate segmentation masks for the bile duct and pancreatic duct. The high-resolution localization sub-unit extracts the high-resolution spatial information of the fused features through the HRNet network and outputs the sub-pixel level coordinates of the key anatomical points.

[0087] Specifically, the segmentation sub-unit adopts a multi-scale decoder structure to gradually restore the spatial resolution of the fused features. The decoder includes multiple upsampling and convolutional layers. Each upsampling layer doubles the size of the feature map, and the convolutional layer performs feature refinement. The last layer uses a 1x1 convolution to reduce the number of channels of the feature map to 2, corresponding to the two categories of the bile duct and the pancreatic duct, and then obtains the pixel-level segmentation mask through softmax activation.

[0088] The high-resolution positioning subunit is designed based on the HRNet network architecture. The network contains 4 parallel branches, which respectively maintain feature maps with input resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32. Information is exchanged between branches through a cross-resolution fusion module. The low-resolution branches capture semantic information, and the high-resolution branches retain spatial details. The output end of the network is connected to a regression head, which predicts the coordinates of sub-pixel-level key points. The regression head consists of two fully connected layers, and the output dimension is twice the number of key points, corresponding to the x and y coordinates respectively.

[0089] Through the above technical solutions, the present application realizes high-precision bile duct and pancreatic duct segmentation and key anatomical point positioning. The segmentation subunit uses multi-scale decoding to restore spatial details, ensuring the integrity and accuracy of the segmentation mask. The high-resolution positioning subunit is based on the parallel multi-branch architecture of HRNet, maintaining high-resolution feature expression during the decoding process and avoiding spatial information loss caused by downsampling. Through cross-resolution feature fusion, both deep semantic information and original image details are utilized to achieve sub-pixel-level precise positioning of key anatomical points. This design compensates for the positioning deviation of the conventional decoding process, provides a high-precision spatial information basis for subsequent parameter measurement, and effectively improves the accuracy and reliability of ERCP image analysis.

[0090] In some of the above solutions of the present application, due to the positioning error of catheter marker points or the correction deviation of image distortion, there is a proportional error between the pixel pitch parameter and the actual physical size, affecting the measurement accuracy of parameters such as bile duct length and diameter.

[0091] In response to this, the present application further proposes a dynamic calibration module, which performs the following steps: extracting the catheter size and pixel pitch parameter from the DICOM metadata; calculating the theoretical pixel length of the catheter marker point in the image according to the actual physical length of the catheter size and the pixel pitch parameter; measuring the actual pixel length of the catheter marker point based on the standardized image data generated by the image preprocessing module; using the ratio of the theoretical pixel length to the actual pixel length as the calibration factor; multiplying the calibration factor by the pixel pitch parameter to generate a corrected pixel pitch parameter, and inputting the corrected pixel pitch parameter into the parameter measurement module.

[0092] Among them, the calculation of the theoretical pixel length of the catheter marking points needs to combine the known physical length of the catheter and the original pixel pitch parameter, and the theoretical number of pixels is obtained by dividing the physical length by the pixel pitch. The actual pixel length measurement can be carried out on the standardized image data output by the image preprocessing module. After locating the boundaries of the marking points using edge detection or gray gradient features, the actual span is calculated. When generating the calibration factor, if there is a difference between the theoretical pixel length and the actual pixel length, the proportional deviation caused by image distortion or positioning error is quantified by the ratio. The corrected pixel pitch parameter is dynamically adjusted by multiplying the calibration factor by the original pixel pitch. For example, when the calibration factor is 0.98 and the original pixel pitch is 0.2 mm, it becomes 0.196 mm after correction.

[0093] Specifically, the dynamic calibration module realizes parameter self-correction through a closed-loop feedback mechanism. In the standardized image data generated by the image preprocessing module, the positions of the catheter marking points have been optimized through noise reduction and distortion correction, but residual errors will still cause the pixel pitch to not match the actual physical size. By comparing the theoretical pixel length corresponding to the known physical length of the catheter with the measured pixel length, the proportional error coefficient can be accurately calculated. For example, if the actual length of the catheter is 10 mm, the theoretical pixel length is 50 pixels, and the measured pixel length is 49 pixels, then the calibration factor is 50 / 49≈1.02, and the corrected pixel pitch is the original value multiplied by 1.02. This process ensures that the pixel pitch parameter used by the parameter measurement module always strictly corresponds to the true physical size, eliminating the cumulative error caused by image distortion or positioning error. Further, this dynamic calibration does not require manual intervention and is automatically executed in each image processing process, ensuring the consistency of the measurement benchmark between different images. Through the multiplication operation of the calibration factor and the pixel pitch parameter, proportional correction at the pixel level is achieved.

[0094] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0095] The dynamic calibration module extracts the catheter size and pixel pitch parameter from the DICOM metadata. For example, the catheter size can be 5 French, and the pixel pitch parameter can be 0.2 mm / pixel. According to the actual physical length of the catheter size and the pixel pitch parameter, the theoretical pixel length of the catheter marking points in the image is calculated. Specifically, if the distance between the catheter marking points is 10 mm, the theoretical pixel length is 50 pixels.

[0096] Furthermore, based on the standardized image data generated by the image preprocessing module, the actual pixel length of the catheter marking points is measured. For example, by using an edge detection algorithm to locate the catheter marking points, the measured actual pixel length is 48 pixels. Thus, the ratio of the theoretical pixel length to the actual pixel length is used as the calibration factor, which is 50 / 48 = 1.0417 in this example.

[0097] Finally, multiply the calibration factor by the pixel pitch parameter to generate a corrected pixel pitch parameter. Specifically, the corrected pixel pitch parameter is 0.2 mm / pixel * 1.0417 = 0.2083 mm / pixel. The corrected pixel pitch parameter is input into the parameter measurement module for subsequent calculation of parameters such as bile duct length and diameter.

[0098] Through the above technical solution, the present application realizes the dynamic calibration of the pixel pitch parameter. By comparing the theoretical pixel length and the actual pixel length of the catheter marking points, the system can automatically detect and compensate for the size deviation caused by distortion or positioning error in the image processing process. This dynamic calibration mechanism does not require manual intervention and can continuously ensure the consistency between the measured parameters and the actual physical size. Thereby, the system significantly improves the measurement accuracy of key parameters such as bile duct length and diameter, providing more reliable data support for doctors' diagnosis and treatment decisions. At the same time, this automated calibration process also improves the operation efficiency of the system and reduces the possibility of human error.

[0099] In some of the above solutions of the present application, the multi-modal fusion model module is used to extract the segmentation masks and key anatomical point coordinates of the bile duct and pancreatic duct. However, when calculating the morphological parameters of the bile duct and pancreatic duct based on these data, there are problems such as deviation in the extraction of the bile duct centerline, calculation error of the angulation degree, and risk of misjudgment of filling defects caused by insufficient segmentation accuracy and loss of spatial information, which cannot meet the parameter measurement accuracy required for clinical diagnosis.

[0100] In response to this, the present application further proposes a parameter measurement module, including a centerline extraction unit, a direction vector calculation unit, an angulation degree calculation unit, a length calculation unit, a diameter analysis unit, and a filling defect detection unit.

[0101] Among them, the centerline extraction unit processes the segmentation mask using a topological thinning algorithm. For example, the Zhang-Suen iterative thinning algorithm is applied to eliminate the skeleton burrs and generate a continuous centerline with a single-pixel width. The direction vector calculation unit performs principal component analysis based on the key anatomical point coordinates, calculates the eigenvectors through the covariance matrix, and defines the direction of the maximum variance as the main direction vector. The angulation degree calculation unit calculates the angle between the main direction vectors through the vector dot product formula to generate an angle value. The length calculation unit uses cubic spline interpolation to fit the centerline and then performs curve integration, and combines the corrected pixel pitch parameter for physical length conversion. The diameter analysis unit generates a dynamic cross-section along the normal direction of the centerline, detects the circular contour through the Hough transform, and calculates the minimum diameter. The filling defect detection unit applies the Otsu threshold segmentation algorithm to identify the low-density area, and then excludes the noise interference through morphological closing operation.

[0102] Specifically, the centerline extraction unit generates smooth and continuous bile duct centerlines and pancreatic duct centerlines through a topological thinning algorithm, providing a geometric reference for subsequent parameter calculations. The direction vector calculation unit extracts the principal direction vector based on principal component analysis, avoiding the angular error caused by the traditional endpoint connection method due to local curvature. The length calculation unit accumulates the arc lengths of each segment of the centerline through curve integration, and combines the corrected pixel pitch parameter output by the dynamic calibration module to convert the pixel distance into the actual physical length. The diameter analysis unit dynamically generates cross-sections along the normal direction of the centerline, and measures the minimum diameter of the full circumference to avoid the overestimation of the diameter caused by projection deformation. The filling defect detection unit excludes the interference of artifacts caused by uneven contrast agent distribution through density difference comparison and morphological rule verification. The corrected pixel pitch parameter provided by the dynamic calibration module is synchronously transmitted to the length calculation unit and the diameter analysis unit to ensure the parameter reference consistency between different measurement units. For example, when the theoretical pixel length of the catheter marker point is 150 pixels and the actual measured value is 148 pixels, the calibration factor 0.987 corrects the original pixel pitch of 0.2 mm to 0.197 mm, and this corrected value is used for all spatial dimension calculations, thereby reducing the parameter calculation error and meeting the parameter accuracy requirements for clinical diagnosis.

[0103] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0104] The parameter measurement module includes a centerline extraction unit, a direction vector calculation unit, an angle calculation unit, a length calculation unit, a diameter analysis unit, and a filling defect detection unit.

[0105] The centerline extraction unit processes the segmentation mask using a topological thinning algorithm. Specifically, first, the segmentation mask is binarized, with the bile duct or pancreatic duct region set as the foreground, and then the boundary pixels are iteratively deleted until no further deletion is possible. In each iteration, the 8-neighborhood of each foreground pixel is checked, and if deleting the pixel does not cause a change in connectivity, it is deleted. Finally, a skeleton with a width of 1 pixel is obtained as the bile duct centerline or pancreatic duct centerline.

[0106] The direction vector calculation unit determines the principal direction vector of the bile duct centerline based on the coordinates of key anatomical points through the principal component analysis algorithm. The specific steps are as follows: First, the coordinates of the key anatomical points are formed into a matrix X, the covariance matrix of X is calculated, and then the covariance matrix is eigen-decomposed to obtain the eigenvalues and eigenvectors. The eigenvector corresponding to the largest eigenvalue is the principal direction vector.

[0107] The angle calculation unit calculates the angle between the principal direction vectors to generate the angle of the distal bile duct. Specifically, the angle between the two principal direction vectors is calculated through the vector dot product formula, that is, cosθ = (a·b) / (|a||b|), where a and b are the two principal direction vectors.

[0108] The length calculation unit performs a curve integral along the center line of the bile duct and calculates the length of the distal bile duct in combination with the pixel pitch parameter. The specific steps are as follows: First, sample equally spaced points along the center line, then calculate the Euclidean distance between adjacent sampling points, and finally add up all the distances and multiply by the pixel pitch parameter to obtain the actual length of the bile duct.

[0109] The diameter analysis unit generates a first dynamic cross-section and a second dynamic cross-section along the normal directions of the center line of the bile duct and the center line of the pancreatic duct, and calculates the minimum diameters of the full circumferences of the first dynamic cross-section and the second dynamic cross-section in combination with the pixel pitch parameter as the diameter of the bile duct and the diameter of the pancreatic duct respectively. The specific steps are as follows: First, sample equally spaced points on the center line, calculate the tangent direction of the center line at each sampling point, then generate a plane perpendicular to the tangent as the dynamic cross-section. For the first dynamic cross-section and the second dynamic cross-section, calculate the intersection lengths of the cross-sections with the bile duct contour and the pancreatic duct contour respectively, and take the minimum value as the diameter of the cross-section. Finally, multiply the diameter in pixel units by the pixel pitch parameter to obtain the diameter in actual physical dimensions.

[0110] The filling defect detection unit detects and calculates the size and quantity of filling defects. The specific steps are as follows: First, compare the contrast agent density distribution of the segmentation mask and the standardized image data, and identify the regions with local density lower than the preset threshold as candidate filling defect regions. Then perform connected component analysis on the candidate regions to determine independent candidate filling defect instances. Further, combine morphological rules to exclude artifact interference and generate valid filling defect instances. Finally, label the maximum transverse diameter of each valid instance and count the number of instances as the number of filling defects.

[0111] Through the above technical solutions, the present application can improve the accuracy of anatomical parameter measurement. The center line extraction unit uses a topological thinning algorithm to eliminate the burr interference generated by conventional skeleton extraction algorithms and generate continuous and anatomically complete center lines of the bile duct and the pancreatic duct. The direction vector calculation unit combines the coordinates of key anatomical points and extracts the main direction vector from the perspective of spatial distribution through the principal component analysis algorithm, overcoming the angular error caused by the traditional method of calculating the direction based on the connection of endpoints. The length calculation unit uses the curve integral method to accumulate the pixel pitch along the center line, which can more accurately reflect the actual physiological curvature length of the bile duct. The diameter analysis unit dynamically generates cross-sections along the normal direction of the center line and finds the minimum diameter of the full circumference, which can avoid the problem of over-inflated diameter measurement caused by the image projection angle. The filling defect detection unit improves the specificity and sensitivity of filling defect recognition by integrating the segmentation mask and the original image data for density difference comparison and combining morphological rules to exclude artifact interference. Therefore, the present application can meet the parameter measurement accuracy required for clinical diagnosis and provide a reliable quantitative analysis basis for doctors.

[0112] In some of the above solutions of the present application, there may be problems such as inaccurate identification of candidate filling defect regions, misjudgment caused by artifact interference, inability to effectively determine independent instances, and inaccurate size conversion.

[0113] In response to this, the present application further proposes that the filling defect detection unit includes a density difference comparison subunit, an instance segmentation and verification subunit, a size annotation and statistics subunit, and a size conversion subunit.

[0114] Among them, the density difference comparison subunit compares the density distribution difference between the segmentation mask and the contrast agent density distribution of the standardized image data. For example, it uses a histogram comparison method or a region growing algorithm to identify regions where the local density is lower than a preset threshold, and the preset threshold can be set to 60%-80% of the average density of the contrast agent. The instance segmentation and verification subunit performs connected region analysis on the candidate regions, filters out small regions using an area threshold, for example, considers regions with an area less than 10 pixels as noise, and at the same time uses morphological closing operations to fill holes, and combines the aspect ratio rule to exclude linear artifacts. For example, candidate regions with an aspect ratio greater than 3:1 are determined as artifacts. The size annotation and statistics subunit extracts the main axis of each instance through skeletonization, and calculates the maximum transverse diameter using the rotational projection method. For example, it uses pixel-level edge detection combined with sub-pixel interpolation technology to improve the measurement accuracy. The size conversion subunit performs physical size conversion based on the pixel pitch parameter, and the pixel pitch parameter can be obtained from the physical size tag in the DICOM metadata, for example, in the range of 0.1 mm / pixel to 0.3 mm / pixel, and maintains the conversion accuracy through the bilinear interpolation algorithm.

[0115] Specifically, the density difference comparison subunit can effectively eliminate the interference caused by uneven contrast agent distribution by comparing the density distribution difference between the segmentation mask and the standardized image. For example, the low-density artifacts formed by contrast agent diffusion in the bile duct edge region. The instance segmentation and verification subunit combines connected region analysis and morphological rules. For example, it determines the boundary of independent instances through eight-neighbor connectivity detection, applies ellipse fitting to verify the rationality of the instance morphology, and excludes candidate regions with a circularity lower than 0.7. When measuring the maximum transverse diameter, the size annotation and statistics subunit uses a multi-directional projection method to obtain accurate dimensions. For example, it projects along four directions of 0°, 45°, 90°, and 135° to calculate the maximum value. During the physical size conversion process, the size conversion subunit combines the corrected pixel pitch parameter provided by the dynamic calibration module. For example, multiplying the calibration factor by the original pixel pitch can reduce the measurement error.

[0116] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0117] The filling defect detection unit includes a density difference comparison subunit, an instance segmentation and verification subunit, a size annotation and statistics subunit, and a size conversion subunit.

[0118] The density difference comparison subunit compares the density distribution differences of the bile duct and pancreatic duct regions in the segmentation mask with the contrast agent density of the standardized image data. Specifically, first, the average gray value of the internal region of the segmentation mask is calculated as the reference density value. Then, the standardized image data is scanned with a sliding window, and the average gray value within each window is calculated. Finally, the window average gray value is compared with the reference density value. If it is lower than 80% of the reference density value, the window region is marked as a candidate filling defect region.

[0119] The instance segmentation and verification subunit performs connected component analysis on the candidate filling defect regions. First, the candidate regions are labeled using the 8-connectivity criterion to obtain preliminary independent candidate filling defect instances. Further, morphological opening operations are applied to remove noise points smaller than 5 pixels. Subsequently, the area, perimeter, and circularity of each candidate instance are calculated. If the area is less than 10 pixels or the circularity is greater than 0.9, it is regarded as an artifact and excluded. Finally, the boundaries of the remaining candidate instances are smoothed to obtain valid filling defect instances.

[0120] The dimension annotation and statistics subunit annotates the maximum transverse diameter of each valid filling defect instance. Specifically, for each instance, the minimum bounding rectangle is fitted, and the long side of the rectangle is taken as the maximum transverse diameter. At the same time, the number of valid filling defect instances is counted as the number of filling defects.

[0121] The dimension conversion subunit converts the maximum transverse diameter of each valid filling defect instance into a filling defect size based on the pixel pitch parameter, where the filling defect size is the physical size of the maximum transverse diameter. For example, if the pixel pitch parameter is 0.2 mm / pixel and the maximum transverse diameter of an instance is 50 pixels, the corresponding filling defect size is 10 mm.

[0122] Through the above technical solutions, the present application achieves accurate detection and quantitative analysis of filling defects. The density difference comparison subunit can accurately identify local density abnormal regions by comparing the contrast agent density distribution differences between the segmentation mask and the standardized image data, avoiding the limitations of single-threshold segmentation. The instance segmentation and verification subunit effectively differentiates independent filling defect instances by combining connected component analysis with morphological rules and excludes the interference of image noise and artifacts. The size annotation and statistics subunit provides quantitative descriptions of the morphological features and distribution density of filling defects through the annotation of the maximum transverse diameter and the statistics of the instance quantity. The size conversion subunit performs physical size conversion based on the pixel spacing parameter to ensure the consistency of the measurement results with the actual anatomical structure. Thus, this solution overcomes problems such as inaccurate identification of candidate filling defect regions, misjudgment caused by artifact interference, inability to effectively determine independent instances, and inaccurate size conversion, providing reliable technical support for the automated analysis of filling defects in ERCP images.

[0123] In some of the above solutions of the present application, the lack of a quantitative judgment standard in the filling defect detection unit leads to late warning and subjective deviation, and the lack of a quantitative threshold determination mechanism for the angulation degree of the distal bile duct and the abnormal diameter of the bile and pancreatic ducts makes it difficult to achieve real-time intraoperative risk warning.

[0124] In response to this, the present application further proposes that the warning signal generation conditions include that the angulation degree of the distal bile duct is lower than the preset angulation degree threshold, the diameter of the bile duct or pancreatic duct exceeds the preset diameter threshold, the size of the filling defect reaches the preset size threshold, and the number of filling defects reaches the preset number threshold.

[0125] Among them, the angulation degree threshold of the distal bile duct is set as an adjustable parameter within the range of 30 degrees to 60 degrees, which is used to identify the intubation risk caused by excessive bending of the bile duct. The bile duct diameter threshold is set to 8 mm, and the pancreatic duct diameter threshold is set to 3 mm. This value is determined based on the pathological criteria of bile and pancreatic duct dilation in clinical guidelines. The filling defect size threshold is configured as an adjustable range of 3 mm to 8 mm, and the filling defect number threshold is set to 3. This parameter combination can effectively distinguish single stones from multiple impacted lesions. Each threshold parameter performs physical size conversion through the modified pixel spacing parameter of the dynamic calibration module to ensure the consistency of the measurement results with the anatomical reality. When any threshold condition is triggered, a warning signal is immediately generated, and the warning level is increased when multiple conditions are met simultaneously.

[0126] Specifically, during the intraoperative real-time analysis process, the angle value output by the angulation calculation unit is directly compared with a preset threshold. When the detected angle is lower than 50 degrees, it indicates that there is an abnormal anatomical curvature of the bile duct, triggering a first-level warning. The diameter analysis unit compares the minimum diameter of the dynamic cross-section with the preset threshold. When the diameter of the bile duct exceeds 8 mm or the diameter of the pancreatic duct exceeds 3 mm, a second-level warning is generated. The physical size data output by the filling defect detection unit is synchronously verified with the quantity statistical result. If it is detected that the size of a single filling defect exceeds 5 mm or there are 3 or more filling defect instances simultaneously, a third-level warning is triggered. The multi-level warning signals are projected onto the operation area in real time through the augmented reality interface. Among them, the first-level warning corresponds to a yellow flashing identifier, the second-level warning is superimposed with a red border, and the third-level warning activates a beeping prompt. All warning triggers are based on the accurate measurement data after dynamic calibration, avoiding misjudgment caused by image distortion or catheter marker errors in traditional methods. At the same time, through the multi-dimensional threshold combination determination mechanism, clinical experience is transformed into quantifiable trigger conditions, improving the accuracy of intraoperative risk identification and shortening the warning response time.

[0127] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The warning signal generation process of the warning judgment module constructs a multi-condition trigger mechanism through preset angulation thresholds, diameter thresholds, size thresholds, and quantity thresholds. During the ERCP operation, it receives in real time the distal bile duct angulation, bile and pancreatic duct diameters, filling defect sizes, and quantity data output by the parameter measurement module, and compares the angulation with the preset angulation threshold in real time. When it is detected that the angulation is lower than the threshold, a warning signal of abnormal bile duct curvature is generated; simultaneously, the diameter of the bile duct or pancreatic duct is compared with the preset diameter threshold. If it exceeds the threshold, a warning signal of duct dilation or stenosis is generated. After the filling defect detection data is analyzed by the connected region analysis, the maximum transverse diameter physical size and total quantity of each instance are extracted and compared with the preset size threshold and quantity threshold respectively. When the threshold conditions are met, a warning of stones or space-occupying lesions is triggered. After the warning signal is generated, a flashing identification frame is superimposed on the corresponding anatomical position of the original image through the real-time interaction module, and a beeping prompt sound is triggered.

[0128] Through the above technical solution, the present application realizes the real-time quantitative judgment of intraoperative anatomical structure abnormalities and filling defect characteristics, transforms the risk assessment driven by subjective experience into a multi-dimensional threshold trigger mechanism, effectively eliminates the influence of the lag of manual judgment and individual differences, and through the synergistic effect of the preset thresholds, it can simultaneously identify multiple risk factors such as abnormal bile duct morphology, duct dilation and contraction, and excessive lesion scale, ensuring that the generation of warning signals has clinical pathological relevance and avoiding the risk of misjudgment of a single parameter.

[0129] In some of the above solutions of this application, the traditional image overlay method has problems such as insufficient differentiation of anatomical structures, interference of parameter display on the surgical field, and inaccurate spatial positioning of risk areas, resulting in difficulties for doctors to quickly identify key anatomical structure differences, inability to intuitively judge the correspondence between parameters and anatomical positions, and inability to accurately mark the location of the lesion when an early warning is triggered.

[0130] In response to this, this application further proposes to enhance the display of the bile duct and pancreatic duct in different colors respectively, display the bile duct length, angulation, diameter, and filling defect size in the form of dynamic floating labels in real time. When an early warning signal is detected, based on the segmentation mask and the coordinates of key anatomical points, the corresponding anatomical structure spatial position is determined, and this spatial position is marked in the original image with a visual marker as a risk marker.

[0131] Among them, the color enhancement display of the bile duct and pancreatic duct can use the blue and green spectral ranges for differentiation, where blue corresponds to the bile duct structure and green corresponds to the pancreatic duct structure, and the hue difference between the two colors needs to be greater than 120 degrees to form a significant contrast; the display position of the dynamic floating label can be automatically adjusted based on the edge area of the image. For example, when there is a bifurcation of the bile duct in the central area of the surgical field, the label can float to the lower right corner area of the image, and the label transparency can be set to 50%-70% to avoid completely blocking the anatomical structure; the spatial positioning of the risk marker is achieved by establishing a mapping relationship between the pixel coordinates of the segmentation mask and the coordinates of key anatomical points. For example, the affine transformation matrix is used to map the mask coordinate system to the original image coordinate system, so as to convert the abstract early warning signal into a specific coordinate position, and the positioning accuracy can reach the sub-pixel level.

[0132] Specifically, during the image enhancement display process, the bile duct and pancreatic duct are given different colors with high contrast, and the anatomical boundary identification is strengthened through the hue difference to avoid misjudgment during the operation; the dynamic floating label adjusts the display position in real time as the surgical field moves, which not only maintains the spatial correspondence between the parameter and the anatomical structure, but also avoids the label continuously blocking the key area; when the early warning signal is triggered, based on the segmentation mask, the contour coordinates of the lesion area are obtained, and combined with the spatial mapping relationship of the coordinates of key anatomical points, a flashing marker box with a position anchor point is generated in the original image. For example, a narrow bile duct segment is accurately circled with a red circular marker box, and the diameter of the marker box can be dynamically adjusted according to the filling defect size. Thus, doctors can intuitively identify the anatomical position corresponding to the abnormal parameter and quickly lock the lesion area through the accurately positioned risk marker, effectively improving the intraoperative decision-making efficiency and operation safety.

[0133] When implementing augmented reality overlay and risk marking, the bile duct and pancreatic duct are rendered in blue and green respectively through the AR module of the OpenCV library, and the color coding follows the DICOM PS3.16 standard. The dynamic floating label is implemented using the floating window control of the Qt framework, and the label position is dynamically adjusted based on the curvature change of the bile duct centerline to ensure that the label maintains a preset pixel spacing from the anatomical structure. When the warning signal is triggered, the contour coordinate data of the segmentation mask is called, and the coordinates of the key anatomical points are mapped to the image coordinate system through affine transformation, and a red semi-transparent circular mark with a diameter of 5 millimeters is generated at the target position, and the error between the center point of the mark and the centroid coordinates of the lesion is controlled within 0.3 millimeters.

[0134] Through the above technical solutions, the visual recognition of the anatomical structure boundary is effectively improved, the clarity of the bile and pancreatic duct demarcation is guaranteed, and at the same time, the automatic avoidance mechanism of the dynamic label reduces the occlusion of the surgical field. The risk marking positioning error based on sub-pixel coordinate mapping is lower than that of traditional methods, realizing the millimeter-level spatial matching of the lesion position and warning information, and significantly shortening the intraoperative decision-making response time.

[0135] In some of the above solutions of this application, a multi-modal fusion model module is proposed to extract the segmentation masks and sub-pixel level key anatomical point coordinates of the bile duct and pancreatic duct from the standardized image data. However, due to the complex structures such as the pre-trained convolutional neural network, attention mechanism, atrous spatial pyramid pooling, and HRNet network included in this model module, the number of model parameters is huge, and there are problems of high computational resource occupancy and limited inference speed in real-time interaction scenarios, making it difficult to meet the requirements of low-latency processing and edge device deployment during clinical interventional treatment.

[0136] In response to this, this application further proposes that an intelligent recognition system for ERCP images based on multi-modal fusion further includes a model optimization unit for compressing the number of parameters of the multi-modal fusion model module through the knowledge distillation algorithm.

[0137] Among them, the core of the knowledge distillation algorithm is to transfer the knowledge possessed by the complex network structures such as the pre-trained convolutional neural network and HRNet network included in the original multi-modal fusion model to the lightweight model. By establishing a teacher-student model framework, the feature map distribution, segmentation mask probability, and key point coordinate prediction results output by the teacher model are used as supervision signals to guide the training process of the student model. This technical means enables the student model to still inherit the teacher model's ability to extract multi-scale anatomical structure features, enhance the recognition of tiny lesions, and achieve high-precision positioning while reducing the number of convolutional layers, channels, or network width, thereby maintaining the segmentation and positioning accuracy on the premise of reducing the model computational complexity.

[0138] By compressing the number of parameters through the above technical solutions, the burden of the model on memory occupation and inference time consumption can be effectively reduced, enabling the system to adapt to mobile terminals or embedded devices, meet the requirements of real-time image processing during ERCP surgery, and at the same time avoid the problem of increased hardware costs caused by an overly large model.

[0139] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An intelligent recognition system for ERCP images based on multimodal fusion, characterized in that, Including: An image acquisition module, which is used to acquire the original ERCP contrast image in DICOM format, perform privacy information desensitization processing on the original image, and retain the pixel spacing and catheter size parameters in the DICOM metadata; An image preprocessing module, which is used to perform noise reduction, distortion correction based on the catheter size parameters, and dynamic range optimization on the desensitized original image to generate standardized image data; A multi-modal fusion model module, which is used to extract the segmentation masks and sub-pixel level key anatomical point coordinates of the bile duct and pancreatic duct from the standardized image data; A parameter measurement module, which is used to calculate the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct based on the segmentation mask, the sub-pixel key anatomical point coordinates, and the pixel spacing parameters; An early warning judgment module, which is used to judge the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size and quantity of filling defects in the bile duct and pancreatic duct according to preset rules, and generate an early warning signal; A real-time interaction module, which is used to superimpose the length of the distal bile duct, the angle of the distal bile duct, the diameter of the bile duct, the diameter of the pancreatic duct, and the size of the filling defects in the bile duct and pancreatic duct on the original image in an augmented reality manner, and perform risk marking according to the early warning signal.

2. The intelligent recognition system for ERCP images based on multimodal fusion according to claim 1, wherein The image preprocessing module includes: A noise reduction unit, which is used to reduce the image noise of the desensitized original image by using a non-local means filtering algorithm to generate a noise-reduced image; A gradient calculation unit, which is used to extract the gray gradient features of the noise-reduced image through an edge detection algorithm; A marker point extraction unit, which is used to locate the catheter marker points based on the catheter size parameters and the gray gradient features; A distortion correction unit, which is used to perform an affine transformation on the noise-reduced image based on the catheter marker points to generate a geometrically corrected image; A dynamic optimization unit, which is used to perform adaptive histogram equalization on the geometrically corrected image to adjust the contrast, and verify whether the sharpness of the adjusted image meets the preset gradient threshold, and output standardized image data.

3. The intelligent recognition system for ERCP images based on multi-modal fusion according to claim 2, characterized in that, The multi-modal fusion model module includes: A feature extraction unit, which is used to perform multi-scale feature extraction on the standardized image data through a pre-trained convolutional neural network to generate feature maps containing different resolution levels; A local feature enhancement unit, which is used to enhance the response of the tiny anatomical structure features in the feature map through an attention mechanism; A global relationship modeling unit, which is used to capture the multi-scale anatomical context information of the feature map through a dilated spatial pyramid pooling operation; A feature fusion unit, which is used to fuse the enhanced tiny anatomical structure features and the multi-scale anatomical context features to generate fused features; A segmentation and localization unit, which is used to output the segmentation masks and sub-pixel level key anatomical point coordinates of the bile duct and pancreatic duct based on the fused features.

4. The intelligent recognition system for ERCP images based on multi-modal fusion according to claim 3, characterized in that, The segmentation and localization unit includes: A segmentation sub-unit, which is used to perform decoding processing on the fused features to generate the segmentation masks of the bile duct and pancreatic duct; A high-resolution positioning subunit is used to extract the high-resolution spatial information of the fused features through an HRNet network and output the coordinates of the sub-pixel-level key anatomical points.

5. An intelligent recognition system for ERCP images based on multi-modal fusion according to claim 4, characterized in that, It further includes a dynamic calibration module, and the dynamic calibration module performs: Extract the catheter size and pixel pitch parameters from the DICOM metadata; Calculate the theoretical pixel length of the catheter marking point in the image according to the actual physical length of the catheter size and the pixel pitch parameters; Measure the actual pixel length of the catheter marking point based on the standardized image data generated by the image preprocessing module; Take the ratio of the theoretical pixel length to the actual pixel length as the calibration factor; Multiply the calibration factor by the pixel pitch parameter to generate a corrected pixel pitch parameter, and input the corrected pixel pitch parameter into the parameter measurement module.

6. An intelligent recognition system for ERCP images based on multimodal fusion according to any one of claims 3 to 5, characterized in that, The parameter measurement module includes: A centerline extraction unit for generating the bile duct centerline and the pancreatic duct centerline through a topological refinement algorithm based on the segmentation mask; A direction vector calculation unit for determining the main direction vector of the bile duct centerline through a principal component analysis algorithm based on the key anatomical point coordinates; An angle calculation unit for calculating the included angle between the main direction vectors to generate the angulation degree of the distal bile duct; A length calculation unit for performing curve integration along the bile duct centerline and calculating the length of the distal bile duct in combination with the pixel pitch parameter; A diameter analysis unit for generating a first dynamic section along the normal direction of the bile duct centerline and a second dynamic section along the normal direction of the pancreatic duct centerline, and calculating the minimum diameter of the full circumference of the first dynamic section as the bile duct diameter according to the pixel pitch parameter, and calculating the minimum diameter of the full circumference of the second dynamic section as the pancreatic duct diameter; A filling defect detection unit for detecting and calculating the size and quantity of filling defects.

7. An intelligent recognition system for ERCP images based on multi-modal fusion according to claim 6, characterized in that, The filling defect detection unit includes: A density difference comparison subunit for comparing the density distribution differences of the contrast agent between the bile duct and pancreatic duct regions in the segmentation mask and the standardized image data, and identifying the regions with local density lower than the preset threshold as candidate filling defect regions; An instance segmentation and verification subunit for performing connected region analysis on the candidate filling defect regions, determining independent candidate filling defect instances, and excluding artifact interference in combination with morphological rules to generate effective filling defect instances; A size annotation and statistics subunit for annotating the maximum transverse diameter of each effective filling defect instance and counting the number of instances as the filling defect quantity; A size conversion subunit for converting the maximum transverse diameter of each effective filling defect instance into a filling defect size based on the pixel pitch parameter, and the filling defect size is the physical size of the maximum transverse diameter.

8. An intelligent recognition system for ERCP images based on multimodal fusion according to claim 7, characterized in that, The warning signal generation conditions of the warning judgment module include: The angulation degree of the distal bile duct is lower than the preset angulation degree threshold; The diameter of the bile duct or pancreatic duct exceeds the preset diameter threshold; The filling defect size reaches the preset size threshold; The number of filling defects reaches the preset quantity threshold.

9. The intelligent recognition system for ERCP images based on multi-modal fusion according to claim 1, wherein The specific operation of superimposing the bile duct length, angulation degree, diameter, and filling defect size on the original image in an augmented reality manner and performing risk marking according to the warning signal is as follows: The bile duct and pancreatic duct are respectively highlighted in different colors; The bile duct length, angulation degree, diameter, and filling defect size are displayed in real time in the form of dynamic floating labels; When a warning signal is detected, based on the segmentation mask and the coordinates of the key anatomical points, the spatial position of the anatomical structure corresponding to the warning signal is determined, and a visual marker is used to mark this spatial position on the original image as a risk mark.

10. An intelligent recognition system for ERCP images based on multimodal fusion according to claim 3, characterized in that, It further includes: A model optimization module for compressing the number of parameters of the multimodal fusion model module through a knowledge distillation algorithm.

Citation Information

Patent Citations

  • Artificial intelligence auxiliary diagnosis model construction system for medical images

    CN113239972A

  • Urban streetscape advertisement image segmentation method

    CN116189180A

  • Deep learning and image processing combined pig carcass backfat thickness measurement position automatic positioning method

    CN116596999A

  • Medical image segmentation method based on multi-scale feature fusion

    US20250095828A1

  • Small sample steel defect detection method based on attention feature pyramid mechanism

    WO2025010883A1

Cited By

  • Photodynamic therapy imaging positioning method and system

    CN120525758A

  • Multi-modal brain image intelligent feature extraction method and system based on deep learning

    CN120747705A

  • Deep learning-based multi-modal brain image intelligent feature extraction method and system

    CN120747705B

  • AKI diagnosis method fusing multi-modal image features

    CN120894311A

  • Oil and gas pipeline magnetic flux leakage image defect identification method based on deep attention mechanism

    CN121259407A