Inverse depth map conversion method and model based on complex clinical scene
Through the extended cycle consistency loss and direction discriminator of the XDCycleGAN model, the problems of realism and depth consistency during image conversion in complex clinical scenarios are solved, high-quality inverse depth maps are generated, the effects of 3D reconstruction and automated diagnosis are improved, and the data acquisition cost and privacy risks are reduced.
Patent Information
- Application Number
- CN202510973392.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-21
Smart Images

Figure CN120823973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an inverse depth map conversion method and model based on complex clinical scenarios. Background Art
[0002] Complex clinical scenarios, such as internal examinations of the human body, including the colon, esophagus, and stomach, require the use of colonoscopy and gastroscopy. For example, colon cancer, one of the most lethal malignancies worldwide, relies heavily on colonoscopy for early screening and accurate diagnosis. Clinical statistics show that approximately 20% of potential precancerous polyps are missed due to the complex intestinal environment (such as dynamic lighting, mucosal reflections, and instrument obstruction) and differences in operator experience, severely limiting screening efficiency.
[0003] Related technologies use 3D reconstruction techniques based on monocular depth estimation to improve lesion localization and navigation accuracy. Model training samples use cross-domain image conversion methods to narrow the gap between synthetic virtual data and real data, but two common problems exist. On the one hand, deep structural information is lost during image conversion. On the other hand, these methods often rely on models pre-trained on clinical data, such as using a pre-trained depth estimator to extract structural features. This makes such methods difficult to implement in scenarios where real clinical annotations are scarce.
[0004] These defects collectively make it difficult for related cross-domain image conversion methods to balance realism and depth consistency in complex clinical scenarios, hindering their actual implementation in 3D reconstruction and automated diagnosis. Summary of the Invention
[0005] This paper proposes an inverse depth map conversion method based on complex clinical scenarios, which can generate highly realistic images while preserving the original depth and geometric structure of the synthesized data.
[0006] In order to achieve the above effects, the present invention provides the following technical solutions: a method and model for inverse depth map conversion based on complex clinical scenarios, comprising the following steps: On the one hand, an inverse depth map conversion method based on complex clinical scenarios is applied to the generator, including: converting first optical image domain data of the complex clinical scene into first inverse depth domain data of the complex clinical scene; Converting the first inverse depth domain data into second optical image domain data, and converting the second optical image domain data into second inverse depth domain data; Minimize the difference between the first inverse depth domain data and the second inverse depth domain data to obtain the third inverse depth domain data
[0007] In a related embodiment, a method for discriminating the image conversion direction in the process from the first optical image domain data to the third inverse depth data is also included, which is applied to the direction discriminator; The method for discriminating the image conversion direction includes: discriminating the input and generated images, distinguishing the conversion of optical image domain data to inverse depth domain data and the conversion of inverse depth domain data to optical image domain data, outputting a pairing authenticity probability, and guiding the generator to accurately align the target domain data distribution.
[0008] In a related embodiment, minimizing the difference between the first inverse depth domain data and the second inverse depth domain data to obtain the third inverse depth domain data is achieved by extending the cycle consistency loss function, which is: ; in is a generator that converts inverse depth domain data b into optical image domain data a, A generator that converts optical image domain data a into inverse depth domain data b. is the data distribution of image domain data, represents the L1 norm; Indicates that data y is input into the generator Among them It is used to map the optical image domain data of complex clinical scenes to the inverse depth domain data of complex clinical scenes. The data distribution of data y sampled from the optical image domain data is , where A is the optical image domain; Where E represents the expected value, which is the probability-weighted average of all possible values of the random variable y; This loss function preserves the deep structural information by constraining the geometric consistency between the initial generation result and the secondary loop reconstruction result in the target domain.
[0009] In a related embodiment, the direction discriminator: , its loss function is: ; Among them, D is defined as the discriminator, is the data distribution of b image domain data; express Indicates that data x is input into the generator Among them It is used to map the inverse depth domain data of complex clinical scenes to the optical image domain data of complex clinical scenes. The data distribution of data x sampled from the inverse depth domain data is , where B is the inverse depth domain; Where E represents the expected value, which is the probability-weighted average of all possible values of the random variable x; By distinguishing the input-generated pairs , optical image domain number a→inverse depth domain data b direction, and , inverse the depth domain data b→optical image domain data a direction, explicitly discriminate the domain conversion direction, and guide the generator to align with the target domain distribution.
[0010] In a related embodiment, the total loss function is: ; in, are the weight coefficients of extended cycle loss, adversarial loss and b identity loss, is the cycle consistency loss, is the extended consistency loss, is the direction discriminator loss, is the generative adversarial loss, It’s a loss of identity.
[0011] Secondly, an inverse depth map conversion model based on complex clinical scenarios includes: Generator : used for mapping optical image domain data of a complex clinical scene to inverse depth domain data of the complex clinical scene; used for executing the step of converting first optical image domain data of the complex clinical scene into first inverse depth domain data of the complex clinical scene; Generator : used for mapping the inverse depth domain data of the complex clinical scene to the optical image domain data of the complex clinical scene; used for executing the step of converting the first inverse depth domain data into the second optical image domain data; Direction discriminator: used to perform the steps of claim 2.
[0012] In a related embodiment, the pre-training method of the model includes: The generator is trained using optical images of complex clinical scenes and inverse depth maps that do not correspond to the optical images. ,The optical images of complex clinical scenes contain the texture features of the ,corresponding scenes; The generator is trained using inverse depth maps of complex clinical scenes and optical images that do not correspond to the inverse depth maps. .
[0013] The present invention provides a method and model for inverse depth map conversion based on complex clinical scenarios, which has the following beneficial effects: (1) The present invention generates first inverse depth domain data and second inverse depth domain data in sequence through an extended cycle from the real optical image domain data of a complex clinical scene, and finally minimizes the difference between the first inverse depth domain data and the second inverse depth domain data to retain their inherent geometric consistency; by directly constraining the geometric consistency between the first inverse depth domain data generated initially in the image depth domain and the second inverse depth domain data reconstructed as a result of the secondary cycle, that is, forcing the two to be consistent in structural information, rather than forcing the reconstructed second optical image domain data to be completely consistent with the first optical image domain data in appearance, thereby avoiding the appearance features of the optical image domain data input from remaining in the image depth domain data; ensuring the realism and depth consistency of the final generated image; retaining key structural information and improving the generalization ability of real scenes: the invention can effectively retain the original depth and geometric information in the synthesized data when performing image conversion, which solves the problem that traditional methods lose fine geometric structures such as the curvature of the colon wall folds and the depth mutation of the lesion boundary during the conversion process, so that the generated image maintains a high degree of realism while not losing structural consistency, thereby improving the generalization ability of the model in real clinical scenes.
[0014] (2) The realistic optical image domain data and the corresponding real synthetic inverse depth domain data generated by the present invention can be used together as new training samples for the monocular depth estimation model, eliminating the dependence on large-scale real optical image domain depth annotation data: the framework proposed by the present invention does not rely on models pre-trained on clinical data. For example, there is no need to use a pre-trained depth estimator to extract structural features. This solves the pain point of the existing technology that is difficult to implement when real clinical annotation data is scarce, and significantly reduces the high cost of data acquisition and potential privacy risks.
[0015] (3) The present invention enhances the accuracy of inter-domain conversion: by introducing a direction discriminator, the model can explicitly discriminate the direction of image conversion (for example, from the optical colon domain image to the depth domain, or vice versa), which guides the generator to more accurately learn and align the data distribution of the target domain, further improving the ability to remove the unique texture and lighting of optical colonoscopy and strengthening the correlation between the two domains. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of traditional cycle consistency loss; Figure 2 This is a schematic diagram of the extended consistency loss of the present invention; Figure 3 Schematic diagram of traditional discriminator; Figure 4 Schematic diagram of the direction discriminator of the present invention; Figure 5 Flowchart for improving XDCycleGAN to generate realistic colon images; Figure 6 A comparison chart of the improved realistic colon image generation effect based on XDCycleGAN. DETAILED DESCRIPTION
[0017] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0018] Complex clinical scenarios, such as internal examinations of the human body, such as the colon, esophagus, and stomach, require the use of colonoscopes, gastroscopes, etc.; related technologies are often based on 3D reconstruction technology based on monocular depth estimation, which is considered a key means to improve lesion localization and navigation accuracy, but its practical application is limited by two major problems: first, real clinical data has privacy issues and lacks dense depth annotations, making it difficult to use directly for supervised learning; second, although existing synthetic data and bionic models can provide depth labels, they cannot truly simulate the biological reflectance characteristics and complex geometric structures of living tissues, resulting in insufficient generalization capabilities of the models in real scenarios.
[0019] Existing studies have attempted to narrow the gap between synthetic and real data through cross-domain image conversion methods, but there are two common problems. On the one hand, deep structural information is lost during image conversion. For example, the cycle consistency constraint in the generative adversarial network method only focuses on global content matching, resulting in the weakening of fine geometric structures in the synthetic data (such as the curvature of the colon wall folds and the depth mutation of the lesion boundary) during the conversion process. On the other hand, existing technologies usually rely on models pre-trained on clinical data. For example, pre-trained depth estimators are required to extract structural features. Such methods are difficult to implement in scenarios where real clinical annotations are scarce. These defects together make it difficult for existing cross-domain image conversion methods to balance realism and depth consistency in complex clinical scenarios, hindering their actual implementation in colon 3D reconstruction and automated diagnosis.
[0020] This embodiment proposes an inverse depth map conversion method based on complex clinical scenarios. The core of the method is to generate highly realistic images while preserving the original depth and geometric structure of the synthesized data. The colon image is used as an example.
[0021] 1. XDCycleGAN model First, let's introduce the architectural principles of the XDCycleGAN model. The XDCycleGAN model aims to solve the problem of lossy image cross-domain conversion between optical colonoscopy (OC) and virtual colonoscopy (VC). In the following description, image a represents the OC image in XDCycleGAN, and image b represents the VC image in XDCycleGAN. This model implements bidirectional domain conversion based on the Cycle-Consistent Adversarial Networks (CycleGAN): G is defined as the generator, Convert a into b image with preserved geometry, while Convert the b-image into a realistic a-image containing patient-specific texture and specular highlights.
[0022] like Figure 1 As shown in Figure 2, when traditional CycleGAN processes the conversion from image a to image b, its standard cycle consistency loss mechanism often forces the network to embed and retain redundant information of domain a in the generated image b, such as texture and lighting details: (1.1) in A generator that converts image b to image a, A generator that converts image a to image b, is the data distribution of image domain data, The goal of traditional CycleGAN is to ensure that the synthesized b-image can be reconstructed back to an image that is as consistent as possible with the original a-image. However, this is contrary to the goal of the a-to-b-image conversion, which is to remove such visual information and extract pure structure.
[0023] like Figure 2 As shown in Figure 2, to overcome the above limitations, XDCycleGAN proposes an extended cycle consistency loss to replace the standard cycle consistency loss: (1.2) The meaning of the letters in this formula is consistent with that in formula (1.1). For a given domain data y, first pass the generator Convert it to b-domain data , then through the generator b domain data Convert back to domain a data , and then use the generator again Convert the converted a domain data to get Finally, the goal is to minimize the difference between the b-image data obtained by the two transformations in the b-domain. This loss is achieved by directly constraining the initial generation result in the b-domain and the secondary cycle reconstruction results The geometric consistency between them is to force the two to be consistent in structural information, rather than forcing the reconstructed image a to be completely consistent with the original image a in appearance, thereby avoiding the residual appearance features of the input a in the domain b.
[0024] In order to further enhance the model's ability to remove the unique texture and lighting of domain a and establish a stronger correlation between domain a and domain b, XDCycleGAN introduces a direction discriminator , the loss function of the discriminator is: (1.3) Among them, D is defined as the discriminator, is the data distribution of the b image domain data, and the meanings of the remaining letters are consistent with those in formula (1.1). Figure 3 As shown in , the traditional discriminator in CycleGAN is direction agnostic, such as Figure 4 As shown, the direction discriminator of the present invention By distinguishing the input-generated pairs (a→b direction) and (b→a direction), explicit discriminant domain conversion direction ( The output is the pairing authenticity probability), thereby guiding the generator to accurately align with the target domain distribution.
[0025] Finally, the total loss function of XDCycleGAN training is: (1.4) in, are the weight coefficients of extended cycle loss, adversarial loss and b identity loss, is the cycle consistency loss (Equation (1.1)), is the extended consistency loss (Eq. (1.2)), is the direction discriminator loss (Formula (1.3)), is the generative adversarial loss (Equation (1.5)), is the identity loss (Equation 12.6)), which is used to preserve the color consistency of b.
[0026] It should be noted that and The same direction is calculated. Experiments show that the model achieves scale-consistent depth inference in the a→b transformation and can synthesize polyp images with complex textures through b-geometry control.
[0027] (1.5) (1.6) 2 Realistic Colon Image Generation Model Based on XDCycleGAN The core technical contribution of XDCycleGAN lies in overcoming the lossy problem of image translation from domain A to domain B. This model not only significantly improves the fidelity of cross-domain image translation but also generates high-quality depth maps.
[0028] The present invention demonstrates that this method can effectively transfer and convert simulated colon images into realistic colon images, and the generated images perfectly preserve the original structural information. This property makes them an ideal data source for training high-precision depth estimation models.
[0029] Specifically, based on the advantages of XDCycleGAN, this paper improves the generation method of realistic colon images. The specific process is as follows: Figure 5 shown.
[0030] First, we train XDCycleGAN using real colonoscopy images containing specific textures and (non-corresponding) inverse depth maps. The goal of this stage is to learn two generators that transform each other: Generator Can map the real colonoscopy image (a domain) to the inverse depth domain, the generator The inverse depth map can be mapped back to the a domain; the specific texture is the texture feature of the corresponding scene on the real optical image, such as the mucosal texture on the colonoscopic image (the surface texture of the colon mucosa, including its folds, grooves and microvilli and other details), vascular texture (the distribution and morphology of blood vessels under the colon mucosa, these blood vessels may appear as linear or reticular structures in the image), crypt openings (openings of crypts on the surface of the colon mucosa, these openings may appear as small depressions or holes in the endoscopic image) and reflections and shadows (due to the lighting conditions in the colon cavity and the smoothness of the mucosal surface, reflections and shadows may appear in the image), etc.
[0031] Next, the colon images in the simulated dataset are fed into the previously trained a-domain to inverse deep domain generator This step aims to strip away the texture information of the image while preserving its core structural information, thereby generating an inverse depth representation of the simulated colon RGB image.
[0032] Finally, the extracted inverse depth map is input into the generator from the inverse depth domain to the a domain , and then synthesize a high-fidelity colon image that retains the original simulation data structure and has real texture features.
[0033] In the experiment, the image dataset used for training includes 1,800 unpaired real colonoscopy images and thousands of high-definition real colonoscopy images open-sourced by the XDCycleGAN model team; the depth map dataset used for training includes the inverse depth map open-sourced by the XDCycleGAN model team.
[0034] Figure 6 The first row shows the original images input to the generator during the inference phase. These images are derived from the SimCol3D simulated colon dataset and, as indicated by the annotations above the images, cover different observation angles of the colon. The second row shows the realistic colon images generated by processing the simulated images obtained by the model trained on the XDCycleGAN dataset. The third row shows the realistic colon images generated by reasoning on the same batch of simulated images obtained by the model trained on high-definition real colonoscopy images.
[0035] The technology of this embodiment can be used not only for colonoscopic images, but also for other complex clinical scenarios, such as optical images of the stomach, esophagus, etc.
[0036] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for inverse depth map conversion based on complex clinical scenarios, characterized in that: Applies to generators, including: converting first optical image domain data of the complex clinical scene into first inverse depth domain data of the complex clinical scene; Converting the first inverse depth domain data into second optical image domain data, and converting the second optical image domain data into second inverse depth domain data; The difference between the first inverse depth domain data and the second inverse depth domain data is minimized to obtain third inverse depth domain data.
2. The inverse depth map conversion method based on complex clinical scenarios according to claim 1 is characterized in that: Also included is a method for discriminating the image conversion direction in the process from the first optical image domain data to the third inverse depth data, which is applied to the direction discriminator; The method for discriminating the image conversion direction includes: discriminating the input and generated images, distinguishing the conversion of optical image domain data to inverse depth domain data and the conversion of inverse depth domain data to optical image domain data, outputting a pairing authenticity probability, and guiding the generator to accurately align the target domain data distribution.
3. The inverse depth map conversion method based on complex clinical scenarios according to claim 1 is characterized in that: Minimizing the difference between the first inverse depth domain data and the second inverse depth domain data to obtain the third inverse depth domain data is achieved by extending the cycle consistency loss function, which is: ; in is a generator that converts inverse depth domain data b into optical image domain data a, A generator that converts optical image domain data a into inverse depth domain data b. is the data distribution of image domain data, represents the L1 norm; Indicates that data y is input into the generator Among them It is used to map the optical image domain data of complex clinical scenes to the inverse depth domain data of complex clinical scenes. The data distribution of data y sampled from the optical image domain data is , where A is the optical image domain; Where E represents the expected value, which is the probability-weighted average of all possible values of the random variable y; This loss function preserves deep structural information by constraining the geometric consistency between the initial generation result and the secondary loop reconstruction result in the target domain.
4. The inverse depth map conversion method based on complex clinical scenarios according to claim 2 is characterized in that: Direction Discriminator: , its loss function is: ; Among them, D is defined as the discriminator, is the data distribution of b image domain data; Indicates that data x is input into the generator Among them It is used to map the inverse depth domain data of complex clinical scenes to the optical image domain data of complex clinical scenes. The data distribution of data x sampled from the inverse depth domain data is , where B is the inverse depth domain; Where E represents the expected value, which is the probability-weighted average of all possible values of the random variable x; By distinguishing the input-generated pairs , optical image domain number a→inverse depth domain data b direction, and , inverse the depth domain data b→optical image domain data a direction, explicitly discriminate the domain conversion direction, and guide the generator to align with the target domain distribution.
5. The inverse depth map conversion method based on complex clinical scenarios according to any one of claims 1 to 4, characterized in that: The total loss function is: ; in, are the weight coefficients of extended cycle loss, adversarial loss and b identity loss, is the cycle consistency loss, is the extended consistency loss, is the direction discriminator loss, is the generative adversarial loss, It’s a loss of identity.
6. A model for inverse depth map conversion based on complex clinical scenarios, characterized by: include: Generator : Used to map optical image domain data of complex clinical scenes to inverse depth domain data of complex clinical scenes; for performing the step of converting first optical image domain data of the complex clinical scene into first inverse depth domain data of the complex clinical scene; Generator : used for mapping the inverse depth domain data of the complex clinical scene to the optical image domain data of the complex clinical scene; used for executing the step of converting the first inverse depth domain data into the second optical image domain data; Direction discriminator: used to perform the steps of claim 2.
7. The inverse depth map conversion model based on complex clinical scenarios according to claim 6, characterized in that: Model pre-training methods include: The generator is trained using optical images of complex clinical scenes and inverse depth maps that do not correspond to the optical images. ,The optical images of complex clinical scenes contain the texture features of the ,corresponding scenes; The generator is trained using inverse depth maps of complex clinical scenes and optical images that do not correspond to the inverse depth maps. .