Three-dimensional image generation method and device, medium and equipment
By capturing two-dimensional images at different focal planes using an endoscope and generating three-dimensional images using a depth model, the problem of providing accurate three-dimensional information in existing technologies is solved, enabling the generation of three-dimensional images from two-dimensional endoscopes and improving the operability of surgery.
Patent Information
- Application Number
- CN202410582667.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-11
- Publication Date
- 2025-11-11
AI Technical Summary
Existing 3D spinal endoscopes cannot provide accurate 3D image information, leading to difficulties in surgical planning and operation. 2D spinal endoscopes cannot provide depth information, and CT images have low resolution and cannot be updated in real time.
Two-dimensional images of the lesion area are captured by endoscopy at different focal planes. The image features and distances of pixels are determined. A three-dimensional image of the lesion area is generated using a depth model. The two-dimensional images and the depth map are then fused.
It enables the acquisition of three-dimensional information of the lesion area through two-dimensional endoscopy, helping doctors to better plan and operate the surgery, and providing accurate depth information.
Smart Images

Figure CN120931476A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, medium and device for generating three-dimensional images. Background Technology
[0002] Currently, with endoscopic minimally invasive surgery overcoming the drawbacks of traditional surgery, such as large trauma and long recovery periods, endoscopy is widely favored in the clinical field. However, due to the large size of 3D endoscopes, they are prone to getting stuck when inserted into the incision, such as 3D spinal endoscopes, which can easily become stuck in the spine. While 2D spinal endoscopes are smaller, they cannot provide 3D images of the inside of the spine. Therefore, neither 3D nor 2D spinal endoscopes can effectively assist doctors in surgical planning and operation.
[0003] In existing technologies, a combination of two-dimensional endoscopy and three-dimensional CT images is typically used to guide surgeons during operations and expand their field of vision. However, because CT images are taken before surgery and have low resolution for soft tissues, the two-dimensional endoscopic view cannot adequately provide surgeons with three-dimensional information about the patient's spine, thus hindering their surgical planning and execution.
[0004] Therefore, this specification provides a method, apparatus, medium, and device for generating three-dimensional images. Summary of the Invention
[0005] This specification provides a method, apparatus, medium, and device for generating three-dimensional images, in order to partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This specification provides a method for generating three-dimensional images, including:
[0008] Acquire two-dimensional images of the lesion area taken by endoscopy at different focal planes;
[0009] Determine the image features corresponding to each pixel in each two-dimensional image;
[0010] Determine the target image from each of the two-dimensional images, and determine at least some pixels in the target image that match pixels in other two-dimensional images;
[0011] Based on the image features of the partial pixels in the target image and the image features of the matching pixels, the distance from the partial pixels in the target image to the endoscope is determined;
[0012] The depth map of the target image is determined based on the distance from the partial pixels to the endoscope;
[0013] The target image and the depth map are fused to generate a three-dimensional image of the lesion area.
[0014] Optionally, acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes, specifically including:
[0015] Based on the preset focal length adjustment range, determine each focal plane for acquiring the two-dimensional image;
[0016] Adjust the camera parameters of the endoscope to acquire two-dimensional images of the lesion area taken at each focal plane.
[0017] Optionally, acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes, specifically including:
[0018] The position of the endoscope for acquiring two-dimensional images is determined based on the preset endoscope movement range.
[0019] The endoscope is moved, and when the endoscope is moved to the endoscope position, a two-dimensional image of the lesion area is acquired at a fixed focusing distance.
[0020] Optionally, the image features corresponding to each pixel in each two-dimensional image are determined separately, specifically including:
[0021] For each two-dimensional image, the degree of defocus of each pixel in the two-dimensional image is determined by the point spread function;
[0022] Based on the image features of the selected pixels in the target image and the image features of the matched pixels, the distance from the selected pixels in the target image to the endoscope is determined, specifically including:
[0023] Based on the degree of defocus of the partial pixels in the target image and the degree of defocus of the matched pixels, the distance from the partial pixels in the target image to the endoscope is determined by a pre-trained defocus ranging model.
[0024] Optionally, the image features corresponding to each pixel in each two-dimensional image are determined separately, specifically including:
[0025] For each two-dimensional image, determine the sharpness of each pixel in that two-dimensional image;
[0026] Based on the image features of the selected pixels in the target image and the image features of the matched pixels, the distance from the selected pixels in the target image to the endoscope is determined, specifically including:
[0027] For the target image, based on the determined sharpness of the partial pixels and the sharpness of the pixels that match the partial pixels in the other two-dimensional images, the focal length at which the partial pixels have the highest sharpness in each two-dimensional image is determined.
[0028] Based on the determined focal length of the selected pixels, the distance from the selected pixels to the endoscope is determined.
[0029] Optionally, a target image is determined from each of the two-dimensional images, and at least some pixels in the target image are determined to match pixels in other two-dimensional images, specifically including:
[0030] Determine the target image from each two-dimensional image, and determine at least a portion of the pixels in the target image;
[0031] For each pixel identified in the target image, determine the coordinates of that pixel;
[0032] In other two-dimensional images, identify the pixel whose coordinates match that of the given pixel, and use that pixel as the matching pixel.
[0033] Optionally, a target image is determined from each of the two-dimensional images, and at least some pixels in the target image are determined to match pixels in other two-dimensional images, specifically including:
[0034] Determine the target image from each two-dimensional image, and determine at least a portion of the pixels in the target image;
[0035] At least some pixels in the target image are matched with feature points in other two-dimensional images to determine the matching pixels in the other two-dimensional images.
[0036] Optionally, the trained deep model includes at least an encoding layer, a fusion layer, and a decoding layer;
[0037] The depth map of the target image is determined based on the distance from the selected pixels to the endoscope, specifically including:
[0038] Each of the two-dimensional images is input into the coding layer of the depth model to obtain the image features of each of the two-dimensional images output by the coding layer;
[0039] The image features of each two-dimensional image are input into the fusion layer to obtain the fused image features of each two-dimensional image;
[0040] The fused image features are input into the decoding layer to obtain the depth map output by the decoding layer;
[0041] The depth map of the target image is determined based on the distance from the partial pixels to the endoscope and the depth map output by the decoding layer.
[0042] This specification provides a three-dimensional image generation apparatus, comprising:
[0043] The acquisition module is used to acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes;
[0044] The first determining module is used to determine the image features corresponding to each pixel in each two-dimensional image;
[0045] The second determining module is used to determine a target image from each of the two-dimensional images, and to determine at least some pixels in the target image that match pixels in other two-dimensional images;
[0046] The third determining module is used to determine the distance from the partial pixels in the target image to the endoscope based on the image features of the partial pixels in the target image and the image features of the matching pixels;
[0047] The fourth determining module is used to determine the depth map of the target image based on the distance from the partial pixels to the endoscope;
[0048] The fusion module is used to fuse the target image and the depth map to generate a three-dimensional image of the lesion area.
[0049] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described three-dimensional image generation method.
[0050] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a three-dimensional image generation method.
[0051] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0052] The three-dimensional image generation method provided in this specification first acquires two-dimensional images of the lesion region taken by an endoscope at different focal planes. For each two-dimensional image, the image features corresponding to each pixel in the two-dimensional image are determined. The image features of at least a portion of the pixels in the two-dimensional image that match pixels in other two-dimensional images are determined. Based on the image features of the portion of pixels in the two-dimensional image and the image features of the matching pixels, the distance from the portion of pixels in the two-dimensional image to the endoscope is determined. Based on the distance from the portion of pixels to the endoscope, a depth map of the two-dimensional image is determined. The two-dimensional image and the depth map are fused to generate a three-dimensional image of the lesion region.
[0053] By using two-dimensional images with different focal planes, depth information of the lesion area is generated, enabling the acquisition of this depth information even through a two-dimensional endoscope. Fusing this depth information with the two-dimensional image yields a three-dimensional image of the lesion area, providing doctors with 3D information about the lesion and better assisting them in surgical planning and execution. Attached Figure Description
[0054] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0055] Figure 1 A flowchart illustrating a three-dimensional image generation method provided in an embodiment of this specification;
[0056] Figure 2 This is a schematic diagram illustrating how a depth model is used to output a depth map, as provided in this specification.
[0057] Figure 3 This specification provides a flowchart for generating a three-dimensional image;
[0058] Figure 4 This is a schematic diagram of a three-dimensional image generation device provided in this specification;
[0059] Figure 5 This specification provides a corresponding Figure 1 A schematic diagram of the structure of an electronic device. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.
[0061] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0062] Figure 1 This is a flowchart illustrating a three-dimensional image generation method provided in an embodiment of this specification, including the following steps:
[0063] S100: Acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes.
[0064] The process of generating a 3D image in this specification typically involves processing image data. In the embodiments described herein, the 3D image generation process can be performed by a server. Of course, this specification does not limit the type of device or platform used to perform the 3D image generation process; for example, personal computers, intelligent surgical robots, and mobile terminals can also be used. For ease of description, the following description uses a server as the executing entity.
[0065] In one or more embodiments of this specification, the server acquires two-dimensional images of the lesion area taken by the endoscope at different focal planes.
[0066] The two-dimensional images acquired by the server are captured by a two-dimensional endoscope, such as a two-dimensional spinal endoscope. For ease of explanation, the following description will use a two-dimensional spinal endoscope as an example; that is, the endoscope referred to in the following instructions is a two-dimensional spinal endoscope. The endoscope can be inserted deep into the wound, placed in the lesion area, and fixed. Then, the focal length is adjusted according to a preset range to determine the focal planes for acquiring two-dimensional images. The camera parameters of the endoscope are adjusted to focus the camera on the determined focal planes, and two-dimensional images of the lesion area are acquired at each focal plane.
[0067] The endoscope can also be inserted deep into the wound. During the process of placing it into the lesion area, when a certain lesion area appears in the field of view within the range that the endoscope can capture, it can be photographed. This is equivalent to adjusting the depth of the endoscope in the lesion area while moving the endoscope and capturing two-dimensional images of different focal planes.
[0068] Specifically, the server determines the endoscope position for acquiring two-dimensional images based on a preset endoscope movement range. Then, the endoscope is moved, and when it reaches the determined position, a two-dimensional image of the lesion area is acquired at a fixed focusing distance.
[0069] Of course, in this specification, regardless of the method used to capture two-dimensional images, the number of two-dimensional images captured is at least two.
[0070] S102: Determine the image features corresponding to each pixel in each two-dimensional image.
[0071] In one or more embodiments of this specification, for each two-dimensional image, the server determines the image features corresponding to each pixel in the two-dimensional image.
[0072] Specifically, the server can evaluate the sharpness of each 2D image, determine the sharpness of each pixel in each 2D image, and use the sharpness as the image feature corresponding to the pixel. Alternatively, it can calculate the degree of defocus in each 2D image using a point spread function, and then determine the degree of defocus of each pixel, using the degree of defocus as the image feature corresponding to the pixel. Of course, the degree of defocus can be represented by the contrast of the 2D image, and this specification does not limit this.
[0073] It is worth noting that the image features corresponding to each pixel in the two-dimensional image determined by the server in this manual can be either sharpness or defocus. This manual does not impose any restrictions on this, and it can be set according to the actual situation.
[0074] S104: Determine the target image from each of the two-dimensional images, and determine at least some pixels in the target image that match pixels in other two-dimensional images.
[0075] In one or more embodiments of this specification, because the two-dimensional images of the lesion area taken by the endoscope with different focal planes may show a situation where a certain object present in one two-dimensional image may not be captured in another. In this case, only a portion of the image in the first two-dimensional image can correspond to a portion of the image in the second two-dimensional image.
[0076] Therefore, the server can determine the target image from the acquired two-dimensional images and identify at least some pixels in the target image that match pixels in other two-dimensional images. Of course, the method for determining the target image can be to randomly select a two-dimensional image, or to display all the two-dimensional images to the endoscope user for selection. This manual does not limit the method for determining the two-dimensional image; it can be set according to the actual situation.
[0077] Specifically, the server determines the target image from each two-dimensional image and determines at least some of the pixels in the target image. Specifically, it may determine all the pixels in the target image or the pixels in the middle area of the target image. This specification does not limit how the pixels are determined.
[0078] Next, the pixels in the target image that match at least some of the pixels are identified in other two-dimensional images. If the two-dimensional images acquired by the server were obtained by moving an endoscope, then feature point matching is performed on at least some of the pixels in the target image with those in other two-dimensional images to determine the matching pixels. That is, when the pixels surrounding a pixel match the pixels surrounding a pixel in another two-dimensional image, that pixel can be considered to match a pixel in the other two-dimensional image.
[0079] Of course, if the 2D images acquired by the server are captured using an endoscope, the target image and at least some pixels can be determined from each 2D image. Then, the coordinates of these pixels in the target image can be determined. These coordinates are then matched with the coordinates of pixels in other 2D images to determine the matching pixels. It's worth noting that determining the matching pixels in other 2D images can also be done by performing feature point matching between at least some pixels in the target image and other 2D images.
[0080] S106: Determine the distance from the partial pixels in the target image to the endoscope based on the image features of the partial pixels in the target image and the image features of the matched pixels.
[0081] In one or more embodiments of this specification, the server determines the distance from some pixels in the target image to the endoscope based on the image features of some pixels in the target image and the image features of matching pixels.
[0082] Specifically, the server can determine the distance from some pixels in the target image to the endoscope based on the degree of defocus of some pixels in the target image and the degree of defocus of some pixels matching the target image, using a pre-trained defocus ranging model.
[0083] Of course, the server can also, for the target image, determine the focal length of the pixels with the highest sharpness across all two-dimensional images based on the sharpness of the identified pixels and the sharpness of the matching pixels in other two-dimensional images. Then, based on the determined focal length of the identified pixels, the distance from these pixels to the endoscope can be determined.
[0084] S108: Determine the depth map of the target image based on the distance from the partial pixels to the endoscope.
[0085] In one or more embodiments of this specification, the server determines the depth map of the target image based on the distance of some pixels in the target image to the endoscope.
[0086] Specifically, the server constructs a two-dimensional matrix based on the resolution of the target image. Then, the distances from some pixels in the target image to the endoscope are added to the two-dimensional matrix to obtain a depth matrix. Finally, the depth map of the target image is determined based on the depth matrix.
[0087] S110: The target image and the depth map are fused to generate a three-dimensional image of the lesion area.
[0088] In one or more embodiments of this specification, the server fuses the target image and the depth map of the target image to generate a three-dimensional image of the lesion area.
[0089] based on Figure 1 The method for generating a three-dimensional image, as shown, first acquires two-dimensional images of the lesion region taken with an endoscope at different focal planes. For each two-dimensional image, the image features corresponding to each pixel in that two-dimensional image are determined. The image features of at least a portion of the pixels in that two-dimensional image that match pixels in other two-dimensional images are determined. Based on the image features of the portion of pixels in the two-dimensional image and the image features of the matching pixels, the distance from the portion of pixels in the two-dimensional image to the endoscope is determined. Based on the distance from the portion of pixels to the endoscope, a depth map of the two-dimensional image is determined. The two-dimensional image and the depth map are fused to generate a three-dimensional image of the lesion region.
[0090] By using two-dimensional images with different focal planes, depth information of the lesion area is generated, enabling the acquisition of this depth information even through a two-dimensional endoscope. Fusing this depth information with the two-dimensional image yields a three-dimensional image of the lesion area, providing doctors with 3D information about the lesion and better assisting them in surgical planning and execution.
[0091] Furthermore, in one or more embodiments of this specification, the endoscope captures each two-dimensional image, wherein the field of view within each two-dimensional image is a gradual process from blurry to clear and then back to blurry. Of course, it can also be an image sequence consisting of multiple blurry two-dimensional images and multiple clear two-dimensional images, or it can be just two blurry two-dimensional images.
[0092] In one or more embodiments of this specification, since the two-dimensional images captured by the endoscope may contain noise and other issues, in order to maximize the usability of the final generated three-dimensional image of the lesion region, image enhancement operations, such as noise reduction and image smoothing, can be performed on the two-dimensional images acquired by the server. Alternatively, in this specification, the server can also use a pre-trained depth completion network to perform depth image enhancement on the depth map, and then fuse the enhanced depth map with the target image to generate a three-dimensional image of the lesion region, thus completing any discontinuous areas that may exist in the depth map, making the generated three-dimensional image more usable.
[0093] In one or more embodiments of this specification, the two-dimensional images captured by the endoscope and obtained by the server can be captured at different focal planes by manually adjusting the focus knob of the endoscope. Alternatively, the server can automatically adjust the endoscope to capture the focus area of the lesion region, thus capturing two-dimensional images at different focal planes. Of course, both manual and automatic adjustments can be made by adjusting only the focus of the endoscope while keeping it fixed. Two-dimensional images of the lesion region can also be captured intermittently while the endoscope is moved.
[0094] In one or more embodiments of this specification, a trained deep model can also be used to determine the depth map of a two-dimensional image. The server takes the acquired two-dimensional images as an image sequence and inputs it as a whole into the trained deep model to obtain the depth map output by the deep model. The trained deep model includes at least an encoding layer, a fusion layer, and a decoding layer.
[0095] Specifically, each 2D image in the image sequence is input into the encoding layer of the trained deep model to obtain the image features of each 2D image output by the encoding layer. These image features are then input into a fusion layer to obtain the fused features of each 2D image. Finally, the fused features are input into a decoding layer to obtain the depth map output by the decoding layer.
[0096] Finally, the server can determine a preliminary depth map of the target image based on the distance of some pixels in the target image to the endoscope. Then, based on the depth map output by the decoding layer, it makes fine adjustments to the preliminary depth map of the target image to determine the final depth map.
[0097] It is worth noting that the depth map obtained from the depth model trained in this manual can also be directly fused with a two-dimensional image to obtain a three-dimensional image for use.
[0098] Figure 2 This diagram illustrates a method for outputting a depth map using a depth model, as provided in this specification. The depth model includes an encoding layer, pooling layer, fusion layer, bottleneck layer, and decoding layer. Each two-dimensional image in the two-dimensional image sequence is input into a sub-encoding layer of the first encoding layer of the depth model. Figure 2 From left to right, using pooling layers as boundaries, the layers are arranged as follows: Layer 1, Layer 2, Layer 3, and Layer 4. In the first encoding layer, each encoding sub-layer extracts image features and inputs them into the first fusion layer. The image features output from each encoding sub-layer in the first layer are then processed through global pooling in the first layer and input into the next encoding sub-layer along with the image features output from that sub-layer. This continues until the final result is input into the bottleneck layer. The output of the bottleneck layer is then input into the fourth fusion layer, which reduces subsequent model computation and saves computational resources. The output of the fourth fusion layer and the result of the first fusion layer are then input into the first decoding layer. Figure 2 The decoding layers from left to right are numbered first, second, and third, until the third decoding layer outputs the depth map.
[0099] In one or more embodiments of this specification, the method for training a deep model may be as follows: the server inputs the acquired two-dimensional images into the encoding layer of the deep model, and the encoding layer outputs the image features of each two-dimensional image. The image features of each two-dimensional image are then input into a fusion layer to obtain the fused image features of each two-dimensional image. Finally, the fused image features of each two-dimensional image are input into a decoding layer to obtain the overall depth map and the overall full-focus map of each two-dimensional image.
[0100] The overall depth map is input into the depth prediction layer to obtain the predicted depth map for each 2D image. The predicted depth maps for each 2D image are then fused together to obtain the predicted full-focus map for each 2D image.
[0101] Based on the overall full-focus image of each 2D image and the predicted depth map of each 2D image, the point spread function convolutional layer of the depth model is used to obtain each predicted 2D image.
[0102] Calculate the fuzzy estimation loss of the predicted full-focus image, as well as the reconstruction loss of each predicted 2D image and each 2D image. Adjust the model parameters of the depth model with the goal of minimizing the sum of the fuzzy estimation loss and the reconstruction loss.
[0103] The trained deep model can directly obtain a depth map from the acquired two-dimensional image. Moreover, the depth map obtained by the neural network-based deep model has more comprehensive depth information and can express more information. It does not have the problem of missing or discontinuous parts of the depth map.
[0104] Figure 3 This is a flowchart illustrating a three-dimensional image generation method provided in this specification. Figure 3 As shown:
[0105] S300: Acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes.
[0106] In one or more embodiments of this specification, the server acquires two-dimensional images of the lesion area taken by the endoscope at different focal planes.
[0107] S302: Input each two-dimensional image into the encoding layer of the trained deep model to obtain the image features of each two-dimensional image output by the encoding layer.
[0108] In one or more embodiments of this specification, for ease of description, an example is given where the deep model has only one encoding layer, a fusion layer, and a decoding layer. The server inputs each acquired two-dimensional image into the encoding layer of the trained deep model to obtain the image features of each two-dimensional image output by the encoding layer.
[0109] S304: Input the image features of each two-dimensional image into the fusion layer in the depth model to obtain the fused image features of each two-dimensional image.
[0110] In one or more embodiments of this specification, the server inputs the image features of each two-dimensional image into the fusion layer in the depth model to obtain the fused image features of each two-dimensional image.
[0111] S306: Input the fused image features into the decoding layer of the depth model to obtain the depth map output by the decoding layer.
[0112] In one or more embodiments of this specification, the server fuses image features into the decoding layer of the depth model to obtain a depth map output by the decoding layer.
[0113] S308: The two-dimensional images are fused with the depth map output by the decoding layer to generate a three-dimensional image of the lesion area.
[0114] In one or more embodiments of this specification, the server fuses each two-dimensional image with the depth map output by the decoding layer to generate a three-dimensional image of the lesion area.
[0115] For details regarding steps S300 to S308 above, please refer to the description of the previous method; they will not be described in detail here.
[0116] In one or more embodiments of this specification, the server can determine each focal plane for acquiring two-dimensional images based on a preset focal length adjustment range. The camera parameters of the endoscope are adjusted to focus the endoscope camera on each determined focal plane, acquiring two-dimensional images of the lesion area captured at each focal plane. Then, the sharpness of each two-dimensional image can be evaluated to determine the sharpness corresponding to each pixel in each two-dimensional image, using the sharpness as the image feature corresponding to the pixel. After determining the target image and at least some pixels from each two-dimensional image, the coordinates of these pixels are determined for at least some pixels in the target image. The coordinates of these pixels are then matched with the coordinates of pixels in other two-dimensional images to determine the matching pixels. Finally, for the target image, based on the sharpness of the determined pixels and the sharpness of the matching pixels in other two-dimensional images, the focal length at which the sharpness of the pixels is highest in each two-dimensional image is determined. The distance from the determined pixels to the endoscope is then determined based on the focal length of the determined pixels. Finally, the depth map of the target image is determined based on the distance from the determined pixels to the endoscope. The target image and its depth map are fused to generate a three-dimensional image of the lesion area.
[0117] In one or more embodiments of this specification, the server can determine the focal planes for acquiring two-dimensional images based on a preset focal length adjustment range. The camera parameters of the endoscope are adjusted to focus the endoscope camera on the determined focal planes, acquiring two-dimensional images of the lesion area at each focal plane. Then, the sharpness of each two-dimensional image can be evaluated to determine the sharpness corresponding to each pixel in each two-dimensional image, using the sharpness as the image feature corresponding to the pixel. Next, at least some pixels in the target image are matched with feature points in other two-dimensional images to determine the matching pixels in the other two-dimensional images. Finally, for the target image, based on the sharpness of the determined partial pixels and the sharpness of the matching pixels in other two-dimensional images, the focal length at which the sharpness of the partial pixels is highest in all two-dimensional images is determined. Then, based on the determined focal length of the partial pixels, the distance from the partial pixels to the endoscope is determined. Based on the distance from the partial pixels in the target image to the endoscope, the depth map of the target image is determined. The target image and the depth map of the target image are fused to generate a three-dimensional image of the lesion area.
[0118] In one or more embodiments of this specification, the server can determine the focal planes for acquiring two-dimensional images based on a preset focal length adjustment range. The camera parameters of the endoscope are adjusted to focus the endoscope camera on the determined focal planes, acquiring two-dimensional images of the lesion area under each focal plane. Then, the degree of defocus in each two-dimensional image is calculated using a point spread function, thereby determining the degree of defocus for each pixel, and using the degree of defocus as the image feature corresponding to the pixel. Next, at least some pixels in the target image are matched with feature points in other two-dimensional images to determine the matching pixels in the other two-dimensional images. Finally, based on the degree of defocus of some pixels in the target image and the degree of defocus of the matching pixels, a pre-trained defocus ranging model is used to determine the distance from some pixels in the target image to the endoscope. Based on the distance from some pixels in the target image to the endoscope, a depth map of the target image is determined. The target image and its depth map are fused to generate a three-dimensional image of the lesion area.
[0119] In one or more embodiments of this specification, the server can determine the focal planes for acquiring two-dimensional images based on a preset focal length adjustment range. The camera parameters of the endoscope are adjusted to focus the endoscope camera on the determined focal planes, acquiring two-dimensional images of the lesion area under each focal plane. Then, the defocus degree of each two-dimensional image is calculated using a point spread function, thereby determining the defocus degree of each pixel, and using the defocus degree as the image feature corresponding to the pixel. After determining the target image and at least some pixels from each two-dimensional image, the coordinates of at least some pixels in the target image are determined. Then, the coordinates of these pixels are matched with the coordinates of pixels in other two-dimensional images to determine the matched pixels. Finally, based on the defocus degree of some pixels in the target image and the defocus degree of the matched pixels, the distance from some pixels in the target image to the endoscope is determined using a pre-trained defocus ranging model. The depth map of the target image is determined based on the distance from some pixels in the target image to the endoscope. The target image and its depth map are fused to generate a three-dimensional image of the lesion area.
[0120] In one or more embodiments of this specification, the server can determine the endoscope position for acquiring two-dimensional images based on a preset endoscope movement range. The endoscope is then moved, and when it reaches the determined endoscope position, a two-dimensional image of the lesion area is acquired at a fixed focusing distance. Subsequently, the sharpness of each two-dimensional image is evaluated to determine the sharpness corresponding to each pixel in each two-dimensional image, and the sharpness is used as the image feature corresponding to the pixel. At least some pixels in the target image are then matched with feature points in other two-dimensional images to determine the matching pixels in those other two-dimensional images. Finally, for the target image, based on the sharpness of the determined partial pixels and the sharpness of the pixels matching those partial pixels in other two-dimensional images, the focal length at which the sharpness of some pixels is highest in all two-dimensional images is determined. Then, based on the determined focal length of some pixels, the distance from some pixels to the endoscope is determined. Based on the distance from some pixels in the target image to the endoscope, the depth map of the target image is determined. The target image and its depth map are fused to generate a three-dimensional image of the lesion area.
[0121] In one or more embodiments of this specification, the server can determine the endoscope position for acquiring two-dimensional images based on a preset endoscope movement range. The endoscope is then moved, and when it reaches the determined endoscope position, a two-dimensional image of the lesion area is acquired at a fixed focusing distance. Subsequently, the defocus degree of each two-dimensional image is calculated using a point spread function to determine the defocus degree of each pixel, which is then used as the image feature corresponding to the pixel. At least some pixels in the target image are then matched with feature points in other two-dimensional images to determine the matching pixels. Finally, based on the defocus degree of some pixels in the target image and the defocus degree of the matching pixels, a pre-trained defocus ranging model is used to determine the distance from some pixels in the target image to the endoscope. Based on the distance from some pixels in the target image to the endoscope, a depth map of the target image is determined. The target image and its depth map are then fused to generate a three-dimensional image of the lesion area.
[0122] The above describes a method for generating a three-dimensional image according to one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding three-dimensional image generating apparatus, such as... Figure 4 As shown.
[0123] Figure 4 This is a schematic diagram of a three-dimensional image generation device provided in this specification, specifically including:
[0124] The acquisition module 400 is used to acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes;
[0125] The first determining module 402 is used to determine the image features corresponding to each pixel in each two-dimensional image;
[0126] The second determining module 404 is used to determine a target image from each of the two-dimensional images, and to determine at least some pixels in the target image that match pixels in other two-dimensional images;
[0127] The third determining module 406 is used to determine the distance from the partial pixels in the target image to the endoscope based on the image features of the partial pixels in the target image and the image features of the matching pixels.
[0128] The fourth determining module 408 is used to determine the depth map of the target image based on the distance from the partial pixels to the endoscope;
[0129] The fusion module 410 is used to fuse the target image and the depth map to generate a three-dimensional image of the lesion area.
[0130] Optionally, the acquisition module 400 is specifically used to determine each focal plane for acquiring two-dimensional images according to a preset focal length adjustment range, adjust the camera parameters of the endoscope, and acquire two-dimensional images of the lesion area captured under each focal plane.
[0131] Optionally, the acquisition module 400 is further configured to determine the endoscope position for acquiring two-dimensional images according to a preset endoscope movement range, move the endoscope, and acquire two-dimensional images of the lesion area at a fixed focusing distance when the endoscope moves to the endoscope position.
[0132] Optionally, the first determining module 402 is specifically used to determine the degree of defocus of each pixel in each two-dimensional image by using a point spread function;
[0133] The third determining module 406 is specifically used to determine the distance from the partial pixels in the target image to the endoscope based on the degree of defocus of the partial pixels in the target image and the degree of defocus of the matched pixels, through a pre-trained defocus ranging model.
[0134] Optionally, the first determining module 402 is further configured to determine the sharpness of each pixel in each two-dimensional image;
[0135] The third determining module 406 is further configured to, for the target image, determine the focal length of the partial pixels when the resolution is highest in each two-dimensional image based on the determined resolution of the partial pixels and the resolution of the pixels that the partial pixels match in the other two-dimensional images, and determine the distance from the partial pixels to the endoscope based on the determined focal length of the partial pixels.
[0136] Optionally, the acquisition module 400 is further configured to determine a target image from each two-dimensional image, and determine at least some pixels in the target image, determine the coordinates of each pixel determined in the target image, and determine a pixel in other two-dimensional images that matches the coordinates of the pixel as the pixel that matches the pixel.
[0137] Optionally, the acquisition module 400 is further configured to determine a target image from each two-dimensional image, determine at least some pixels in the target image, perform feature point matching of at least some pixels in the target image with other two-dimensional images, and determine the matching pixels in the other two-dimensional images.
[0138] Optionally, the trained deep model includes at least an encoding layer, a fusion layer, and a decoding layer;
[0139] The fourth determining module 408 is specifically used to input each two-dimensional image into the encoding layer of the depth model to obtain the image features of each two-dimensional image output by the encoding layer, input the image features of each two-dimensional image into the fusion layer to obtain the fused image features of each two-dimensional image, input the fused image features into the decoding layer to obtain the depth map output by the decoding layer, and determine the depth map of the target image based on the distance from the partial pixels to the endoscope and the depth map output by the decoding layer.
[0140] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 A method for generating three-dimensional images is provided.
[0141] This instruction manual also provides Figure 5 The diagram shows a schematic structural representation of the electronic device. Figure 5 As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The aforementioned three-dimensional image generation method.
[0142] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0143] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0144] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0145] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0146] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0147] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0148] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0149] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0150] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0151] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0152] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0153] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0154] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0155] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0157] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0158] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for generating a three-dimensional image, characterized in that, include: Acquire two-dimensional images of the lesion area taken by endoscopy at different focal planes; Determine the image features corresponding to each pixel in each two-dimensional image; Determine the target image from each of the two-dimensional images, and determine at least some pixels in the target image that match pixels in other two-dimensional images; Based on the image features of the partial pixels in the target image and the image features of the matching pixels, the distance from the partial pixels in the target image to the endoscope is determined; The depth map of the target image is determined based on the distance from the partial pixels to the endoscope; The target image and the depth map are fused to generate a three-dimensional image of the lesion area.
2. The method as described in claim 1, characterized in that, Acquiring two-dimensional images of the lesion region taken endoscopically at different focal planes, specifically including: Based on the preset focal length adjustment range, determine each focal plane for acquiring the two-dimensional image; Adjust the camera parameters of the endoscope to acquire two-dimensional images of the lesion area taken at each focal plane.
3. The method as described in claim 1, characterized in that, Acquiring two-dimensional images of the lesion region taken endoscopically at different focal planes, specifically including: The position of the endoscope for acquiring two-dimensional images is determined based on the preset endoscope movement range. The endoscope is moved, and when the endoscope is moved to the endoscope position, a two-dimensional image of the lesion area is acquired at a fixed focusing distance.
4. The method as described in claim 1, characterized in that, Determine the image features corresponding to each pixel in each two-dimensional image, specifically including: For each two-dimensional image, the degree of defocus of each pixel in the two-dimensional image is determined by the point spread function; Based on the image features of the selected pixels in the target image and the image features of the matched pixels, the distance from the selected pixels in the target image to the endoscope is determined, specifically including: Based on the degree of defocus of the partial pixels in the target image and the degree of defocus of the matched pixels, the distance from the partial pixels in the target image to the endoscope is determined by a pre-trained defocus ranging model.
5. The method as described in claim 1, characterized in that, Determine the image features corresponding to each pixel in each two-dimensional image, specifically including: For each two-dimensional image, determine the sharpness of each pixel in that two-dimensional image; Based on the image features of the selected pixels in the target image and the image features of the matched pixels, the distance from the selected pixels in the target image to the endoscope is determined, specifically including: For the target image, based on the determined sharpness of the partial pixels and the sharpness of the pixels that match the partial pixels in the other two-dimensional images, the focal length at which the partial pixels have the highest sharpness in each two-dimensional image is determined. Based on the determined focal length of the selected pixels, the distance from the selected pixels to the endoscope is determined.
6. The method as described in claim 2, characterized in that, Determining a target image from each of the two-dimensional images, and determining at least some pixels in the target image that match pixels in other two-dimensional images, specifically includes: Determine the target image from each two-dimensional image, and determine at least a portion of the pixels in the target image; For each pixel identified in the target image, determine the coordinates of that pixel; In other two-dimensional images, identify the pixel whose coordinates match that of the given pixel, and use that pixel as the matching pixel.
7. The method as described in claim 2 or 3, characterized in that, Determining a target image from each of the two-dimensional images, and determining at least some pixels in the target image that match pixels in other two-dimensional images, specifically includes: Determine the target image from each two-dimensional image, and determine at least a portion of the pixels in the target image; At least some pixels in the target image are matched with feature points in other two-dimensional images to determine the matching pixels in the other two-dimensional images.
8. The method as described in claim 1, characterized in that, A trained deep model includes at least an encoding layer, a fusion layer, and a decoding layer; The depth map of the target image is determined based on the distance from the selected pixels to the endoscope, specifically including: Each of the two-dimensional images is input into the coding layer of the depth model to obtain the image features of each of the two-dimensional images output by the coding layer; The image features of each two-dimensional image are input into the fusion layer to obtain the fused image features of each two-dimensional image; The fused image features are input into the decoding layer to obtain the depth map output by the decoding layer; The depth map of the target image is determined based on the distance from the partial pixels to the endoscope and the depth map output by the decoding layer.
9. A three-dimensional image generation device, characterized in that, include: The acquisition module is used to acquire two-dimensional images of the lesion area taken by the endoscope at different focal planes; The first determining module is used to determine the image features corresponding to each pixel in each two-dimensional image; The second determining module is used to determine a target image from each of the two-dimensional images, and to determine at least some pixels in the target image that match pixels in other two-dimensional images; The third determining module is used to determine the distance from the partial pixels in the target image to the endoscope based on the image features of the partial pixels in the target image and the image features of the matching pixels; The fourth determining module is used to determine the depth map of the target image based on the distance from the partial pixels to the endoscope; The fusion module is used to fuse the target image and the depth map to generate a three-dimensional image of the lesion area.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 8.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Device and method for measuring scene depth
CN102997891A
Coding re-focusing calculation shooting method and device
CN103209307A
A method and system for three-dimensional reconstruction
CN109087395A
Non-contact measurement system and method
CN113808019A
Endoscope medical image three-dimensional display method and system
CN117409928A