A method and system for coloring black and white mode video of laser knife surgery

By building a black and white video color recovery system for laser knife surgery based on an adversarial neural network, the problem of abnormal picture exposure during laser knife surgery is solved, real-time color image recovery of laser knife surgery video is realized, color recovery performance is improved, and computational amount and memory consumption are reduced.

CN115170385BActive Publication Date: 2025-05-06BLUERAY MEDICAL TECHNOLOGIES LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210774461.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-01
Publication Date
2025-05-06
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

During laser knife surgery, the screen exposure is abnormal when the laser knife is working under the endoscopic, and effective color video images cannot be obtained.

Method used

A deep learning algorithm based on adversarial neural network is used to build a black and white video color recovery system for laser knife surgery. By constructing a generator and discriminator network model, image feature extraction is used using octave convolution and residual modules, and color video recovery is achieved through HSI fusion.

Benefits of technology

Real-time color image recovery of laser knife surgical videos is realized, the color recovery performance of surgical videos is improved, the calculation amount and memory consumption are reduced, and the color recovery of different video specifications is adapted to color recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170385B_ABST
    Figure CN115170385B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for colorizing a black-and-white mode video of laser knife surgery, comprising: 1) collecting color videos of medical endoscopic surgery, selecting key frame color images to convert into key frame grayscale images of original resolution, and making a video colorization sample library; 2) constructing and training a network model to obtain a pre-trained color image; 3) calculating the loss between the pre-trained color image and the surgical color image to obtain a loss function; 4) optimizing the network model; 5) real-time collection of black-and-white mode videos in laser knife surgery, and then inputting the black-and-white images into the network model in sequence, and performing HSI fusion in sequence on the obtained low-resolution color image and the original resolution grayscale image to obtain a color video. The present invention solves the technical problem of abnormal exposure of the screen when the laser knife is working under an endoscope, and has the effects of low calculation amount, low memory consumption, and fast model color recovery speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of laser knife surgery black and white mode video color restoration and medical-engineering integration, and in particular relates to a laser knife surgery black and white mode video colorization method and system based on an adversarial neural network. Background Art

[0002] With the aging of the population and changes in social and environmental factors, the incidence of cancer, polyps, chronic inflammation, etc. in various cavity organs (including urinary tract, digestive tract, female reproductive tract, nasopharyngeal tract, respiratory tract) of human beings shows a continuous upward trend. Early detection and early treatment are expected to cure such patients, especially the consensus of various early cancers. Although most of the above-mentioned early cancer treatment methods of cavity organs have entered the minimally invasive era, electrosurgical instruments (electric knives) are still the mainstream weapons in the hands of doctors. Generally speaking, due to the common problems of low cutting accuracy, large surgical side effects, and high technical requirements of electric knives, most minimally invasive technologies are limited to the promotion and application of county medical institutions, and cannot serve most ordinary patients well. Although the existing endoscopic submucosal dissection (ESD) system has been widely used in minimally invasive treatment of early cancer of the digestive tract, namely esophagus, stomach, and colon, the design of the relevant minimally invasive instruments is only suitable for digestive endoscopes and cannot be used with minimally invasive equipment for other cavity diseases. The expensive equipment and consumables also make primary medical institutions and ordinary patients discouraged.

[0003] Laser knives have shown obvious advantages in the treatment of some diseases, such as high cutting efficiency, less bleeding, and less secondary damage, and are constantly replacing traditional surgical instruments. Currently, the most representative laser energy platforms on the market are visible light lasers and near-infrared lasers, which can reduce tissue thermal damage and achieve rapid recovery after surgery. However, these devices generally have disadvantages such as high prices, bulky equipment, and high working noise, which seriously affect their promotion and application in grassroots hospitals.

[0004] During laser surgery, the endoscope lens will be overexposed because the powerful laser is so bright. For traditional endoscopes, we can use laser pulse scalpels to solve this problem. We use laser pulses to perform surgery. During the intervals between laser pulses, the camera can work normally to obtain normal surgical images because there is no laser influence. Then use the normal image obtained to restore the image when the laser pulse is working, so as to obtain a normal surgical video without laser. However, in some endoscopic cases, even if the laser knife works for a short time, the image of the color camera used is still exposed and no effective image can be obtained. Summary of the invention

[0005] In order to solve the problem that the image exposure is abnormal and effective images cannot be obtained when the laser knife is working under an endoscope, the present invention proposes a method and system for colorizing black and white mode videos of laser knife surgery, develops a deep learning algorithm using a convolutional neural network for color restoration of black and white mode videos of laser knife surgery systems, and develops a real-time imaging transcription system that can be adapted to a flexible electronic endoscope and is used for color restoration of black and white mode videos of laser knife surgery.

[0006] In order to achieve the above object, the technical solution provided by the present invention is:

[0007] A method for coloring a black-and-white video of laser knife surgery, which is special in that it includes the following steps:

[0008] Step 1) Get the video colorization sample library:

[0009] 1.1) Collect color videos of medical endoscopic surgery;

[0010] 1.2) Selecting a key frame color image from the color video, and converting the key frame color image into a key frame grayscale image of original resolution; the key frame color image is a color image containing the surgical site;

[0011] 1.3) extracting the surgical scenes from the key frame color image and the key frame grayscale image as sub-images by downsampling, and using the sub-images to create a video colorization sample library; the surgical scenes do not contain the display information in the key frame color image;

[0012] Step 2) constructing and training a network model to obtain a pre-trained colored image;

[0013] Step 3) Calculate the loss between the pre-trained colored image and the surgical color image to obtain a loss function;

[0014] Step 4) Optimize the network model:

[0015] Using the loss function as the optimization objective function, allowing the network model to participate in the gradient back propagation process in network optimization, thereby achieving optimization of the network model;

[0016] Step 5) Coloring of surgical images:

[0017] The black-and-white mode video of the laser knife surgery is collected in real time, and then the down-sampled multiple frames of black-and-white images are sequentially input into the network model, and the low-resolution colored image obtained by the network model and the original resolution grayscale image are sequentially subjected to HSI fusion to obtain a color video composed of multiple frames of original resolution color images.

[0018] The above step 2 specifically includes:

[0019] 2.1) Construct a laser knife surgery black and white mode video colorization network model based on adversarial neural network;

[0020] The network model includes a generator network model and a discriminator network model; the generator network model includes an input layer, a second combination module and an output layer connected in sequence; the discriminator network model is a Resnet model;

[0021] The second combination module includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence;

[0022] The convolution module contains three convolutional layers connected in sequence, and the parameters of the three convolutional layers are 32×9×9, stride=1; 64×3×3, stride=2; 128×3×3, stride=2;

[0023] The residual module includes five residual layers connected in sequence, and the equivalent convolutional layer structure of each residual layer is two 3×3 convolutional layers connected in sequence, a Batch Norm layer, and a ReLU layer;

[0024] The deconvolution module contains three deconvolution layers connected in sequence. The parameters of the three deconvolution layers are 64×3×3, 32×3×3, 3×9×9, stride=1;

[0025] 2.2) Use the grayscale images of the video colorization sample library generated in step 1 to train the constructed network model to obtain pre-trained colorization images.

[0026] The above discriminator network model is ResNet-18, and the corresponding feature extraction network includes at least the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer; wherein,

[0027] The size of the first convolution kernel conv1 is 7×7×64, the stride is 2, the Max-Pooling is 2×2, and the stride is 2;

[0028] The second layer of convolution kernel conv2 is 2 64×64×256 kernels with stride 1;

[0029] The third convolution kernel conv3 is 2 128×128×512;

[0030] The fourth layer convolution kernel conv4 is 2 256×256×1024;

[0031] The fifth convolution kernel conv5 is 512×512×2048.

[0032] The above loss function is specifically the loss function for colorizing black and white laser knife surgery videos:

[0033]

[0034] Among them, H represents the number of network layers;

[0035] l BCE The function expression is:

[0036]

[0037] l KL The function expression is:

[0038]

[0039] Where (i, j) is the pixel coordinate;

[0040] (M,N) is the width and height of the image;

[0041] G(i,j) is a true value;

[0042] S(i,j) is the predicted target color restored pixel value;

[0043] is the weight of the BCE loss function;

[0044] is the weight of the KL loss function.

[0045] The above step 1.3 is to extract a sub-image of a specified size by locating the template of the surgical screen.

[0046] The resolution of the above sub-image is preferably 100×100.

[0047] A laser knife surgery black and white mode video coloring system, which is special in that it includes the following modules:

[0048] The data set acquisition module is used to collect medical endoscopic surgery videos, select key frames with rich colors from the video stream, and obtain their grayscale images through processing as the true value and image input of the training set to produce a training sample data set;

[0049] The data set preprocessing module is used to extract the core content of the image from the prepared training sample data set to solve the color restoration problem encountered during the test, and to perform unified sampling processing on the training sample data set to solve the color restoration efficiency problem encountered during the test;

[0050] The network model training module is used to construct a laser knife surgery video color restoration network model based on a generative adversarial network, and use the training samples generated in the data set preprocessing module to train the constructed video color restoration network model to generate a predicted color image; the video color restoration network model includes a generator and a discriminator; the network model structure of the generator includes an input layer, a second combination module, and an output layer connected in sequence; the network model of the discriminator is a Resnet model;

[0051] A loss function calculation module, used to calculate the loss between the pre-trained color restoration result and the true value of color endoscopic surgery;

[0052] The network optimization module is used to use the loss function as the optimization objective function, so that the video color restoration network model participates in the gradient back propagation process in the network optimization, and realizes the optimization of the color restoration of the black and white mode video of the laser knife surgery;

[0053] The multi-resolution fusion module is used to fuse low-resolution color surgical images and restore them to color surgical images of original resolution size.

[0054] The second combined module specifically includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence;

[0055] The second combination module includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence;

[0056] The convolution module contains three convolutional layers connected in sequence, and the parameters of the three convolutional layers are 32×9×9, stride=1; 64×3×3, stride=2; 128×3×3, stride=2;

[0057] The residual module includes five residual layers connected in sequence, and the equivalent convolutional layer structure of each residual layer is two 3×3 convolutional layers connected in sequence, a Batch Norm layer, and a ReLU layer;

[0058] The deconvolution module contains three deconvolution layers connected in sequence. The parameters of the three deconvolution layers are 64×3×3, 32×3×3, 3×9×9, stride=1.

[0059] Compared with the prior art, the advantages and beneficial effects of the present invention are:

[0060] 1. Since the existing laser knife surgery directly uses color video, all 60 frames of images contain lasers, and the images synthesized from the three images are all overexposed, so the required normal surgical images cannot be obtained. The present invention collects clean black and white mode videos in laser knife surgery in real time, and then sequentially inputs multiple frames of downsampled black and white images with small post-processing amount into the network model, and sequentially performs HSI fusion on the low-resolution colored images obtained by the network model and the original resolution grayscale images, and finally obtains a color video composed of multiple frames of original resolution color images.

[0061] 2. The present invention uses a residual network to connect the convolution module and the deconvolution module. Compared with the non-residual network, the model training speed including the residual network structure is faster.

[0062] 3. The present invention uses octave convolution to decompose the output feature map of the convolution layer into high-frequency and low-frequency feature maps stored in different groups, which can safely reduce the spatial resolution of the low-frequency group and reduce spatial redundancy. At the same time, low-frequency convolution operations on low-frequency information can effectively expand the receptive field in the pixel space. Therefore, compared with traditional methods, it can further reduce the amount of calculation and memory consumption.

[0063] 4. The present invention first processes the surgical video, including adaptive picture extraction, downsampling and other operations, to obtain a unified low-resolution color restoration data set for the training of the color restoration network, and then performs HSI fusion of the obtained low-resolution color result with the black-and-white image of the original resolution to obtain the color surgical image of the original resolution. Training the network with a low-resolution data set can improve the speed of network model training and model color restoration. At the same time, using low-resolution images as model input is more conducive to the learning of low-dimensional color features, which improves the color restoration performance. The laser knife surgery black-and-white mode video colorization method and system based on adversarial neural network of the present invention can meet the requirements of real-time surgical video color restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a flow chart of an embodiment of the present invention;

[0065] Figure 2 This is a schematic diagram of the network structure in the present invention;

[0066] Figure 3 Schematic diagram of the structure of a residual layer in the residual module;

[0067] Figure 4 Schematic diagram of another structure of a residual layer in a residual module;

[0068] Figure 5 Schematic diagram of the octave convolution structure. DETAILED DESCRIPTION

[0069] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0070] The present invention replaces the traditional convolution with octave convolution and applies it to the black-and-white mode video color restoration network, thereby proposing a new laser knife surgery black-and-white mode video colorization network based on an adversarial neural network for color restoration. This new color restoration network is used on a preprocessed data set to better utilize local and global context information to improve the color restoration effect. This new laser knife surgery black-and-white mode video colorization network based on an adversarial neural network extracts features through convolution operations to extract color feature information of endoscopic surgery data at different scales. The convolution module uses octave convolution instead of ordinary convolution to extract local features. Compared with traditional methods, the method using octave convolution can further reduce the amount of calculation and memory consumption, while improving the performance of color restoration, and further improves the network performance by connecting the residual module with the convolution module and the deconvolution module to meet the performance requirements of fact processing.

[0071] The specific steps include:

[0072] 1) Dataset acquisition: Collect color medical endoscopic surgery videos, select key frames from the video stream to make training samples, and obtain black and white grayscale images of key frames to achieve color restoration of black and white mode videos of laser knife surgery;

[0073] The collected clinical endoscopic surgery color video is screened to obtain key frames containing rich color information, and the key frames are subjected to surgical screen extraction to produce a suitable surgical screen mask for extracting the surgical screen content, thereby preventing other invalid areas from interfering with the surgical video color restoration. In the implementation process of the present invention, the training set image is a sub-image of 100×100 size, and the sub-image of a specified size is extracted by locating the surgical screen mask as a training sample to overcome the problem of inconsistent specifications of different videos.

[0074] 2) Training the network model: construct a new laser knife surgery black and white mode video colorization model based on an adversarial neural network, use the training samples generated in step 2 to train the constructed network model, and generate color laser knife surgery images;

[0075] refer to Figure 2 The laser knife surgery black and white mode video colorization network based on the adversarial neural network constructed by the present invention is further described in detail:

[0076] The laser knife surgery black and white mode video colorization network model based on adversarial neural network has two parts: one is the generator G, and the other is the discriminator D. The network model structure of the generator includes the following connected in sequence: input layer - second combination module - output layer; the network model of the discriminator is Resnet;

[0077] The second module of the generator model G includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence; the convolution module includes three convolution layers connected in sequence, and the parameters are 32×9×9, stride=1; 64×3×3, stride=2; 128×3×3, stride=2; the residual module includes five residual layers connected in sequence, and its structure is as follows Figure 3 or Figure 4 As shown in Figure 1, each residual layer contains two 3×3 convolutional layers with the same number of filters on each layer; the deconvolution module contains three deconvolutional layers connected in sequence, with parameters of 64×3×3, 32×3×3, 3×9×9, stride=1.

[0078] Construct the discriminator model D to build ResNet-18, that is, use 18 layers for feature extraction. The feature extraction network includes at least the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer. The parameters are set as follows: the size of the first convolution kernel conv1 is 7×7×64, the stride is 2, the Max-Pooling is 2×2, and the stride is 2; conv2 is 2 64×64×256, the stride is 1, conv3 is 2 128×128×512; conv4 is 2 256×256×1024, and conv5 is 512×512×2048.

[0079] Figure 5 This is the implementation details of the octave convolution, which consists of four calculation paths corresponding to four items: f(X H ; W H →H )、upsamplef(X L ; W L→H )、f(X L ; W L→L )、fpoolX H ,2;W H→L ,The two solid paths correspond to the information update of high-frequency and low-frequency feature maps, and the two dotted paths facilitate the information exchange between the two octave convolutions.

[0080] Octave convolution is similar to the decomposition of spatial frequency components of natural images. It decomposes the output feature map of the convolution layer into high- and low-frequency feature maps stored in different groups. Therefore, it can safely reduce the spatial resolution of the low-frequency group and reduce spatial redundancy by sharing information between adjacent positions. In addition, the low-frequency convolution operation of the low-frequency information can effectively expand the receptive field in the pixel space. Therefore, the use of octave convolution can further reduce the computational and memory overhead of the network.

[0081] 3) Calculate the loss function: Calculate the loss between the color restoration result predicted by the pre-trained surgical image and the true value of the color surgical image;

[0082] During the training process, the present invention adopts a hierarchical training supervision strategy to replace the standard top-level supervision training and deep supervision scheme. The loss function of the black and white mode video colorization network in the present invention is as follows:

[0083]

[0084] Where H represents the number of network layers, l BCE , l KL The function expressions are:

[0085]

[0086] Where (i, j) is the pixel coordinate;

[0087] (M,N) is the width and height of the image;

[0088] G(i,j) and S(i,j) are the true value and predicted target color restored pixel value respectively;

[0089] and are the weights of the BCE loss function and the KL loss function respectively.

[0090] For each layer, we used the standard BCE loss function and KL loss function to calculate the loss. By adding a pair of probability prediction matching losses (i.e., KL loss function) between any two layers, multi-layer interactions between different layers are promoted. The optimization objectives of the loss functions at different layers are consistent, thus ensuring the robustness and generalization of the model.

[0091] 4) Optimize the network: Use the loss function as the optimization objective function, and make the convolutional neural network participate in the gradient back propagation process in the network optimization to achieve the optimization of the color restoration of the laser knife surgery system video.

[0092] 5) Multi-resolution fusion: The low-resolution color surgical image obtained by the color restoration network is fused with the original resolution surgical grayscale image through HSI to obtain the color surgical image with the original resolution.

Claims

1. A method for colorizing a black-and-white video of laser knife surgery, characterized in that: The following steps are involved: Step 1) Get the video colorization sample library: 1.1) Collect color videos of medical endoscopic surgery; 1.2) Selecting a key frame color image from the color video, and converting the key frame color image into a key frame grayscale image of original resolution; the key frame color image is a color image containing the surgical site; 1.3) extracting the surgical scenes from the key frame color image and the key frame grayscale image as sub-images by downsampling, and using the sub-images to create a video colorization sample library; the surgical scenes do not contain the display information in the key frame color image; Step 2) construct and train the network model to obtain the pre-trained colorized image, specifically including: 2.1) Construct a laser knife surgery black and white mode video colorization network model based on adversarial neural network; The network model includes a generator network model and a discriminator network model; the generator network model includes an input layer, a second combination module and an output layer connected in sequence; the discriminator network model is a Resnet model; The second combination module includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence; The convolution module contains three convolutional layers connected in sequence, and the parameters of the three convolutional layers are 32×9×9, stride=1; 64×3×3, stride=2; 128×3×3, stride=2; The residual module includes five residual layers connected in sequence, and the equivalent convolutional layer structure of each residual layer is two 3×3 convolutional layers connected in sequence, a Batch Norm layer, and a ReLU layer; The deconvolution module contains three deconvolution layers connected in sequence. The parameters of the three deconvolution layers are 64×3×3, 32×3×3, 3×9×9, stride=1; 2.2) Using the grayscale images of the video colorization sample library generated in step 1 to train the constructed network model, to obtain a pre-trained colorization image; Step 3) Calculate the loss between the pre-trained colored image and the surgical color image to obtain a loss function; Step 4) Optimize the network model: Using the loss function as the optimization objective function, allowing the network model to participate in the gradient back propagation process in network optimization, thereby achieving optimization of the network model; Step 5) Coloring of surgical images: The black-and-white mode video of the laser knife surgery is collected in real time, and then the down-sampled multiple frames of black-and-white images are sequentially input into the network model, and the low-resolution colored image obtained by the network model and the original resolution grayscale image are sequentially subjected to HSI fusion to obtain a color video composed of multiple frames of original resolution color images.

2. According to claim 1, a method for colorizing a black and white mode video of laser knife surgery, characterized in that: The discriminator network model is ResNet-18, and the corresponding feature extraction network includes at least the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer; wherein, The size of the first convolution kernel conv1 is 7×7×64, the stride is 2, the Max-Pooling is 2×2, and the stride is 2; The second layer of convolution kernel conv2 is 2 64×64×256 kernels with stride 1; The third convolution kernel conv3 is 2 128×128×512; The fourth layer convolution kernel conv4 is 2 256×256×1024; The fifth convolution kernel conv5 is 512×512×2048.

3. According to claim 1, a method for colorizing a black and white mode video of laser knife surgery, characterized in that: The loss function is the colorization loss function of black and white laser knife surgery video: Among them, H represents the number of network layers; l BCE The function expression is: l KL The function expression is: Where (i, j) is the pixel coordinate; (M,N) is the width and height of the image; G(i,j) is a true value; S(i,j) is the predicted target color restored pixel value; is the weight of the BCE loss function; is the weight of the KL loss function.

4. A method for colorizing a black-and-white video of laser knife surgery according to any one of claims 1 to 3, characterized in that: Step 1.3 is to extract a sub-image of a specified size by locating the template of the surgical screen.

5. According to claim 4, a method for colorizing a black and white mode video of laser knife surgery, characterized in that: The resolution of the sub-image is 100×100.

6. A laser knife surgery black and white mode video colorization system, characterized in that: Includes the following modules: The data set acquisition module is used to collect medical endoscopic surgery videos, select key frames with rich colors from the video stream, and obtain their grayscale images through processing as the true value and image input of the training set to produce a training sample data set; The data set preprocessing module is used to extract the core content of the image from the prepared training sample data set to solve the color restoration problem encountered during the test, and to perform unified sampling processing on the training sample data set to solve the color restoration efficiency problem encountered during the test; The network model training module is used to construct a laser knife surgery video color restoration network model based on a generative adversarial network, and use the training samples generated in the data set preprocessing module to train the constructed video color restoration network model to generate a predicted color image; the video color restoration network model includes a generator and a discriminator; the network model structure of the generator includes an input layer, a second combination module, and an output layer connected in sequence; the network model of the discriminator is a Resnet model; A loss function calculation module, used to calculate the loss between the pre-trained color restoration result and the true value of color endoscopic surgery; The network optimization module is used to use the loss function as the optimization objective function, so that the video color restoration network model participates in the gradient back propagation process in the network optimization, and realizes the optimization of the color restoration of the black and white mode video of the laser knife surgery; The multi-resolution fusion module is used to fuse low-resolution color surgical images and restore them to color surgical images of original resolution size.

7. According to claim 6, a laser knife surgery black and white mode video colorization system is characterized by: The second combination module includes a reflection filling layer, a convolution module, a residual module layer and a deconvolution module connected in sequence; The convolution module contains three convolutional layers connected in sequence, and the parameters of the three convolutional layers are 32×9×9, stride=1; 64×3×3, stride=2; 128×3×3, stride=2; The residual module includes five residual layers connected in sequence, and the equivalent convolutional layer structure of each residual layer is two 3×3 convolutional layers connected in sequence, a Batch Norm layer, and a ReLU layer; The deconvolution module contains three deconvolution layers connected in sequence. The parameters of the three deconvolution layers are 64×3×3, 32×3×3, 3×9×9, stride=1.

Citation Information

Patent Citations

  • Coloring method and device for black and white video, equipment and storage medium

    CN112884866A

  • Black-and-white video coloring method and device, storage medium and terminal

    CN113421312A