Model training based on recurrent laryngeal nerve, recurrent laryngeal nerve identification method and device
Patent Information
- Application Number
- CN202210790922.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-07-05
AI Technical Summary
当前甲状腺手术中,一般由喉返神经监测仪辅助医生完成手术,但是医生需频繁操作喉返神经监测仪,这样导致医生的手术过程被频繁打断,对手术的连贯性不够友好
[0048]As can be seen from the above, in the solution provided by the embodiments of the present invention, the recurrent laryngeal nerve recognition model to be trained includes a residual network, a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer. These network layers process the sample endoscopic image sequentially, thereby segmenting the recurrent laryngeal nerve region from the sample endoscopic image. When training the recurrent laryngeal nerve recognition model, not only the aforementioned recurrent laryngeal nerve region is referenced, but also the sample markers. Since the sample markers identify the actual location of the recurrent laryngeal nerve in the sample endoscopic image, the differences between the recurrent laryngeal nerve region and the sample markers can be determined to identify the differences in the recurrent laryngeal nerve recognition model when segmenting the recurrent laryngeal nerve region. Therefore, the model parameters of the recurrent laryngeal nerve recognition model can be adjusted based on these differences, enabling the recurrent laryngeal nerve recognition model to learn the characteristics of the actual recurrent laryngeal nerve region, thereby improving the accuracy of segmenting the recurrent laryngeal nerve region. In other words, a model capable of segmenting the recurrent laryngeal nerve region based on an image is trained.
Smart Images

Figure CN117392475B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model training, recurrent laryngeal nerve recognition method and apparatus based on the recurrent laryngeal nerve. Background Technology
[0002] Recurrent laryngeal nerve injury is a common complication of thyroid surgery. If the recurrent laryngeal nerve is damaged, it can lead to hoarseness, difficulty breathing, or even suffocation. Furthermore, the nerve is difficult to regenerate after rupture, making the consequences of recurrent laryngeal nerve injury during surgery extremely serious. Currently, recurrent laryngeal nerve monitoring is generally used to assist surgeons in thyroid surgery. However, frequent operation of this monitoring device disrupts the surgical process and compromises its continuity.
[0003] In view of the above, and given the rapid development and widespread application of artificial intelligence technology in recent years, it is necessary to provide a model training method based on the recurrent laryngeal nerve to obtain a model that can identify the recurrent laryngeal nerve using images. Summary of the Invention
[0004] The purpose of this invention is to provide a model training method and apparatus for recurrent laryngeal nerve recognition, based on the recurrent laryngeal nerve, so as to obtain a model capable of recognizing the recurrent laryngeal nerve using images. The specific technical solution is as follows:
[0005] According to one aspect of the present invention, a model training method based on the recurrent laryngeal nerve is provided, the method comprising:
[0006] Obtain endoscopic images of the sample and identify sample markers for the recurrent laryngeal nerve region within the endoscopic images;
[0007] The sample endoscope image is input into the residual network of the recurrent laryngeal nerve recognition model to be trained to obtain the image features of the sample endoscope image. The recurrent laryngeal nerve recognition model further includes: a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer.
[0008] The image features are input into the pooling layer for global average pooling to obtain multi-scale pooling results.
[0009] The pooling results at each scale are input into the first convolutional layer, and the output of the first convolutional layer is input into the first upsampling layer to obtain a first feature with the same size as the image feature.
[0010] The first features and the image features are input into the connection layer and connected to obtain the second features;
[0011] The second feature is input into the second convolutional layer, and the output of the second convolutional layer is input into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the endoscopic image of the sample;
[0012] The recurrent laryngeal nerve recognition model is trained based on the obtained recurrent laryngeal nerve region and sample labels.
[0013] In one embodiment of the present invention, the sample label is a binary image with the same size as the endoscopic image of the sample;
[0014] The output information of the second upsampling layer is a binary image with the same size as the sample endoscope image.
[0015] In one embodiment of the present invention, training the recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels includes:
[0016] Based on the obtained recurrent laryngeal nerve region and sample labels, the loss value of the recurrent laryngeal nerve recognition model is obtained;
[0017] Based on the preset initial learning rate and the loss value, the model parameters of the recurrent laryngeal nerve recognition model are adjusted according to the poly learning rate decay strategy and the gradient descent criterion set by the Adam optimizer, thereby achieving model training.
[0018] In one embodiment of the present invention, before obtaining the sample image, the method further includes:
[0019] Obtain raw endoscopic images;
[0020] The scaling ratio of the original endoscopic image is obtained, and the scaling ratio is determined based on the size of the original endoscopic image and the preset size of the input image of the recurrent laryngeal nerve recognition model;
[0021] The original endoscopic image is scaled according to the specified ratio to obtain a first image;
[0022] If the size of the first image is smaller than the preset size, the first image is subjected to pixel expansion processing according to the preset size to obtain the second image;
[0023] The pixel values of each pixel in the second image are normalized to obtain the third image;
[0024] The sample image is obtained based on the third image.
[0025] In one embodiment of the present invention, obtaining the sample image based on the third image includes:
[0026] Obtain the mean and variance of the pixel values of each pixel in the third image;
[0027] The difference is obtained by subtracting the mean value from the pixel value of each pixel in the third image, and then dividing the difference by the variance to obtain the sample image.
[0028] According to another aspect of the present invention, a method for identifying the recurrent laryngeal nerve is provided, the method comprising:
[0029] Obtain an image of the endoscope to be identified;
[0030] The endoscope image to be identified is input into a pre-trained recurrent laryngeal nerve recognition model to identify the recurrent laryngeal nerve, thereby obtaining the recurrent laryngeal nerve region in the endoscope image to be identified. The recurrent laryngeal nerve recognition model is a model trained according to any one of the above-mentioned recurrent laryngeal nerve-based model training methods.
[0031] According to another aspect of the present invention, a model training device based on the recurrent laryngeal nerve is provided, the device comprising:
[0032] The sample acquisition module is used to acquire endoscopic images of samples and to identify sample markers in the recurrent laryngeal nerve region within the endoscopic images of the samples.
[0033] The feature acquisition module is used to input the sample endoscope image into the residual network of the recurrent laryngeal nerve recognition model to be trained, and obtain the image features of the sample endoscope image. The recurrent laryngeal nerve recognition model further includes: a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer.
[0034] The feature pooling module is used to input the image features into the pooling layer for global average pooling to obtain multi-scale pooling results.
[0035] The pooling result processing module is used to input the pooling results of each scale into the first convolutional layer, and input the output result of the first convolutional layer into the first upsampling layer to obtain a first feature with the same size as the image feature;
[0036] A feature connection module is used to input each first feature and the image feature into the connection layer for connection to obtain a second feature;
[0037] The region segmentation module is used to input the second feature into the second convolutional layer and input the output of the second convolutional layer into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the endoscopic image of the sample;
[0038] The model training module is used to train the recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels.
[0039] According to another aspect of the present invention, a recurrent laryngeal nerve recognition device is provided, the device comprising:
[0040] The first image acquisition module is used to acquire the image of the endoscope to be identified;
[0041] The region recognition module is used to input the endoscope image to be recognized into a pre-trained recurrent laryngeal nerve recognition model to recognize the recurrent laryngeal nerve and obtain the recurrent laryngeal nerve region in the endoscope image to be recognized. The recurrent laryngeal nerve recognition model is a model trained by the recurrent laryngeal nerve-based model training device described in the foregoing embodiments.
[0042] According to another aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0043] Memory, used to store computer programs;
[0044] When the processor executes the program stored in the memory, it implements the method steps of the model training method or the recurrent laryngeal nerve recognition method based on the recurrent laryngeal nerve described in the above embodiments.
[0045] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, the method steps of the model training method or the recurrent laryngeal nerve recognition method based on the above embodiments are implemented.
[0046] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the method steps of the recurrent laryngeal nerve-based model training method or recurrent laryngeal nerve recognition method described in the above embodiments.
[0047] Beneficial effects of the embodiments of the present invention:
[0048] As can be seen from the above, in the solution provided by the embodiments of the present invention, the recurrent laryngeal nerve recognition model to be trained includes a residual network, a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer. These network layers process the sample endoscopic image sequentially, thereby segmenting the recurrent laryngeal nerve region from the sample endoscopic image. When training the recurrent laryngeal nerve recognition model, not only the aforementioned recurrent laryngeal nerve region is referenced, but also the sample markers. Since the sample markers identify the actual location of the recurrent laryngeal nerve in the sample endoscopic image, the differences between the recurrent laryngeal nerve region and the sample markers can be determined to identify the differences in the recurrent laryngeal nerve recognition model when segmenting the recurrent laryngeal nerve region. Therefore, the model parameters of the recurrent laryngeal nerve recognition model can be adjusted based on these differences, enabling the recurrent laryngeal nerve recognition model to learn the characteristics of the actual recurrent laryngeal nerve region, thereby improving the accuracy of segmenting the recurrent laryngeal nerve region. In other words, a model capable of segmenting the recurrent laryngeal nerve region based on an image is trained.
[0049] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0051] Figure 1 This is a schematic diagram of the structure of a recurrent laryngeal nerve recognition model provided in an embodiment of the present invention;
[0052] Figure 2 A schematic flowchart of a model training method based on the recurrent laryngeal nerve provided in an embodiment of the present invention;
[0053] Figure 3 A schematic diagram of a sample labeling method provided in an embodiment of the present invention;
[0054] Figure 4 This is a flowchart illustrating a method for obtaining sample images according to an embodiment of the present invention.
[0055] Figure 5 A flowchart illustrating a recurrent laryngeal nerve identification method provided in an embodiment of the present invention;
[0056] Figure 6 A schematic diagram of the recurrent laryngeal nerve provided in an embodiment of the present invention;
[0057] Figure 7This is a schematic diagram of a model training device based on the recurrent laryngeal nerve provided in an embodiment of the present invention;
[0058] Figure 8 This is a schematic diagram of the structure of a recurrent laryngeal nerve recognition device provided in an embodiment of the present invention;
[0059] Figure 9 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.
[0061] The recurrent laryngeal nerve is a vital nerve in the human body, and damage to it during surgery can have very serious consequences. To reduce the probability of recurrent laryngeal nerve injury, a recurrent laryngeal nerve monitor is typically used to assist surgeons during the procedure. However, this requires frequent operation of the monitor by the surgeon, which is not conducive to the continuity of the surgery. Therefore, a solution is needed that can reduce the frequency of operation by the surgeon while still effectively identifying the recurrent laryngeal nerve. In view of this, embodiments of the present invention provide a model training method and device based on the recurrent laryngeal nerve for recurrent laryngeal nerve identification.
[0062] First, combine Figure 1 and Figure 2 The model training method based on the recurrent laryngeal nerve provided in the embodiments of the present invention will be described.
[0063] See Figure 1 This provides a schematic diagram of the structure of a recurrent laryngeal nerve recognition model. From... Figure 1 As can be seen from the diagram, the recurrent laryngeal nerve recognition model includes: a residual network, a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer. Figure 1 The solid lines in the diagram illustrate the connections between the various network layers, while the dashed lines represent the data transmitted between them.
[0064] See Figure 2 A flowchart of a model training method based on the recurrent laryngeal nerve is provided. The method includes the following steps S101-S107.
[0065] Step S101: Obtain the sample endoscopic image and the sample marker that identifies the recurrent laryngeal nerve region in the sample endoscopic image.
[0066] Endoscopic images are images captured through an endoscope. Some endoscopic images contain the recurrent laryngeal nerve region; these are called positive sample images. Other endoscopic images may not contain the recurrent laryngeal nerve region; these are called negative sample images. When training the recurrent laryngeal nerve recognition model, both positive and negative sample images are used. This allows the model to learn the characteristics of images containing the recurrent laryngeal nerve region as well as those not containing it, resulting in a more robust recurrent laryngeal nerve recognition model.
[0067] Other methods for obtaining endoscopic images of samples can be found in subsequent articles. Figure 4 The illustrated embodiment will not be described in detail here.
[0068] The image format of the endoscopic sample can be RGB format, or other image formats.
[0069] The aforementioned sample markings can be obtained by experienced physicians from identifying the endoscopic images of the samples.
[0070] Specifically, sample labels can be represented in different forms.
[0071] In one case, the sample marker can be a binary image with the same size as the endoscopic image of the sample, containing both the recurrent laryngeal nerve region and the non-recurrent laryngeal nerve region. See also Figure 3 The image shows a binary image corresponding to a sample label. It can be seen from the image that the black area is the recurrent laryngeal nerve region, while the white area is the non-recurrent laryngeal nerve region.
[0072] In another scenario, the sample markers could be contour information of the recurrent laryngeal nerve, such as the coordinates of pixels on the contour.
[0073] Step S102: Input the sample endoscope image into the residual network of the recurrent laryngeal nerve recognition model to be trained to obtain the image features of the sample endoscope image.
[0074] In other words, the input of the residual network is the sample endoscopic image, and the output is the image features.
[0075] Residual networks can contain multiple network layers that can be cascaded to extract features from endoscopic images. Each network layer can refer to the features extracted by the preceding network layers, for example, by superimposing the residuals between the features extracted by the preceding networks onto the features extracted by the current network layer. This makes the extracted features more accurate and can extract deep image features from the endoscopic images.
[0076] In one implementation, the residual network can be ResNet50 (a 50-layer residual network).
[0077] Step S103: Input the image features into the pooling layer for global average pooling to obtain multi-scale pooling results.
[0078] In other words, the input to the pooling layer is image features, and the output is the multi-scale pooling result.
[0079] The aforementioned pooling layer can transform multidimensional image features into feature vectors through global average pooling operations, thereby obtaining multi-scale pooling results. These multi-scale pooling results can reflect the global contextual information of the endoscopic image at multiple scales, which is beneficial for subsequent segmentation of the recurrent laryngeal nerve region.
[0080] The number of scales for the above pooling results can be set according to the actual needs of the scenario. For example, the number of scales can be set to 3, 4, 5, etc.
[0081] Step S104: Input the pooling results of each scale into the first convolutional layer, and input the output of the first convolutional layer into the first upsampling layer to obtain the first feature with the same size as the image feature.
[0082] In other words, the input to the first convolutional layer is the pooling results at each scale, and the output is the first convolution result obtained by convolving the pooling results at each scale. That is, each scale pooling result corresponds to a first convolution result. The input to the first upsampling layer is each first convolution result, and the output is the first feature corresponding to each first convolution result.
[0083] In one implementation, such as Figure 1 As shown, the recurrent laryngeal nerve recognition model can contain multiple first convolutional layers, each of which performs convolution processing on the pooling results at one scale.
[0084] For example, each first convolutional layer can use a 1x1 convolutional kernel, so that each first convolutional layer performs a 1x1 convolution transformation on the pooling results at each scale. The 1x1 convolution transformation can integrate the pooling results at different scales. By controlling the number of convolutional kernels, the dimensionality of the pooling results at different scales can be reduced, so that the pooling results at each scale have the same dimension.
[0085] Similar to the first convolutional layer described above, the recurrent laryngeal nerve recognition model can also contain multiple first upsampling layers. The number of first upsampling layers is equal to the number of first convolutional layers, with each first upsampling layer corresponding to a first convolutional layer. The upsampling process is applied to the first convolutional result output by the corresponding first convolutional layer. For example... Figure 1 As shown, three first upsampling layers are illustrated.
[0086] Each first upsampling layer outputs the same first feature size. However, since the first convolution result outputs by different first convolution layers have different sizes, the upsampling ratio of each first upsampling layer is different.
[0087] Step S105: Input each first feature and image feature into the connection layer and connect them to obtain the second feature.
[0088] In other words, the input to the connection layer is each of the first features and image features, and the output is the second feature.
[0089] Specifically, the connection order between each first feature and image feature can be preset, and then each feature can be connected end to end in the above connection order to obtain the second feature.
[0090] The connection order can be based on the scale of the pooling results corresponding to the first feature from small to large, and the image features can be connected to the tail of the last first feature.
[0091] The connection order described above can also be such that the image features are placed at the beginning, and then each first feature is connected to the end of the image features in descending order of the scale of the pooling result corresponding to the first feature.
[0092] Step S106: Input the second feature into the second convolutional layer, and input the output of the second convolutional layer into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the sample endoscopic image.
[0093] In other words, the input to the second convolutional layer is the second feature, and the output is the result of the second convolution. The input to the second upsampling layer is the result of the second convolution, and the output is the segmentation result of the recurrent laryngeal nerve region.
[0094] In one implementation, when the sample is labeled as a binary image and its size is the same as that of the endoscopic image of the sample, the output information of the second upsampling layer can be: a binary image with the same size as the endoscopic image of the sample, which marks the segmented recurrent laryngeal nerve region. In this case, the binary image can be similar to the one described above. Figure 3 The binary image shown ensures that the sample labels and the information output from the second upsampling layer are of the same size and type. This allows for effective data alignment when training the recurrent laryngeal nerve recognition model in subsequent steps, facilitating the calculation of the model's loss value.
[0095] Of course, the output information of the second upsampling layer can also be the regional information of the segmented recurrent laryngeal nerve region, such as the regional contour and regional location information of the recurrent laryngeal nerve region in the sample endoscopic image.
[0096] From step S102 to this step, each network layer in the recurrent laryngeal nerve recognition model sequentially processes the sample endoscope image in different ways, extracting features from the sample endoscope image at different scales and angles. The extracted features can reflect the features of the recurrent laryngeal nerve in the image from different scales and angles. Therefore, after feature extraction is performed again by the second convolutional layer in this step, the recurrent laryngeal nerve region can be segmented from the sample endoscope image.
[0097] Step S107: Based on the obtained recurrent laryngeal nerve region and sample labels, train the recurrent laryngeal nerve recognition model.
[0098] After obtaining the recurrent laryngeal nerve region, the difference between the recurrent laryngeal nerve region and the sample label can be obtained. This difference can reflect the difference between the recurrent laryngeal neural network model and the actual situation when performing recurrent laryngeal nerve segmentation. Therefore, the model parameters of the recurrent laryngeal nerve recognition model can be adjusted based on the above difference and the preset loss function so that the output result of the recurrent laryngeal nerve recognition model is closer to the above sample label.
[0099] Specifically, the loss function mentioned above can be the cross-entropy loss function.
[0100] In addition, the specific training method for the recurrent laryngeal nerve recognition model can be found in the subsequent embodiments, which will not be detailed here.
[0101] Those skilled in the art will understand that when training the recurrent laryngeal nerve recognition model, multiple images can be input into the model for each batch of training. The number of images output in each batch can be eight.
[0102] As can be seen from the above, in the solution provided by the embodiments of the present invention, the recurrent laryngeal nerve recognition model to be trained includes a residual network, a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer. These network layers process the sample endoscopic image sequentially, thereby segmenting the recurrent laryngeal nerve region from the sample endoscopic image. When training the recurrent laryngeal nerve recognition model, not only the aforementioned recurrent laryngeal nerve region is referenced, but also the sample markers. Since the sample markers identify the actual location of the recurrent laryngeal nerve in the sample endoscopic image, the differences between the recurrent laryngeal nerve region and the sample markers can be determined to identify the differences in the recurrent laryngeal nerve recognition model when segmenting the recurrent laryngeal nerve region. Therefore, the model parameters of the recurrent laryngeal nerve recognition model can be adjusted based on these differences, enabling the recurrent laryngeal nerve recognition model to learn the characteristics of the actual recurrent laryngeal nerve region, thereby improving the accuracy of segmenting the recurrent laryngeal nerve region. In other words, a model capable of segmenting the recurrent laryngeal nerve region based on an image is trained.
[0103] In one embodiment of the present invention, during the feature extraction process of the aforementioned residual network on the sample endoscope image, the sample endoscope image can also be downsampled, and the downsampling factor is denoted as the first factor. Subsequently, a second upsampling layer performs upsampling, and the upsampling factor is denoted as the second factor, with the first factor and the second factor being equal.
[0104] For example, if the first multiplier is 8, then the second upsampling layer needs to go through 3 upsampling steps, each with an upsampling multiplier of 2, and the final upsampling multiplier is 8, which means the second multiplier is 8.
[0105] The model training steps mentioned in step S107 above will be explained below.
[0106] In one embodiment of the present invention, the recurrent laryngeal nerve recognition model can be trained by following steps one and two.
[0107] Step 1: Based on the obtained recurrent laryngeal nerve region and sample labels, obtain the loss value of the recurrent laryngeal nerve recognition model.
[0108] In one implementation, the above loss value can be obtained using the cross-entropy loss function. The specific calculation formula is as follows:
[0109]
[0110] Where loss represents the aforementioned loss value, M represents the number of categories, and y c This indicates whether the obtained recurrent laryngeal nerve region is the same as the recurrent laryngeal nerve region identified by the sample marker. If they are the same, the value is 1; otherwise, the value is 0. c This represents the probability that the sample endoscopic image belongs to category c, that is, the probability that the sample endoscopic image contains the recurrent laryngeal nerve, or the probability that the sample endoscopic image does not contain the recurrent laryngeal nerve.
[0111] Since the regions in the endoscopic images of the samples provided in this embodiment of the invention can be divided into two categories, the recurrent laryngeal nerve region and the non-recurrent laryngeal nerve region, the value of M is 2.
[0112] Step 2: Based on the preset initial learning rate and loss value, the model parameters of the recurrent laryngeal nerve recognition model are adjusted according to the poly learning rate decay strategy and the gradient descent criterion set by the Adam optimizer to achieve model training.
[0113] The poly learning rate decay strategy is an exponential transformation strategy, and the specific formula is shown below:
[0114]
[0115] Where lr represents the new learning rate, base_lr is the initial learning rate, epoch is the number of iterations, max_epoch is the maximum number of iterations, and power is used to control the shape of the change curve, which is usually greater than 1.
[0116] Compared to other optimization methods, the Adam optimizer improves the gradient descent criterion through gradient moving average and bias correction.
[0117] The preset initial learning rate can be 0.01.
[0118] As can be seen from the above, the solution provided in this implementation uses a poly learning rate decay strategy and a gradient descent criterion based on the Adam optimizer during model training, which enables the trained model to converge quickly, thereby improving the model training efficiency.
[0119] The method for obtaining the sample endoscopic image mentioned in step S101 above will be explained below.
[0120] In one embodiment of the present invention, see Figure 4 A flowchart of a method for obtaining sample images is provided, which includes the following steps S401-S406.
[0121] Step S401: Obtain the raw endoscopic image.
[0122] The original endoscopic image can be an image acquired via endoscopy during a previous surgical procedure.
[0123] Step S402: Obtain the scaling ratio of the original endoscopic image.
[0124] To reduce the size of the recurrent laryngeal nerve recognition model, the size of the input image can be set, meaning the model only processes images of a specific size. Since images acquired by different endoscopes may have different sizes, the original endoscopic images can be scaled to obtain an image with the same size as the input image of the model.
[0125] In view of the above, the above ratio is determined based on the size of the original endoscopic image and the preset size of the input image of the recurrent laryngeal nerve recognition model.
[0126] Step S403: Scale the original endoscopic image according to the above ratio to obtain the first image.
[0127] Step S404: If the size of the first image is smaller than the preset size, perform pixel expansion processing on the first image according to the preset size to obtain the second image.
[0128] Since some scaling algorithms only support scaling images according to a specific ratio, resulting in the size of the first image being smaller than the preset size, in order to ensure that the size of the first image meets the size requirements of the recurrent laryngeal nerve recognition model for the input image, the first image is subjected to pixel expansion processing in this step to obtain a second image with the same size as the preset size.
[0129] For example, if the resolution of the original endoscope image is 1920x1080 and the preset size is 480×272, the above ratio can be 1 / 4. After scaling according to this ratio, the size of the first image is 480×270, which is smaller than the preset size. Then, the first image is pixel expanded in the height direction to obtain an image with a size of 480×272.
[0130] Step S405: Normalize the pixel values of each pixel in the second image to obtain the third image.
[0131] Step S406: Obtain the sample image based on the third image.
[0132] In one implementation, the mean and variance of the pixel values of each pixel in the third image can be obtained. The mean is subtracted from the pixel values of each pixel in the third image to obtain the difference, and the difference is divided by the variance to obtain the sample image. This can normalize the pixel values of each pixel in the third image to a certain range, making the pixel values smoother and reducing the occurrence of outliers.
[0133] As can be seen from the above, the solution provided in this embodiment performs size transformation and normalization on the original endoscopic image, which not only makes the size of the obtained sample endoscopic images uniform, but also reduces the influence of irregular lighting on the sample endoscopic images. This makes the sample endoscopic images used in the training of the recurrent laryngeal nerve recognition model of high quality and suitable for the model's requirements for input images, which is conducive to reducing the model size and improving the model training efficiency.
[0134] Corresponding to the above-mentioned model training method based on the recurrent laryngeal nerve, this embodiment of the invention also provides a method for identifying the recurrent laryngeal nerve.
[0135] See Figure 5 A flowchart of a method for identifying the recurrent laryngeal nerve is provided, which includes the following steps S501-S502.
[0136] Step S501: Obtain the image of the endoscope to be identified.
[0137] Step S502: Input the endoscope image to be identified into the pre-trained recurrent laryngeal nerve recognition model to identify the recurrent laryngeal nerve and obtain the recurrent laryngeal nerve region in the endoscope image to be identified.
[0138] The recurrent laryngeal nerve recognition model is a model trained according to the scheme provided in any of the above embodiments.
[0139] As can be seen from the previous embodiments, the output information of the recurrent laryngeal nerve recognition model can be a binary image. In this case, the binary image can be processed based on morphological operations to remove imbalances such as edge spikes in the recurrent laryngeal nerve region of the binary image; the image after the opening operation can also be processed by the closing operation to eliminate small fragmented areas.
[0140] The opening operation described above can be performed by first eroding the image and then dilating it. Similarly, the closing operation described above can be performed by first dilating the image and then eroding it.
[0141] In addition, when performing the above processing on the endoscopic images to be identified, a 5*5 kernel can be used.
[0142] In one embodiment of the present invention, after the recurrent laryngeal nerve recognition model segments the recurrent laryngeal nerve region, the recurrent laryngeal nerve region can be superimposed on the endoscopic image to be recognized, and the superimposed image can be displayed to the doctor, so that the doctor can intuitively view the recognized recurrent laryngeal nerve region. See [link to relevant documentation]. Figure 6 , Figure 6 A schematic diagram of a recurrent laryngeal nerve is shown.
[0143] As can be seen from the above, the recurrent laryngeal nerve recognition scheme provided in this embodiment of the invention is implemented using a pre-trained recurrent laryngeal nerve recognition model. Since the recurrent laryngeal nerve recognition model was trained using a large number of samples and learned the characteristics of the recurrent laryngeal nerve, applying the above model for recurrent laryngeal nerve recognition can accurately identify the recurrent laryngeal nerve region in the image.
[0144] Furthermore, once the recurrent laryngeal nerve is exposed during surgery, endoscopic images can be easily acquired using an endoscope. This allows for the identification of the recurrent laryngeal nerve using the aforementioned model, eliminating the need for the surgeon to operate the equipment during the procedure and minimizing interruptions. Additionally, once the recurrent laryngeal nerve is identified, it can be promptly alerted to the surgeon via sound or other means, further reducing surgical risks.
[0145] Corresponding to the above-mentioned model training method based on the recurrent laryngeal nerve, this embodiment of the invention also provides a model training device based on the recurrent laryngeal nerve.
[0146] In one embodiment of the present invention, see Figure 7 , Figure 7 This is a schematic diagram of a model training device based on the recurrent laryngeal nerve provided in an embodiment of the present invention. The device includes:
[0147] The sample acquisition module 701 is used to acquire endoscopic images of samples and to identify sample markers that identify the recurrent laryngeal nerve region in the endoscopic images of samples.
[0148] The feature acquisition module 702 is used to input the above-mentioned sample endoscope image into the residual network of the recurrent laryngeal nerve recognition model to be trained, and obtain the image features of the above-mentioned sample endoscope image. The above-mentioned recurrent laryngeal nerve recognition model further includes: pooling layer, first convolutional layer, first upsampling layer, connection layer, second convolutional layer and second upsampling layer.
[0149] The feature pooling module 703 is used to input the above image features into the above pooling layer for global average pooling to obtain multi-scale pooling results.
[0150] Pooling result processing module 704 is used to input the pooling results of each scale into the first convolutional layer, and input the output of the first convolutional layer into the first upsampling layer to obtain a first feature with the same size as the image feature.
[0151] The feature connection module 705 is used to input each first feature and the above-mentioned image features into the above-mentioned connection layer for connection to obtain the second feature;
[0152] The region segmentation module 706 is used to input the second feature into the second convolutional layer and input the output of the second convolutional layer into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the sample endoscopic image.
[0153] The model training module 707 is used to train the above-mentioned recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels.
[0154] As can be seen from the above, in the solution provided by the embodiments of the present invention, the recurrent laryngeal nerve recognition model to be trained includes a residual network, a pooling layer, a first convolutional layer, a first upsampling layer, a connection layer, a second convolutional layer, and a second upsampling layer. These network layers process the sample endoscopic image sequentially, thereby segmenting the recurrent laryngeal nerve region from the sample endoscopic image. When training the recurrent laryngeal nerve recognition model, not only the aforementioned recurrent laryngeal nerve region is referenced, but also the sample markers. Since the sample markers identify the actual location of the recurrent laryngeal nerve in the sample endoscopic image, the differences between the recurrent laryngeal nerve region and the sample markers can be determined to identify the differences in the recurrent laryngeal nerve recognition model when segmenting the recurrent laryngeal nerve region. Therefore, the model parameters of the recurrent laryngeal nerve recognition model can be adjusted based on these differences, enabling the recurrent laryngeal nerve recognition model to learn the characteristics of the actual recurrent laryngeal nerve region, thereby improving the accuracy of segmenting the recurrent laryngeal nerve region. In other words, a model capable of segmenting the recurrent laryngeal nerve region based on an image is trained.
[0155] In one embodiment of the present invention, the sample is labeled as a binary image with the same size as the aforementioned endoscopic image of the sample;
[0156] The output information of the second upsampling layer is a binary image with the same size as the sample endoscope image mentioned above.
[0157] As can be seen from the above, the solution provided in this embodiment can ensure that the sample labels and the information output by the second upsampling layer have the same size and type. In this way, when training the recurrent laryngeal nerve recognition model based on these two types of information in subsequent steps, data alignment can be effectively achieved, which facilitates the calculation of the model loss value.
[0158] In one embodiment of the present invention, the model training module 707 includes:
[0159] The loss value acquisition unit is used to obtain the loss value of the above recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels.
[0160] The parameter adjustment unit is used to adjust the model parameters of the above recurrent laryngeal nerve recognition model according to the preset initial learning rate and the above loss value, following the poly learning rate decay strategy and based on the gradient descent criterion set by the Adam optimizer, so as to achieve model training.
[0161] As can be seen from the above, the solution provided in this implementation uses a poly learning rate decay strategy and a gradient descent criterion based on the Adam optimizer during model training, which enables the trained model to converge quickly, thereby improving the model training efficiency.
[0162] In one embodiment of the present invention, the above-mentioned model training device based on the recurrent laryngeal nerve further includes:
[0163] The second image acquisition module is used to acquire raw endoscopic images.
[0164] The scaling module is used to obtain the scaling ratio of the original endoscopic image.
[0165] The image scaling module is used to scale the original endoscopic image according to the above ratio to obtain the first image.
[0166] The pixel expansion module is used to perform pixel expansion processing on the first image according to the preset size when the size of the first image is smaller than the preset size, so as to obtain the second image.
[0167] The image normalization module is used to normalize the pixel values of each pixel in the second image to obtain the third image.
[0168] The third image acquisition module is used to obtain the sample image based on the third image.
[0169] As can be seen from the above, the solution provided in this embodiment performs size transformation and normalization on the original endoscopic image, which not only makes the size of the obtained sample endoscopic images uniform, but also reduces the influence of irregular lighting on the sample endoscopic images. This makes the sample endoscopic images used in the training of the recurrent laryngeal nerve recognition model of high quality and suitable for the model's requirements for input images, which is conducive to reducing the model size and improving the model training efficiency.
[0170] In one embodiment of the present invention, the sample image acquisition module includes:
[0171] The mean and variance acquisition unit is used to obtain the mean and variance of the pixel values of each pixel in the third image mentioned above.
[0172] The mean-variance normalization unit is used to subtract the mean from the pixel value of each pixel in the third image to obtain the difference, and then divide the obtained difference by the variance to obtain the sample image.
[0173] As can be seen from the above, in the solution provided in this embodiment, the pixel values of each pixel in the third image can be normalized to a certain range, so that the pixel values can be smoothed and the occurrence of singular values can be reduced.
[0174] In one embodiment of the present invention, see Figure 8 , Figure 8 This is a schematic diagram of a recurrent laryngeal nerve recognition device provided in an embodiment of the present invention. The device includes:
[0175] The first image acquisition module 801 is used to acquire the image of the endoscope to be identified.
[0176] The region recognition module 802 is used to input the above-mentioned endoscope image to be recognized into a pre-trained recurrent laryngeal nerve recognition model to perform recurrent laryngeal nerve recognition, thereby obtaining the recurrent laryngeal nerve region in the above-mentioned endoscope image to be recognized. The recurrent laryngeal nerve recognition model is a model trained by the recurrent laryngeal nerve-based model training device according to the above embodiment.
[0177] As can be seen from the above, the recurrent laryngeal nerve recognition scheme provided in this embodiment of the invention is implemented using a pre-trained recurrent laryngeal nerve recognition model. Since the recurrent laryngeal nerve recognition model was trained using a large number of samples and learned the characteristics of the recurrent laryngeal nerve, applying the above model for recurrent laryngeal nerve recognition can accurately identify the recurrent laryngeal nerve region in the image.
[0178] Furthermore, once the recurrent laryngeal nerve is exposed during surgery, endoscopic images can be easily acquired using an endoscope. This allows for the identification of the recurrent laryngeal nerve using the aforementioned model, eliminating the need for the surgeon to operate the equipment during the procedure and minimizing interruptions. Additionally, once the recurrent laryngeal nerve is identified, it can be promptly alerted to the surgeon via sound or other means, further reducing surgical risks.
[0179] This invention also provides an electronic device, such as... Figure 9 As shown, it includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904. The processor 901, communication interface 902, and memory 903 communicate with each other via the communication bus 904.
[0180] Memory 903 is used to store computer programs;
[0181] When the processor 901 executes the program stored in the memory 903, it implements the model training method or the recurrent laryngeal nerve recognition method described in the above method embodiments.
[0182] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0183] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0184] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0185] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0186] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described recurrent laryngeal nerve-based model training methods or recurrent laryngeal nerve recognition methods.
[0187] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the recurrent laryngeal nerve-based model training or recurrent laryngeal nerve recognition methods described in the above embodiments.
[0188] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0189] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0190] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0191] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A model training method based on the recurrent laryngeal nerve, characterized in that, The method includes: Obtain sample endoscopic images and sample markers that identify the recurrent laryngeal nerve region in the sample endoscopic images, wherein the sample endoscopic images refer to images captured by an endoscope; The sample endoscope image is input into the residual network of the recurrent laryngeal nerve recognition model to be trained to obtain the image features of the sample endoscope image. The recurrent laryngeal nerve recognition model further includes: pooling layer, multiple first convolutional layers, multiple first upsampling layers, connection layer, second convolutional layer and second upsampling layer. Each first convolutional layer performs convolution processing on the pooling result of one scale. The number of first upsampling layers is the same as the number of first convolutional layers. One first upsampling layer corresponds to one first convolutional layer. The image features are input into the pooling layer for global average pooling to obtain multi-scale pooling results. The pooling results at each scale are input into each first convolutional layer, and the output results of each first convolutional layer are input into each first upsampling layer to obtain a first feature with the same size as the image feature. Each first feature and the image feature are input into the connection layer and connected to obtain a second feature. The connection order is as follows: the first features are connected in order of increasing scale of the pooling results corresponding to each first feature, and the image feature is connected to the end of the last first feature; or the image feature is placed at the beginning and each first feature is connected to the end of the image feature in order of decreasing scale of the pooling results corresponding to each first feature. The second feature is input into the second convolutional layer, and the output of the second convolutional layer is input into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the endoscopic image of the sample; The recurrent laryngeal nerve recognition model is trained based on the obtained recurrent laryngeal nerve region and sample labels.
2. The method according to claim 1, characterized in that, The sample is labeled as a binary image with the same size as the endoscopic image of the sample. The output information of the second upsampling layer is a binary image with the same size as the sample endoscope image.
3. The method according to claim 1, characterized in that, The training of the recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels includes: Based on the obtained recurrent laryngeal nerve region and sample labels, the loss value of the recurrent laryngeal nerve recognition model is obtained; Based on the preset initial learning rate and the loss value, the model parameters of the recurrent laryngeal nerve recognition model are adjusted according to the poly learning rate decay strategy and the gradient descent criterion set by the Adam optimizer, thereby achieving model training.
4. The method according to any one of claims 1-3, characterized in that, Before obtaining the sample image, the process also includes: Obtain raw endoscopic images; The scaling ratio of the original endoscopic image is obtained, and the scaling ratio is determined based on the size of the original endoscopic image and the preset size of the input image of the recurrent laryngeal nerve recognition model; The original endoscopic image is scaled according to the specified ratio to obtain a first image; If the size of the first image is smaller than the preset size, the first image is subjected to pixel expansion processing according to the preset size to obtain the second image; The pixel values of each pixel in the second image are normalized to obtain the third image; The sample image is obtained based on the third image.
5. The method according to claim 4, characterized in that, Obtaining the sample image based on the third image includes: Obtain the mean and variance of the pixel values of each pixel in the third image; The difference is obtained by subtracting the mean value from the pixel value of each pixel in the third image, and then dividing the difference by the variance to obtain the sample image.
6. A method for identifying the recurrent laryngeal nerve, characterized in that, The method includes: Obtain the image of the endoscope to be identified; The endoscope image to be identified is input into a pre-trained recurrent laryngeal nerve recognition model to identify the recurrent laryngeal nerve, thereby obtaining the recurrent laryngeal nerve region in the endoscope image to be identified. The recurrent laryngeal nerve recognition model is a model trained according to any one of claims 1-5.
7. A model training device based on the recurrent laryngeal nerve, characterized in that, The device includes: The sample acquisition module is used to acquire endoscopic images of samples and sample markers that identify the recurrent laryngeal nerve region in the endoscopic images of samples, wherein the endoscopic images of samples refer to images captured by an endoscope; The feature acquisition module is used to input the sample endoscope image into the residual network of the recurrent laryngeal nerve recognition model to be trained, and obtain the image features of the sample endoscope image. The recurrent laryngeal nerve recognition model further includes: pooling layer, multiple first convolutional layers, multiple first upsampling layers, connection layer, second convolutional layer and second upsampling layer, wherein each first convolutional layer performs convolution processing on the pooling result of one scale, the number of first upsampling layers is the same as the number of first convolutional layers, and one first upsampling layer corresponds to one first convolutional layer. The feature pooling module is used to input the image features into the pooling layer for global average pooling to obtain multi-scale pooling results. The pooling result processing module is used to input the pooling results of each scale into each first convolutional layer, and input the output results of each first convolutional layer into each first upsampling layer to obtain a first feature with the same size as the image feature; The feature connection module is used to input each first feature and the image feature into the connection layer for connection to obtain a second feature. The connection order is as follows: the first features are connected in order of increasing scale of the pooling results corresponding to each first feature, and the image feature is connected to the tail of the last first feature; or the image feature is placed at the beginning and the first features are connected to the tail of the image feature in order of decreasing scale of the pooling results corresponding to each first feature. The region segmentation module is used to input the second feature into the second convolutional layer and input the output of the second convolutional layer into the second upsampling layer to obtain the recurrent laryngeal nerve region segmented from the endoscopic image of the sample; The model training module is used to train the recurrent laryngeal nerve recognition model based on the obtained recurrent laryngeal nerve region and sample labels.
8. A recurrent laryngeal nerve recognition device, characterized in that, The device includes: The first image acquisition module is used to acquire the image of the endoscope to be identified; The region recognition module is used to input the endoscope image to be recognized into a pre-trained recurrent laryngeal nerve recognition model to perform recurrent laryngeal nerve recognition, thereby obtaining the recurrent laryngeal nerve region in the endoscope image to be recognized, wherein the recurrent laryngeal nerve recognition model is a model trained by the device according to claim 7.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method according to any one of claims 1-5 or 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-5 or 6.
Citation Information
Patent Citations
Image segmentation method, apparatus, and storage medium
CN109410185A