Endoscopic high-resolution image adaptive blocking method, system, device and medium
Through the adaptive chunking method, the MLP model and Adam optimizer are used to solve the shortcomings of fixed-size chunking method in endoscopic high-resolution image processing, and achieve more efficient and higher quality image chunking.
Patent Information
- Application Number
- CN202411691055.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The fixed-size blocking method is used in existing endoscopic high-resolution image processing tasks, which cannot take into account both processing efficiency and image quality, resulting in loss of details or excessive calculation burden.
An endoscopic high-resolution image adaptive chunking method is proposed. By acquiring the image data set, extracting the training set and performing feature extraction, an MLP model is constructed to predict the chunking width and height, and a model training is used by Adam optimizer to realize adaptive chunking.
The processing efficiency and image quality of the endoscope's high-resolution image are improved, and the blocking can be accurately and efficiently for different regions to achieve optimal efficiency and optimal quality image block segmentation.
Smart Images

Figure CN119169033B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical equipment and image processing technology, and in particular to an endoscope high-resolution image adaptive blocking method, system, device and medium. Background Art
[0002] During endoscopy, the use of high-resolution images can provide an extremely clear view, which helps in accurate diagnosis. In high-resolution image processing tasks (such as 4K or 8K images), it is often necessary to divide the image into smaller blocks to facilitate subsequent processing operations.
[0003] However, due to the large differences in detail complexity in different areas of the image, the traditional fixed-size block method often cannot balance processing efficiency and image quality. On the one hand, blocks that are too large may not be able to capture subtle lesions, resulting in loss of details; on the other hand, blocks that are too small will increase the computational burden and lead to low processing efficiency. This fixed block method cannot meet the needs of dynamic adjustment, which restricts the effective application of endoscopic images in real-time diagnosis. Summary of the invention
[0004] In view of this, the present invention provides a method, system, device and medium for adaptively segmenting endoscopic high-resolution images to solve the problem of not being able to balance processing efficiency and image quality when using fixed-size segmentation in existing endoscopic high-resolution image processing tasks.
[0005] The present invention provides an endoscope high-resolution image adaptive blocking method, the method comprising:
[0006] Acquire an endoscopic image data set, extract a training set from the endoscopic image data set, and perform feature extraction on the training set to obtain a feature set;
[0007] Constructing an initial MLP model before training, using the feature set as the input of the initial MLP model, using the block width and the block height as the output of the initial MLP model, and using the Adam optimizer as the parameter optimizer of the initial MLP model;
[0008] Defining a loss function of the initial MLP model, and training the initial MLP model using the feature set, the loss function, and the Adam optimizer to obtain a trained target MLP model;
[0009] An endoscopic image to be segmented is acquired, and segmentation prediction is performed on the endoscopic image to be segmented using the target MLP model to obtain a plurality of target image blocks.
[0010] Optionally, the training set includes a plurality of training images;
[0011] Perform feature extraction on the training set to obtain a feature set, including:
[0012] Select any one training image from the training set, and divide the selected training image into a plurality of image regions;
[0013] Select an image region from the selected training image, and calculate the horizontal gradient and the vertical gradient of each pixel of the image region selected from the selected training image;
[0014] According to the horizontal direction gradient and the vertical direction gradient of each pixel point of the image area selected in the selected training image, the texture energy of each pixel point of the image area selected in the selected training image is calculated respectively;
[0015] Calculate the standard deviation of the selected image area in the selected training image in the RGB color channel;
[0016] Constructing a feature vector of each pixel of the image region selected in the selected training image according to the texture energy of each pixel of the image region selected in the selected training image and the standard deviation of the image region selected in the selected training image in the RGB color channel;
[0017] Traversing each image region in the selected training image, and constructing the feature vector of each pixel point of each image region in the selected training image in the same way; and obtaining the feature vector set corresponding to the selected training image according to all the feature vectors of all the image regions in the selected training image;
[0018] Traversing each training image in the endoscopic image data set, and obtaining a feature vector set corresponding to each training image in a one-to-one manner according to the same method;
[0019] The feature set is obtained according to all feature vector sets in all training images.
[0020] Optionally, the selected training image is the ath training image in the endoscopic image data set, and the selected image region is the bth image region in the ath training image;
[0021] Then the horizontal gradient and vertical gradient of each pixel in the image area selected in the selected training image are calculated respectively, including:
[0022] The Sobel operator is used to calculate the horizontal gradient and vertical gradient of each pixel in the b-th image area in the a-th training image.
[0023] The specific formula for calculating the horizontal gradient and vertical gradient of the bth image area in the ath training image at the pixel point (x, y) is:
[0024] ;
[0025] in, and are the horizontal and vertical gradients of the b-th image region in the a-th training image at the pixel point (x, y), i and j are the horizontal and vertical indexes of the Sobel operator in the window (3, 3), respectively. and are the pixel grayscale values at pixel point (x, y) and pixel point (x+i, y+j) in the bth image area in the ath training image, respectively. and They are the horizontal convolution kernel and vertical convolution kernel of the b-th image area in the a-th training image at the pixel point (i, j) in the window (3, 3), respectively.
[0026] Optionally, the specific formula for calculating the texture energy of the b-th image region in the a-th training image at the pixel point (x, y) is:
[0027] ;
[0028] in, is the texture energy of the bth image region in the ath training image at the pixel point (x, y), N is the total number of pixels in the bth image region in the ath training image, N x and N y are the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the b-th image area in the a-th training image, respectively. and are the horizontal and vertical gradients of the bth image region in the ath training image at the pixel point (x+i, y+j), respectively.
[0029] Optionally, the specific formula for calculating the standard deviation of the b-th image region in the a-th training image in the RGB color channel is:
[0030] ;
[0031] Among them, the RGB color channels include R channel, G channel and B channel. , and are the R channel standard deviation, G channel standard deviation and B channel standard deviation of the b-th image region in the a-th training image, respectively. , and are the R channel color value, G channel color value, and B channel color value of the k-th pixel in the b-th image area in the a-th training image, respectively. , and are the average color value of the R channel, the average color value of the G channel, and the average color value of the B channel of the bth image area in the ath training image, respectively.
[0032] Optionally, the constructed feature vector of the b-th image region at the pixel point (x, y) in the a-th training image is:
[0033] ;
[0034] in, is the feature vector of the bth image region at the pixel point (x, y) in the ath training image.
[0035] Optionally, the constructed initial MLP model includes an input layer, n hidden layers and an output layer;
[0036] In the constructed initial MLP model, the feature set is the input of the input layer, the output of the input layer is the input of the first hidden layer, and the activation signal is transmitted between adjacent hidden layers through model parameters; the output of the last hidden layer is the input of the output layer, and the block width and block height are both outputs of the output layer;
[0037] The input layer includes a plurality of neurons; in the input layer, the feature set is the input of the first neuron in the input layer, and each of the remaining neurons receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, wherein the output of the last neuron in the input layer is the input of the first hidden layer; the number of neurons in the input layer is equal to the dimension of the feature vector of each pixel in each training image;
[0038] In each hidden layer, the hidden layer includes multiple neurons; the output of the last neuron in the input layer is the input of the first neuron in the first hidden layer, and each of the remaining neurons in each hidden layer receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, and the output of the last neuron in the last hidden layer is the input of the output layer; the number of neurons in each hidden layer is equal or unequal.
[0039] Optionally, the output of the first hidden layer of n hidden layers is:
[0040] ;
[0041] in, is the output of the first hidden layer, The input transmitted from the input layer to the first hidden layer is the feature vector in the feature set; and are the network weights and biases of the first hidden layer, is the weighted sum between the input and bias of the first hidden layer, represents the activation function;
[0042] The output of the mth hidden layer among n hidden layers is:
[0043] ;
[0044] in, is the output of the mth hidden layer, and are the network weights and biases of the mth hidden layer, is the weighted sum between the input and bias of the mth hidden layer, is the output of the m-1th hidden layer, m satisfies 2≤m≤n;
[0045] The output of the output layer is:
[0046] ;
[0047] in, is the output of the output layer, and are the block width and block height respectively; is the output of the nth hidden layer, and are the network weights and bias of the output layer, respectively.
[0048] Optionally, the loss function of the initial MLP model is specifically the mean square error of the initial MLP model; the model parameters of the initial MLP model specifically include network weights and biases;
[0049] Training the initial MLP model using the feature set, the loss function and the Adam optimizer includes:
[0050] The feature set is input into the initial MLP model, and the initial MLP model is trained with the minimization of the mean square error as a training objective. During the training process of the initial MLP model, the network weights and biases are iteratively optimized using an Adam optimizer.
[0051] Optionally, after training the initial MLP model using the feature set, the loss function and the Adam optimizer, the method further includes:
[0052] extracting a validation set from the endoscopic image dataset;
[0053] Using the validation set to evaluate the trained intermediate MLP model;
[0054] If the evaluation passes, the intermediate MLP model is determined as the trained target MLP model; if the evaluation fails, the intermediate MLP model is continued to be trained using the feature set, the loss function and the Adam optimizer until the intermediate MLP model obtained by continued training passes the evaluation, and the intermediate MLP model that passes the evaluation is determined as the trained target MLP model.
[0055] Optionally, the target MLP model is used to perform block prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks, including:
[0056] Dividing the endoscope image to be divided into multiple initial image blocks, each initial image block has an initial width and an initial height;
[0057] Perform feature extraction on each initial image block respectively to obtain a feature vector of each pixel point of each initial image block;
[0058] Input all feature vectors of each initial image block into the target MLP model to predict the predicted width and predicted height corresponding to each initial image block;
[0059] The predicted width and the predicted height corresponding to each initial image block are used to adjust the initial width and the initial height of each initial image block respectively, so as to obtain a target image block corresponding to each initial image block.
[0060] Optionally, after obtaining the multiple target image blocks, the method further includes:
[0061] According to the preset image processing task, each of the target image blocks is subjected to image processing to obtain a processed image block corresponding to each of the target image blocks;
[0062] All processed image blocks are synthesized to obtain a target endoscopic image corresponding to the endoscopic image to be segmented.
[0063] In addition, the present invention also provides an endoscope high-resolution image adaptive blocking system, which is applied to the above-mentioned endoscope high-resolution image adaptive blocking method, and the system comprises:
[0064] An image acquisition module, used for acquiring an endoscopic image data set;
[0065] A feature extraction module, used to extract a training set from the endoscopic image data set, and perform feature extraction on the training set to obtain a feature set;
[0066] A model building module, used to build an initial MLP model before training, using the feature set as the input of the initial MLP model, using the block width and block height as the output of the initial MLP model, and using an Adam optimizer as a parameter optimizer of the initial MLP model;
[0067] A model training module, used to define a loss function of the initial MLP model, and train the initial MLP model using the feature set, the loss function and the Adam optimizer to obtain a trained target MLP model;
[0068] The adaptive block module is used to obtain the endoscopic image to be segmented, and use the target MLP model to perform block prediction on the endoscopic image to be segmented to obtain multiple target image blocks.
[0069] In addition, the present invention also provides an endoscope high-resolution image adaptive blocking device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method steps in the aforementioned endoscope high-resolution image adaptive blocking method when executed.
[0070] In addition, the present invention also provides a computer storage medium, which includes: at least one instruction, which implements the method steps in the aforementioned endoscopic high-resolution image adaptive blocking method when the instruction is executed by a computer.
[0071] The beneficial effects of the present invention are as follows: by extracting the training set from the acquired endoscopic image data set and performing feature extraction, the key features directly related to the high-resolution endoscopic image can be extracted, which is convenient for subsequent model training, provides a data basis for model training, and can also simplify the feature space, reduce the computational complexity, and improve the efficiency and accuracy of subsequent model training; the MLP model is also called the multi-layer perceptron model, and the framework of the MLP model is constructed, with the feature set as its input, the block width and block height as its output, and the Adam optimizer as its parameter optimizer. Through the above-mentioned model framework, a large number of nonlinear details and features in the high-resolution endoscopic image can be more effectively captured, and it can adapt to complex, High-dimensional data processing environment, at the same time, the model framework is also highly flexible and scalable, and can flexibly adapt to different endoscopic high-resolution image processing environments; when trained using the constructed model framework, the target MLP model obtained can efficiently and accurately process and adapt to the block task of endoscopic high-resolution images, and the processing quality and speed of the block task are effectively improved; using the target MLP model to predict the block of the endoscopic image to be segmented, the block width and block height of each area can be accurately and efficiently predicted for different areas, realizing the adaptive segmentation of the image, and obtaining multiple target image blocks of optimal quality with optimal efficiency, which is convenient for subsequent image operations;
[0072] The method, system, device and storage medium for adaptive segmentation of endoscopic high-resolution images of the present invention are based on the MLP model framework and Adam optimizer, and can efficiently and accurately adaptively segment different areas of endoscopic high-resolution images, thereby obtaining multiple image blocks of optimal quality for endoscopic high-resolution images with optimal efficiency, and providing strong support for endoscopic medical image analysis and diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:
[0074] Figure 1 A flow chart of an endoscope high-resolution image adaptive blocking method in Embodiment 1 of the present invention is shown;
[0075] Figure 2 A flow chart showing feature extraction of a training set in Embodiment 1 of the present invention is shown;
[0076] Figure 3 A flowchart of training an initial MLP model and evaluating the model using a feature set, a loss function, and an Adam optimizer in Embodiment 1 of the present invention is shown;
[0077] Figure 4A flowchart of using the target MLP model to perform block prediction in Embodiment 1 of the present invention is shown;
[0078] Figure 5 A flowchart of another method for adaptively segmenting endoscope high-resolution images in Embodiment 1 of the present invention is shown;
[0079] Figure 6 The structure diagram of an endoscope high-resolution image adaptive blocking system in the second embodiment of the present invention is shown. DETAILED DESCRIPTION
[0080] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0081] Embodiment 1
[0082] This embodiment provides an endoscope high-resolution image adaptive block method, such as Figure 1 As shown, the method includes:
[0083] S1: acquiring an endoscopic image data set, extracting a training set from the endoscopic image data set, and performing feature extraction on the training set to obtain a feature set;
[0084] S2: constructing an initial MLP model before training, taking the feature set as the input of the initial MLP model, taking the block width and the block height as the output of the initial MLP model, and taking the Adam optimizer as the parameter optimizer of the initial MLP model;
[0085] S3: defining a loss function of the initial MLP model, and training the initial MLP model using the feature set, the loss function and the Adam optimizer to obtain a trained target MLP model;
[0086] S4: Acquire the endoscopic image to be segmented, and use the target MLP model to perform segmentation prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks.
[0087] In this embodiment, by extracting the training set from the acquired endoscopic image data set and performing feature extraction, the key features directly related to the endoscopic high-resolution image can be extracted, which is convenient for subsequent model training, provides a data basis for model training, and can also simplify the feature space and reduce the computational complexity, thereby improving the efficiency and accuracy of subsequent model training; the MLP model is also called the multi-layer perceptron model, and the framework of the MLP model is constructed, with the feature set as its input, the block width and the block height as its output, and the Adam optimizer as its parameter optimizer. Through the above-mentioned model framework, a large number of nonlinear details and features in the endoscopic high-resolution image can be more effectively captured, and it can adapt to complex and high-resolution images. dimensional data processing environment, and the model framework is also highly flexible and scalable, and can flexibly adapt to different endoscopic high-resolution image processing environments; when trained using the constructed model framework, the target MLP model obtained can efficiently and accurately process and adapt to the segmentation task of endoscopic high-resolution images, and the processing quality and speed of the segmentation task are effectively improved; using the target MLP model to perform segmentation prediction on the segmented endoscopic image, the segmentation width and segmentation height of each area can be accurately and efficiently predicted for different areas, thereby realizing adaptive segmentation of the image and obtaining multiple target image blocks of optimal quality with optimal efficiency, which is convenient for subsequent image operations.
[0088] The adaptive segmentation method for high-resolution endoscopic images of this embodiment is based on the MLP model framework and the Adam optimizer, and can efficiently and accurately adaptively segment different areas of the high-resolution endoscopic images, thereby obtaining multiple image blocks of optimal quality for the high-resolution endoscopic images with optimal efficiency, and providing strong support for endoscopic medical image analysis and diagnosis.
[0089] Each step of the method for adaptively segmenting high-resolution endoscope images in this embodiment is described in detail below.
[0090] Specifically, in this embodiment S1, the endoscopic image data set can be obtained through a large number of working images collected by the endoscope in the historical period, and these working images can be in working scenes such as hospitals, clinics or research institutions; the endoscopic image data set can also be obtained in a big data network through search methods such as web crawlers, which is not limited in this embodiment. These images need to be collected from different parts, different disease types, and different lighting conditions, so that the entire data set can be diverse, so as to improve the generalization ability of subsequent models.
[0091] Specifically, in this embodiment S1, after the endoscopic image data set is acquired, the images in the data set are subjected to standardization processing (including but not limited to scaling, cropping, denoising and normalization) and labeling processing (referring to labeling the categories of the images), and then the processed endoscopic image data set is divided into a training set, a validation set and a test set according to a preset division ratio.
[0092] Preferably, in this embodiment S1, the training set includes a plurality of training images;
[0093] like Figure 2 As shown, feature extraction is performed on the training set to obtain a feature set, including:
[0094] S11: selecting any one training image from the training set, and dividing the selected training image into a plurality of image regions;
[0095] S12: Select an image region from the selected training image, and calculate the horizontal gradient and the vertical gradient of each pixel of the image region selected from the selected training image;
[0096] S13: Calculate the texture energy of each pixel of the image region selected in the selected training image according to the horizontal gradient and the vertical gradient of each pixel of the image region selected in the selected training image;
[0097] S14: Calculate the standard deviation of the image area selected in the selected training image in the RGB color channel;
[0098] S15: constructing a feature vector of each pixel of the image region selected in the selected training image according to the texture energy of each pixel of the image region selected in the selected training image and the standard deviation of the image region selected in the selected training image in the RGB color channel;
[0099] S16: traverse each image region in the selected training image, and construct a feature vector of each pixel point of each image region in the selected training image according to the same method; and obtain a feature vector set corresponding to the selected training image according to all feature vectors of all image regions in the selected training image;
[0100] S17: traverse each training image in the endoscope image data set, and obtain a feature vector set corresponding to each training image in a one-to-one manner according to the same method;
[0101] S18: Obtain the feature set according to all feature vector sets in all training images.
[0102] In the process of feature extraction, for each training image in the training set, it is first divided into multiple regions, and feature extraction is performed in each region. This can capture the local features of different regions of the endoscope high-resolution image, facilitate the grasp of more comprehensive image features, and help improve the accuracy of subsequent models.
[0103] For the extraction of local features in different regions, the gradient method (i.e., calculating the horizontal gradient value and the vertical gradient value respectively) is used on the first aspect to accurately capture the edge and texture changes in the image, with strong robustness and high efficiency, which facilitates the subsequent more accurate and reliable calculation of texture energy based on gradient values; and constructing feature vectors based on texture energy can more effectively describe the texture characteristics of the image, and provide rich information for subsequent image processing and analysis, which further helps to improve the accuracy of subsequent models and ensure that adaptive blocking can be performed for different regions in the end.
[0104] Regarding the extraction of local features in different areas, on the second hand, since the color distribution of different tissues or lesion areas in the endoscopic image may be different, by calculating the standard deviation of different areas in the RGB color channel, the discrete degree of pixel values in the local area in each color channel can be accurately reflected. By capturing these discrete degrees and using them to construct feature vectors, it can also provide rich information for subsequent image processing and analysis, which helps to improve the accuracy of subsequent models and ensure that adaptive blocking can be performed for different areas in the end.
[0105] For the convenience of explanation, in this embodiment, it is assumed that the selected training image is the ath training image in the endoscopic image data set, and the selected image region is the bth image region in the ath training image.
[0106] Specifically, in S12, the horizontal gradient and the vertical gradient of each pixel of the image area selected in the selected training image are calculated respectively, including:
[0107] The Sobel operator is used to calculate the horizontal gradient and vertical gradient of each pixel in the b-th image area in the a-th training image.
[0108] The specific formula for calculating the horizontal gradient and vertical gradient of the bth image area in the ath training image at the pixel point (x, y) is:
[0109] ;
[0110] in, and are the horizontal and vertical gradients of the b-th image region in the a-th training image at the pixel point (x, y), i and j are the horizontal and vertical indexes of the Sobel operator in the window (3, 3), respectively. and are the pixel grayscale values at pixel point (x, y) and pixel point (x+i, y+j) in the bth image area in the ath training image, respectively. and They are the horizontal convolution kernel and vertical convolution kernel of the b-th image area in the a-th training image at the pixel point (i, j) in the window (3, 3), respectively.
[0111] The Sobel operator is a gradient-based edge detection operator. It uses the grayscale value difference of adjacent pixels above, below, left and right of a pixel in an image to calculate the gradient, thereby detecting the edge in the image. The horizontal gradient and vertical gradient at each pixel in the local area of the training image are calculated by the above method, which can accurately capture the rich edge features (such as tissue boundaries, lesion areas, etc.) in the endoscopic image. The Sobel operator uses a convolution kernel with a window size of 3×3, which can smooth the noise to a certain extent, reduce the impact of noise on edge detection, and has strong anti-noise ability. In addition, the operator only needs to perform a convolution kernel operation on the image to calculate the gradient value, which has low calculation difficulty and high calculation efficiency.
[0112] The horizontal gradient and vertical gradient of each pixel in other image areas in the training image are calculated according to the above method, which will not be repeated here.
[0113] Specifically, in S13, the specific formula for calculating the texture energy of the b-th image region in the a-th training image at the pixel point (x, y) is:
[0114] ;
[0115] in, is the texture energy of the bth image region in the ath training image at the pixel point (x, y), N is the total number of pixels in the bth image region in the ath training image, N x and N y are the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the b-th image area in the a-th training image, respectively. and are the horizontal and vertical gradients of the bth image region in the ath training image at the pixel point (x+i, y+j), respectively.
[0116] The above-mentioned calculation method based on horizontal gradient and vertical gradient is used to calculate texture energy, which can accurately reflect the texture information of the image and enhance the sensitivity to texture changes. The calculation is simple and efficient, which is convenient for subsequent combination with other features to improve the recognition performance of the model and has wide applicability.
[0117] Similarly, the texture energy of each pixel in other image areas in the training image is calculated according to the above method, which will not be repeated here.
[0118] Specifically, in S14, the specific formula for calculating the standard deviation of the b-th image region in the a-th training image in the RGB color channel is:
[0119] ;
[0120] Among them, the RGB color channels include R channel, G channel and B channel. , and are the R channel standard deviation, G channel standard deviation and B channel standard deviation of the b-th image region in the a-th training image, respectively. , and are the R channel color value, G channel color value, and B channel color value of the k-th pixel in the b-th image area in the a-th training image, respectively. , and are the average color value of the R channel, the average color value of the G channel, and the average color value of the B channel of the bth image area in the ath training image, respectively.
[0121] In the above process of calculating the standard deviation of different color channels, the average color value of all pixels in the area in different color channels is first calculated, and then the standard deviation of the corresponding color channel can be obtained based on the difference between the color values of different color channels and the average color value of the corresponding color channel. The calculation principle is simple, the calculation difficulty is small, and the degree of color change in the image area can be accurately measured.
[0122] Similarly, the standard deviations of other image regions in the training image in the three color channels are calculated according to the above method, which will not be repeated here.
[0123] Specifically, in S15, the feature vector of the b-th image region in the a-th training image at the pixel point (x, y) is constructed as:
[0124] ;
[0125] in, is the feature vector of the bth image region at the pixel point (x, y) in the ath training image.
[0126] Combining texture energy and color channel standard deviation can fully capture the structure and color information of the image, provide rich features for image adaptive segmentation, and thus ensure the accuracy of the adaptive segmentation task of the model obtained by subsequent training; the feature vector constructed based on texture energy and color channel standard deviation can show strong robustness to interference factors such as lighting changes, image rotation and translation; and this feature vector does not require complex image processing algorithms, does not consume a lot of computing resources, facilitates real-time image analysis, and is computationally efficient.
[0127] Similarly, the feature vectors of each image region in the training image at other pixel points are constructed according to the above method, which will not be repeated here. In S16, all feature vectors in all image regions in the a-th training image constitute the feature vector set corresponding to the training image; in S18, the feature vector set of all training images constitutes the feature set of the training set.
[0128] Preferably, in this embodiment S2, the constructed initial MLP model includes an input layer, n hidden layers and an output layer;
[0129] In the constructed initial MLP model, the feature set is the input of the input layer, the output of the input layer is the input of the first hidden layer, and the activation signal is transmitted between adjacent hidden layers through model parameters; the output of the last hidden layer is the input of the output layer, and the block width and block height are both outputs of the output layer;
[0130] The input layer includes a plurality of neurons; in the input layer, the feature set is the input of the first neuron in the input layer, and each of the remaining neurons receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, wherein the output of the last neuron in the input layer is the input of the first hidden layer; the number of neurons in the input layer is equal to the dimension of the feature vector of each pixel in each training image;
[0131] In each hidden layer, the hidden layer includes multiple neurons; the output of the last neuron in the input layer is the input of the first neuron in the first hidden layer, and each of the remaining neurons in each hidden layer receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, and the output of the last neuron in the last hidden layer is the input of the output layer; the number of neurons in each hidden layer is equal or unequal.
[0132] The MLP (Multilayer Perceptron) model, also known as the multilayer perceptron model, is a basic and widely used neural network model in deep learning. In the above initial MLP model, an input layer is used as the first layer of the model framework, which is responsible for receiving external input data (that is, responsible for receiving the feature set); n hidden layers are located between the input layer and the output layer, responsible for extracting the potential features of the input data and performing nonlinear transformations. The neurons in the hidden layer receive the output from the previous layer of neurons as input through weighted connections, and generate their own output as the input of the next layer of neurons, and output the input of the last layer of neurons to the output layer; the output layer, as the last layer of the model framework, is responsible for generating the final prediction results, that is, predicting the block width and block height.
[0133] Among them, the number of neurons in the input layer is the same as the dimension of the feature vector of each pixel (specifically 4 dimensions, namely 4 dimensions corresponding to texture energy, R channel standard deviation, G channel standard deviation and B channel standard deviation respectively), which can ensure that each input feature has a corresponding neuron to receive and process it, avoiding information redundancy caused by too many neurons or information loss caused by too few neurons; it can also facilitate the neural network in the initial MLP model to process input data more efficiently and calculate the loss function of the model more accurately, thereby speeding up the training speed of subsequent models and improving the training effect.
[0134] In the input layer, the feature vector of each pixel is They are input into the input layer in turn, and then transmitted to the first hidden layer; through the processing of each neuron in the first hidden layer, the output of the first hidden layer is:
[0135] ;
[0136] in, is the output of the first hidden layer, The input transmitted from the input layer to the first hidden layer is the feature vector in the feature set; and are the network weights and biases of the first hidden layer, is the weighted sum between the input and bias of the first hidden layer, represents the activation function;
[0137] The output of the first hidden layer is input into the second hidden layer. After being processed by each neuron in the second hidden layer, the output of the second hidden layer is:
[0138] ;
[0139] in, is the output of the second hidden layer, and are the network weights and biases of the second hidden layer, is the weighted sum between the input and bias of the second hidden layer;
[0140] The output of the second hidden layer is input into the third hidden layer. After being processed by each neuron in the third hidden layer, the output of the third hidden layer is:
[0141] ;
[0142] in, is the output of the third hidden layer, and are the network weights and biases of the third hidden layer, is the weighted sum between the input and bias of the third hidden layer;
[0143] By analogy, the output of the mth hidden layer is:
[0144] ;
[0145] in, is the output of the mth hidden layer, and are the network weights and biases of the mth hidden layer, is the weighted sum between the input and bias of the mth hidden layer, is the output of the m-1th hidden layer, m satisfies 2≤m≤n;
[0146] Until it is transmitted to the nth hidden layer, the output of the nth hidden layer (that is, the last hidden layer) is:
[0147] ;
[0148] in, is the output of the nth hidden layer, and are the network weights and biases of the nth hidden layer, is the weighted sum between the input and bias of the nth hidden layer;
[0149] The output of the last hidden layer is input into the output layer to generate the prediction result (i.e. the predicted block width and block height). The output of the output layer is:
[0150] ;
[0151] in, is the output of the output layer, and are the block width and block height respectively; is the output of the nth hidden layer, and are the network weights and bias of the output layer, respectively.
[0152] The output of the above output layer is and That is, they are the predicted block width and block height respectively.
[0153] In the architecture of the initial MLP model, the loss function of the initial MLP model is defined as the mean square error of the initial MLP model (i.e., the model takes minimizing the mean square error as the training objective), and its parameter optimizer is defined as the Adam optimizer; wherein, the model parameters of the initial MLP model specifically include network weights and biases, i.e., the network weights and biases of the aforementioned hidden layers and output layers (including the aforementioned , ,…… , as well as , ,…… , wait).
[0154] Among them, the mean square error of the initial MLP model is:
[0155] ;
[0156] Among them, MSE is the mean square error, K is the number of samples input into the initial MLP model (i.e., the total number of image regions), and are the true block width and true block height of the kth sample input into the initial MLP model, and are the block width and block height of the kth sample predicted by the initial MLP model.
[0157] Preferably, if Figure 3 As shown, in this embodiment S3, the initial MLP model is trained using the feature set, the loss function and the Adam optimizer, including:
[0158] S31: Input the feature set into the initial MLP model, take minimizing the mean square error as the training goal, train the initial MLP model, and during the training process of the initial MLP model, use the Adam optimizer to iteratively optimize the network weights and biases.
[0159] Since the mean square error (MSE) calculates the average of the squares of the differences between the predicted values and the true values, it can intuitively reflect the average difference between the predicted values and the actual values. This is very useful for evaluating the accuracy of the endoscopic image segmentation task, because by using it as the loss function of the model, we can clearly understand the prediction performance of the model in each segment. In the model's adaptive segmentation of endoscopic images, the good mathematical properties of MSE also help to efficiently adjust the weights to minimize the loss function. In addition, since MSE amplifies the impact of large errors by squaring the errors, it is more sensitive to errors. This makes the model pay more attention to reducing predictions that are far from the true values during training, which helps to improve the prediction accuracy of the model. For the endoscopic image segmentation task, it can enable the model to more accurately capture the detailed information in the image, thereby improving the accuracy of adaptive segmentation.
[0160] The Adam optimizer is an efficient adaptive learning rate optimization algorithm that combines the ideas of momentum and adaptive learning rate. It adjusts the learning rate by calculating the first-order moment estimate and the second-order moment estimate of the gradient and performing exponentially weighted moving average. This adaptive adjustment strategy enables the Adam optimizer to adjust the step size more flexibly during training, thereby achieving faster and more stable convergence on complex or noisy data sets.
[0161] The specific process of optimizing the model parameters by the Adam optimizer in this embodiment is as follows:
[0162] 1. Initialization parameters
[0163] (1) Initialize all model parameters, including network weights and biases, and represent them uniformly as θ. The model parameters after initialization are θ 0 ;
[0164] (2) Initialize the first-order moment estimate m 0 = 0 and the second moment estimate v 0 =0;
[0165] (3) Set the learning rate α, for example, to 0.001;
[0166] (4) Set the exponential decay rate β of the first-order moment estimate 1 = 0.9 and the exponential decay rate β of the second-order moment estimate 2 =0.999;
[0167] (5) Set a small constant =10 -8 , used to prevent division by zero.
[0168] 2. Update process
[0169] (1) In each training step (time step is set to t), the gradient of the loss function with respect to the model parameters is first calculated:
[0170] ;
[0171] Among them, g t is the gradient of the loss function with respect to the model parameters in the tth round, θ t is the model parameter of the tth round, is the loss function, is the gradient function;
[0172] (2) Update the first-order moment estimate and the second-order moment estimate respectively:
[0173] ;
[0174] Among them, m t and v t are the first-order moment estimate and the second-order moment estimate of the tth round, m t-1 and v t-1 They are the first-order moment estimate and the second-order moment estimate for round t-1 respectively;
[0175] (3) Deviation correction:
[0176] Because m t and v t In the initial stage, it will be biased towards 0, so it needs to be bias-corrected separately;
[0177] ;
[0178] in, and are the first-order moment estimate and second-order moment estimate of the tth round after bias correction, respectively;
[0179] (4) Parameter update:
[0180] The first-order moment estimate and second-order moment estimate of the tth round after bias correction update the model parameters:
[0181] ;
[0182] Among them, θ t+1 is the model parameter of the t+1th round.
[0183] The specific training process of the initial MLP model in this embodiment is as follows:
[0184] (1) Batch training
[0185] Divide the dataset into multiple mini-batches, each containing s samples (for example, s=32);
[0186] In each training cycle (epoch), the following operations are performed on each mini-batch of data:
[0187] Forward propagation: calculate the predicted output for the current input;
[0188] Calculate the loss: Use the above MSE calculation formula to calculate the mean square error between the predicted value and the true value;
[0189] Back propagation: Using the aforementioned g t Calculate the gradient of the loss function with respect to each model parameter using the calculation formula for ;
[0190] Update parameters: Use the Adam optimizer to update the network weights and biases according to the previous update steps;
[0191] (2) Convergence conditions:
[0192] The training process will continue for multiple epochs until one of the following conditions is met:
[0193] When the preset number of epochs (for example, 100 epochs) is reached, the training ends;
[0194] Alternatively, when the loss function changes very little over multiple epochs (for example, the change is less than a set threshold), the model is considered to have converged and training is terminated.
[0195] Preferably, if Figure 3 As shown, in this embodiment S3, after the initial MLP model is trained using the feature set, the loss function and the Adam optimizer, the method further includes:
[0196] S32: extracting a validation set from the endoscopic image dataset;
[0197] S33: Evaluate the intermediate MLP model obtained through training using the validation set;
[0198] S34: If the evaluation passes, the intermediate MLP model is determined as the trained target MLP model; if the evaluation fails, the intermediate MLP model is continued to be trained using the feature set, the loss function and the Adam optimizer until the intermediate MLP model obtained by continued training passes the evaluation, and the intermediate MLP model that passes the evaluation is determined as the trained target MLP model.
[0199] After the model training is completed, the validation set extracted from the endoscopic image data set is used for model evaluation, which can evaluate the generalization ability of the intermediate MLP model obtained through training, and find out whether the model has overfitting and prevent overfitting; the performance of the model under different parameter configurations can be compared, which helps to select the optimal model; the hyperparameters of the model (such as learning rate, number of iterations, number of networks, etc.) can be further adjusted to facilitate finding the hyperparameter configuration that achieves the best model performance. Only when the model evaluation passes can the intermediate MLP model obtained through training be determined as the required model (i.e., the target MLP model), otherwise it is necessary to continue iterative training until the intermediate MLP model obtained through continued training passes the evaluation, and then the final target MLP model is obtained.
[0200] Specifically, in the process of using the validation set to evaluate the model, one or more of the accuracy, recall, F1 score, mean square error or other parameters can be set as evaluation indicators. A threshold is set for each evaluation indicator. When the threshold is reached, the evaluation indicator is considered to have passed, otherwise it is considered to have failed. When all selected evaluation indicators are passed, the model evaluation is passed, and when at least one selected evaluation indicator fails, the model evaluation fails.
[0201] The above method for calculating the evaluation index adopts a conventional method, and the specific details are not repeated here.
[0202] Preferably, after the model evaluation is passed to obtain the target MLP model, Figure 4 As shown, in this embodiment S4, the target MLP model is used to perform block prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks, including:
[0203] S41: Divide the endoscope image to be divided into multiple initial image blocks, each of which has an initial width and an initial height;
[0204] S42: extracting features from each initial image block to obtain a feature vector of each pixel of each initial image block;
[0205] S43: inputting all feature vectors of each initial image block into the target MLP model to predict a predicted width and a predicted height corresponding to each initial image block;
[0206] S44: using the predicted width and predicted height corresponding to each initial image block, respectively adjusting the initial width and the initial height of each initial image block to obtain a target image block corresponding to each initial image block.
[0207] In the process of executing the segmentation task of the endoscopic image to be segmented, the image is first divided into multiple initial image blocks with initial width and initial height (for example, a 3840x2160 endoscopic image to be segmented is segmented according to the initial width and initial height with a step size of 16 pixels to obtain multiple initial image blocks), and a method similar to the feature extraction in the training set is used to extract local area features to obtain the feature vector of each pixel point of each initial image block, so as to facilitate the subsequent accurate use of the target MLP model obtained in the aforementioned training process to perform the adaptive segmentation task, and efficiently predict the accurate predicted width and predicted height of each initial image block; finally, the predicted width and predicted height are used to adjust the initial width and initial height of each initial image block, so as to achieve the size adjustment of each initial image block, and then obtain multiple target image blocks after adaptive segmentation, and complete the adaptive segmentation of the high-resolution endoscopic image.
[0208] Preferably, if Figure 5 As shown, in this embodiment S4, after obtaining multiple target image blocks, the following steps are also included:
[0209] S5: performing image processing on each of the target image blocks according to a preset image processing task to obtain a processed image block corresponding to each of the target image blocks;
[0210] S6: synthesizing all processed image blocks to obtain a target endoscopic image corresponding to the endoscopic image to be segmented.
[0211] When the endoscopic image to be segmented completes adaptive segmentation to obtain multiple target image blocks, each target image block can be processed (such as filtering, enhancing, etc.) according to the conventional processing task of the endoscopic high-resolution image (i.e., the preset image processing task), and then all the processed target image blocks can be synthesized to obtain the target endoscopic image for subsequent endoscopic medical image analysis and diagnosis, providing strong support for subsequent endoscopic medical image analysis and diagnosis.
[0212] Embodiment 2
[0213] An endoscope high-resolution image adaptive blocking system is applied to the endoscope high-resolution image adaptive blocking method of embodiment 1, such as Figure 6 As shown, the system comprises:
[0214] An image acquisition module, used for acquiring an endoscopic image data set;
[0215] A feature extraction module, used to extract a training set from the endoscopic image data set, and perform feature extraction on the training set to obtain a feature set;
[0216] A model building module, used to build an initial MLP model before training, using the feature set as the input of the initial MLP model, using the block width and block height as the output of the initial MLP model, and using an Adam optimizer as a parameter optimizer of the initial MLP model;
[0217] A model training module, used to define a loss function of the initial MLP model, and train the initial MLP model using the feature set, the loss function and the Adam optimizer to obtain a trained target MLP model;
[0218] The adaptive block module is used to obtain the endoscopic image to be segmented, and use the target MLP model to perform block prediction on the endoscopic image to be segmented to obtain multiple target image blocks.
[0219] In this embodiment, a feature extraction module is used to extract a training set from the acquired endoscopic image data set and perform feature extraction, which can extract key features directly related to the endoscopic high-resolution image, facilitate subsequent model training, provide a data basis for model training, and simplify the feature space, reduce the computational complexity, and improve the efficiency and accuracy of subsequent model training; the MLP model is also called a multi-layer perceptron model. The model building module is used to build the framework of the MLP model, with the feature set as its input, the block width and block height as its output, and the Adam optimizer as its parameter optimizer. Through the above model framework, a large number of nonlinear details and features in the endoscopic high-resolution image can be more effectively captured, and it can adapt to complex, high-dimensional data. According to the processing environment, the model framework is also highly flexible and scalable, and can flexibly adapt to different endoscopic high-resolution image processing environments; when the model training module is used to train with the constructed model framework, the target MLP model obtained can efficiently and accurately process and adapt to the segmentation task of endoscopic high-resolution images, and the processing quality and speed of the segmentation task are effectively improved; finally, through the adaptive segmentation module, the target MLP model is used to perform segmentation prediction on the endoscopic image to be segmented, and the segmentation width and segmentation height of each area can be accurately and efficiently predicted for different areas, thereby realizing adaptive segmentation of the image, and obtaining multiple target image blocks of optimal quality with optimal efficiency, which is convenient for subsequent image operations.
[0220] The endoscope high-resolution image adaptive blocking system of this embodiment, based on the MLP model framework and Adam optimizer, can efficiently and accurately perform adaptive blocking on different areas of the endoscope high-resolution image, obtain multiple image blocks of optimal quality of the endoscope high-resolution image with optimal efficiency, and provide strong support for endoscopic medical image analysis and diagnosis.
[0221] The functions of each module in the endoscope high-resolution image adaptive segmentation system described in this embodiment are the same as the method steps of the endoscope high-resolution image adaptive segmentation method described in Example 1. Therefore, for the details not covered in this embodiment, please refer to Example 1 and Example 2. Figure 1 and Figure 5 The detailed description will not be repeated here.
[0222] Embodiment 3
[0223] This embodiment also provides an endoscope high-resolution image adaptive blocking device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the computer program is executed, the method steps in the endoscope high-resolution image adaptive blocking method of embodiment 1 are implemented.
[0224] Through a computer program stored in the memory and running on the processor, based on the MLP model framework and Adam optimizer, it is possible to efficiently and accurately adaptively block different areas of the endoscopic high-resolution image, and obtain multiple image blocks of the optimal quality of the endoscopic high-resolution image with optimal efficiency, providing strong support for endoscopic medical image analysis and diagnosis.
[0225] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of a computer device, and uses various interfaces and lines to connect various parts of the entire computer device.
[0226] The memory can be used to store computer programs and / or models. The processor realizes various functions of the computer device by running or executing the computer programs and / or models stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, video data, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0227] It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by a computer program. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0228] These computer programs may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0229] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0230] This embodiment further provides a computer storage medium, which includes: at least one instruction, which implements the method steps in the endoscope high-resolution image adaptive blocking method of embodiment 1 when the instruction is executed by a computer.
[0231] By executing a computer storage medium containing at least one instruction, based on the MLP model framework and the Adam optimizer, efficient and accurate adaptive segmentation can be performed on different areas of the endoscopic high-resolution image, and multiple image blocks of optimal quality of the endoscopic high-resolution image can be obtained with optimal efficiency, providing strong support for endoscopic medical image analysis and diagnosis.
[0232] Similarly, for details not yet included in this embodiment, see Embodiment 1, Embodiment 2 and Figures 1 to 6 The detailed description will not be repeated here.
[0233] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. An adaptive segmentation method for high-resolution endoscope images, characterized in that: The method comprises: Acquire an endoscopic image data set, extract a training set from the endoscopic image data set, and perform feature extraction on the training set to obtain a feature set; Constructing an initial MLP model before training, using the feature set as the input of the initial MLP model, using the block width and the block height as the output of the initial MLP model, and using the Adam optimizer as the parameter optimizer of the initial MLP model; Defining a loss function of the initial MLP model, and training the initial MLP model using the feature set, the loss function, and the Adam optimizer to obtain a trained target MLP model; Acquire an endoscopic image to be segmented, and use the target MLP model to perform segmentation prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks; The training set includes a plurality of training images; Perform feature extraction on the training set to obtain a feature set, including: Select any one training image from the training set, and divide the selected training image into a plurality of image regions; Select an image region from the selected training image, and calculate the horizontal gradient and the vertical gradient of each pixel of the image region selected from the selected training image; According to the horizontal direction gradient and the vertical direction gradient of each pixel point of the image area selected in the selected training image, the texture energy of each pixel point of the image area selected in the selected training image is calculated respectively; Calculate the standard deviation of the selected image area in the selected training image in the RGB color channel; Constructing a feature vector of each pixel of the image region selected in the selected training image according to the texture energy of each pixel of the image region selected in the selected training image and the standard deviation of the image region selected in the selected training image in the RGB color channel; Traversing each image region in the selected training image, and constructing the feature vector of each pixel point of each image region in the selected training image in the same way; and obtaining the feature vector set corresponding to the selected training image according to all the feature vectors of all the image regions in the selected training image; Traversing each training image in the endoscopic image data set, and obtaining a feature vector set corresponding to each training image in a one-to-one manner according to the same method; The feature set is obtained according to all feature vector sets in all training images.
2. The method for adaptively segmenting endoscopic high-resolution images according to claim 1, characterized in that: Assume that the selected training image is the ath training image in the endoscope image data set, and the selected image region is the bth image region in the ath training image; Then the horizontal gradient and vertical gradient of each pixel in the image area selected in the selected training image are calculated respectively, including: The Sobel operator is used to calculate the horizontal gradient and vertical gradient of each pixel in the b-th image area in the a-th training image. The specific formula for calculating the horizontal gradient and vertical gradient of the bth image area in the ath training image at the pixel point (x, y) is: ; in, and are the horizontal and vertical gradients of the b-th image region in the a-th training image at the pixel point (x, y), i and j are the horizontal and vertical indexes of the Sobel operator in the window (3, 3), respectively. and are the pixel grayscale values at pixel point (x, y) and pixel point (x+i, y+j) in the bth image area in the ath training image, respectively. and They are the horizontal convolution kernel and vertical convolution kernel of the b-th image area in the a-th training image at the pixel point (i, j) in the window (3, 3), respectively.
3. The method for adaptively segmenting endoscopic high-resolution images according to claim 2, characterized in that: The specific formula for calculating the texture energy of the bth image region at the pixel point (x, y) in the ath training image is: ; in, is the texture energy of the bth image region in the ath training image at the pixel point (x, y), N is the total number of pixels in the bth image region in the ath training image, N x and N y are the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the b-th image area in the a-th training image, respectively. and are the horizontal and vertical gradients of the bth image region in the ath training image at the pixel point (x+i, y+j), respectively.
4. The method for adaptively segmenting endoscopic high-resolution images according to claim 3, characterized in that: The specific formula for calculating the standard deviation of the bth image area in the ath training image in the RGB color channel is: ; Among them, the RGB color channels include R channel, G channel and B channel. , and are the R channel standard deviation, G channel standard deviation and B channel standard deviation of the b-th image region in the a-th training image, respectively. , and are the R channel color value, G channel color value, and B channel color value of the k-th pixel in the b-th image area in the a-th training image, respectively. , and are the average color value of the R channel, the average color value of the G channel, and the average color value of the B channel of the bth image area in the ath training image, respectively.
5. The method for adaptively segmenting endoscopic high-resolution images according to claim 4, characterized in that: The constructed feature vector of the bth image region at the pixel point (x, y) in the ath training image is: ; in, is the feature vector of the bth image region at the pixel point (x, y) in the ath training image.
6. The method for adaptively segmenting endoscopic high-resolution images according to claim 1, characterized in that: The constructed initial MLP model includes an input layer, n hidden layers and an output layer; In the constructed initial MLP model, the feature set is the input of the input layer, the output of the input layer is the input of the first hidden layer, and the activation signal is transmitted between adjacent hidden layers through model parameters; the output of the last hidden layer is the input of the output layer, and the block width and block height are both outputs of the output layer; The input layer includes a plurality of neurons; in the input layer, the feature set is the input of the first neuron in the input layer, and each of the remaining neurons receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, wherein the output of the last neuron in the input layer is the input of the first hidden layer; the number of neurons in the input layer is equal to the dimension of the feature vector of each pixel in each training image; In each hidden layer, the hidden layer includes multiple neurons; the output of the last neuron in the input layer is the input of the first neuron in the first hidden layer, and each of the remaining neurons in each hidden layer receives the output from the previous neuron as input through a weighted connection, and uses its own output as the input of the next neuron, and the output of the last neuron in the last hidden layer is the input of the output layer; the number of neurons in each hidden layer is equal or unequal.
7. The method for adaptively segmenting endoscopic high-resolution images according to claim 6, characterized in that: The output of the first hidden layer among the n hidden layers is: ; in, is the output of the first hidden layer, The input transmitted from the input layer to the first hidden layer is the feature vector in the feature set; and are the network weights and biases of the first hidden layer, is the weighted sum between the input and bias of the first hidden layer, represents the activation function; The output of the mth hidden layer among n hidden layers is: ; in, is the output of the mth hidden layer, and are the network weights and biases of the mth hidden layer, is the weighted sum between the input and bias of the mth hidden layer, is the output of the m-1th hidden layer, m satisfies 2≤m≤n; The output of the output layer is: ; in, is the output of the output layer, and are the block width and block height respectively; is the output of the nth hidden layer, and are the network weights and bias of the output layer, respectively.
8. The method for adaptively segmenting endoscopic high-resolution images according to claim 1, characterized in that: The loss function of the initial MLP model is specifically the mean square error of the initial MLP model; the model parameters of the initial MLP model specifically include network weights and biases; Training the initial MLP model using the feature set, the loss function and the Adam optimizer includes: The feature set is input into the initial MLP model, and the initial MLP model is trained with the minimization of the mean square error as a training objective. During the training process of the initial MLP model, the network weights and biases are iteratively optimized using an Adam optimizer.
9. The method for adaptively segmenting endoscopic high-resolution images according to claim 1, characterized in that: After training the initial MLP model using the feature set, the loss function, and the Adam optimizer, the method further includes: extracting a validation set from the endoscopic image dataset; Using the validation set to evaluate the trained intermediate MLP model; If the evaluation passes, the intermediate MLP model is determined as the trained target MLP model; if the evaluation fails, the intermediate MLP model is continued to be trained using the feature set, the loss function and the Adam optimizer until the intermediate MLP model obtained by continued training passes the evaluation, and the intermediate MLP model that passes the evaluation is determined as the trained target MLP model.
10. The method for adaptively segmenting endoscopic high-resolution images according to claim 1, characterized in that: The target MLP model is used to perform block prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks, including: Dividing the endoscope image to be divided into multiple initial image blocks, each initial image block has an initial width and an initial height; Perform feature extraction on each initial image block respectively to obtain a feature vector of each pixel point of each initial image block; Input all feature vectors of each initial image block into the target MLP model to predict the predicted width and predicted height corresponding to each initial image block; The predicted width and the predicted height corresponding to each initial image block are used to adjust the initial width and the initial height of each initial image block respectively, so as to obtain a target image block corresponding to each initial image block.
11. The method for adaptively segmenting an endoscope high-resolution image according to any one of claims 1 to 10, characterized in that: After obtaining multiple target image blocks, it also includes: According to the preset image processing task, each of the target image blocks is subjected to image processing to obtain a processed image block corresponding to each of the target image blocks; All processed image blocks are synthesized to obtain a target endoscopic image corresponding to the endoscopic image to be segmented.
12. An endoscope high-resolution image adaptive blocking system, characterized in that: Applied to the method for adaptively segmenting an endoscopic high-resolution image according to any one of claims 1 to 11, the system comprises: An image acquisition module, used for acquiring an endoscopic image data set; A feature extraction module, used to extract a training set from the endoscopic image data set, and perform feature extraction on the training set to obtain a feature set; A model building module, used to build an initial MLP model before training, using the feature set as the input of the initial MLP model, using the block width and block height as the output of the initial MLP model, and using an Adam optimizer as a parameter optimizer of the initial MLP model; A model training module, used to define a loss function of the initial MLP model, and train the initial MLP model using the feature set, the loss function and the Adam optimizer to obtain a trained target MLP model; An adaptive block module is used to obtain an endoscopic image to be segmented, and use the target MLP model to perform block prediction on the endoscopic image to be segmented to obtain a plurality of target image blocks; The training set includes a plurality of training images; The feature extraction module extracts features from the training set to obtain a feature set, including: Select any one training image from the training set, and divide the selected training image into a plurality of image regions; Select an image region from the selected training image, and calculate the horizontal gradient and the vertical gradient of each pixel of the image region selected from the selected training image; According to the horizontal direction gradient and the vertical direction gradient of each pixel point of the image area selected in the selected training image, the texture energy of each pixel point of the image area selected in the selected training image is calculated respectively; Calculate the standard deviation of the selected image area in the selected training image in the RGB color channel; Constructing a feature vector of each pixel of the image region selected in the selected training image according to the texture energy of each pixel of the image region selected in the selected training image and the standard deviation of the image region selected in the selected training image in the RGB color channel; Traversing each image region in the selected training image, and constructing the feature vector of each pixel point of each image region in the selected training image in the same way; and obtaining the feature vector set corresponding to the selected training image according to all the feature vectors of all the image regions in the selected training image; Traversing each training image in the endoscopic image data set, and obtaining a feature vector set corresponding to each training image in a one-to-one manner according to the same method; The feature set is obtained according to all feature vector sets in all training images.
13. An endoscope high-resolution image adaptive blocking device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method steps in the method for adaptively segmenting endoscopic high-resolution images as claimed in any one of claims 1 to 11 when executed.
14. A computer storage medium, characterized in that: The computer storage medium comprises: at least one instruction, which, when executed by a computer, implements the method steps in the method for adaptively segmenting endoscopic high-resolution images according to any one of claims 1 to 11.
Citation Information
Patent Citations
A pedestrian rerecognition method based on reinforcement learning adaptive partitioning
CN109086672A