WCE image scene classification method, device and equipment based on reinforcement learning

By combining deep learning with hybrid reinforcement learning, and utilizing convolutional neural networks and recurrent neural networks for WCE image scene classification, the problem of insufficient classification accuracy in existing technologies is solved. This achieves efficient and accurate capsule endoscopy image classification, reduces reading time, and improves detection accuracy and efficiency.

CN120411602BActive Publication Date: 2026-02-03SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510442680.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2026-02-03
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Existing WCE image classification algorithms lack sufficient classification accuracy, resulting in large classification errors. There is currently no effective solution, which affects the ability to classify and predict image scenes.

Method used

We employ a framework combining deep learning and hybrid reinforcement learning. We utilize pre-defined convolutional neural networks and recurrent neural networks for feature extraction and image enhancement, combine Q-Learning algorithm to optimize image classification, improve classification accuracy through hybrid reinforcement learning methods, and evaluate the credibility of classification results using a probability thresholding method.

Benefits of technology

It improves the accuracy and efficiency of WCE image scene classification, reduces reading time, alleviates the workload of doctors, and automatically adjusts image enhancement correction test results to ensure optimized image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411602B_ABST
    Figure CN120411602B_ABST
Patent Text Reader

Abstract

The application provides a WCE image scene classification method and device based on reinforcement learning and equipment, and relates to the technical field of computer analysis of medical images, and the method comprises the following steps: performing feature extraction on each current image frame in all current capsule endoscopy images to obtain a current image frame feature sequence; inputting all current image frame feature sequences into a preset recurrent neural network model to obtain a first probability prediction result of each current image frame; performing image processing on a secondary classification image frame to obtain a processed image frame, performing image enhancement on the processed image frame according to a target action to obtain an enhanced image frame; inputting the enhanced image frame into the preset recurrent neural network model to obtain a second probability prediction result of each enhanced image frame; and determining a target prediction result according to the first probability prediction result and the second probability prediction result. The application can reduce the reading time of WCE, reduce the work burden and improve the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer analysis technology for medical images, and in particular to a WCE image scene classification method, apparatus, and device based on reinforcement learning. Background Technology

[0002] Wireless Capsule Endoscopy (WCE) is a novel non-invasive gastrointestinal examination tool, also known simply as capsule endoscopy. Clinically, after the patient swallows the capsule endoscope, it moves forward under its own weight and the peristaltic movement of the gastrointestinal tract. The capsule endoscope imaging system automatically captures the entire digestive tract scene at a rate of two frames per second. The captured image data is wirelessly transmitted to an external receiving terminal. During an average 7-8 hour digestive tract examination, tens of thousands of color images of the digestive tract scene can be recorded. The human digestive system consists of a series of different organs, including the stomach, small intestine, and large intestine. The small intestine is further divided into the duodenum, jejunum, and ileum. Therefore, each WCE image contains a scene of the digestive tract organs.

[0003] Due to the slow imaging speed of capsule endoscopy and the peristaltic characteristics of the digestive tract, the image acquisition time for a single case is long and the number of sequences is enormous. Even for experienced clinicians, screening images containing suspicious lesions from capsule endoscopy data requires 3-5 hours of focused, accurate interpretation of a single case – a demanding and tedious task. Therefore, automatic classification of the stomach, small intestine, and large intestine can reduce WCE reading time and alleviate workload. However, existing classification algorithms lack sufficient accuracy, resulting in classification errors, and there are no effective strategies to address these errors, leading to poor image scene classification and prediction capabilities. Currently, there is no technical solution that can solve these problems, and no reinforcement learning-based WCE image scene classification method, device, or equipment exists. Summary of the Invention

[0004] This invention provides a WCE image scene classification method, device, and equipment based on reinforcement learning. It utilizes a framework combining deep learning and hybrid reinforcement learning to efficiently and accurately classify various scenes in capsule endoscopy images.

[0005] In a first aspect, the present invention provides a WCE image scene classification method based on reinforcement learning, comprising:

[0006] The features of each current image frame in all current capsule endoscopy images are extracted using a pre-defined convolutional neural network model to obtain the feature sequence of each current image frame. The feature sequences of all current image frames are then input into a pre-defined recurrent neural network model to obtain the first probability prediction result of each current image frame output by the pre-defined recurrent neural network model.

[0007] All current image frames whose first probability prediction result is less than a preset threshold are identified as secondary classification image frames. For any secondary classification image frame, image processing is performed on the secondary classification image frame to obtain a processed image frame. The target action is selected using the Q-Learning algorithm. Based on the target action, the processed image frame is enhanced to obtain an enhanced image frame.

[0008] The enhanced image frame is input into the preset recurrent neural network model to obtain the second probability prediction result of each enhanced image frame output by the preset recurrent neural network model;

[0009] Based on the first probability prediction result and the second probability prediction result, determine the target prediction result for each current image frame in all current capsule endoscopy images;

[0010] The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0011] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, before performing feature extraction on each current image frame in all current capsule endoscopy images using a preset convolutional neural network model, the method further includes:

[0012] Acquire raw capsule endoscopy images using capsule endoscopy;

[0013] Each original capsule endoscopy image is stretched and compressed at the edges to obtain all current capsule endoscopy images.

[0014] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, before inputting all current image frame feature sequences into a preset recurrent neural network model, the method further includes:

[0015] The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample image frame feature sequences. The sample image frame feature sequences are then input into the initial recurrent neural network model to obtain sample probability prediction results.

[0016] Sample actions are sampled using a multinomial distribution to obtain the sample results corresponding to the sample actions. The current reward value is obtained by using the sample label results and the sample results corresponding to the sample actions. The current state is determined based on the current reward value and the reward value before sampling. The first Q value is continuously updated using the current state.

[0017] The expected cumulative return is determined by using the action and reward value corresponding to the maximum first Q value. The gradient of the expected cumulative return is calculated, and the model parameters corresponding to the initial recurrent neural network model are updated using a preset formula to obtain the target model parameters. The preset recurrent neural network model is then determined based on the target model parameters.

[0018] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, the step of extracting features from all sample image frames using a preset convolutional neural network model to obtain a sample image frame feature sequence includes:

[0019] The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample extracted features.

[0020] For each sample image frame, it is stored in HDF5 format, and the patient name, total number of frames, and sample label results of the case corresponding to the extracted features of the sample are recorded.

[0021] The feature sequence of the sample image frames is V = {X1, X2, ..., X...} T}, where X t Let V represent the feature vector of frame t, where T is the total number of frames and V represents the case.

[0022] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, the step of updating the model parameters corresponding to the initial recurrent neural network model using a preset formula to obtain the target model parameters includes:

[0023]

[0024] Where α is the learning rate, β1 and β2 are the hyperparameters of the balancing weights, and L CE Let L be the cross-entropy loss function. weight Here, J is the regularization term, and J is the expected cumulative return. Indicates (-J+β1L) CE +β2L weight The gradient is obtained by taking the global derivative of the parameter θ in the equation.

[0025] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, the first probability prediction result is the predicted probability of different scene classifications, and the scene classifications include stomach, duodenum, jejunum, ileum and large intestine.

[0026] The step of determining all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames includes:

[0027] For any current image frame, if the predicted probability of each scene classification in the first probability prediction result is less than the preset threshold, the current image frame is determined to be a secondary classification image frame.

[0028] Iterate through all current image frames until all secondary classification image frames are determined.

[0029] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, the image processing of the secondary classification image frame includes rotating the image and scaling the image;

[0030] The process of selecting a target action using the Q-Learning algorithm and then enhancing the processed image frame based on the target action to obtain an enhanced image frame includes:

[0031] For any processed image frame, feature extraction is performed on the processed image frame using a preset convolutional neural network model to obtain the processed image features, and the first standard deviation of the current processed image frame is calculated.

[0032] Calculate the second standard deviation of the secondary classification image frame corresponding to the current processed image frame, and determine the target reward value based on the first standard deviation and the second standard deviation;

[0033] The second Q value is updated according to the target reward value, and the action corresponding to the maximum second Q value is used to process the secondary classification image frame to obtain the enhanced image frame.

[0034] Traverse all processed image frames to obtain all enhanced image frames.

[0035] According to the reinforcement learning-based WCE image scene classification method provided by the present invention, the step of determining the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result includes:

[0036] For any enhanced image frame, if the probability prediction result is greater than the preset threshold, the second probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame; otherwise, the first probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame.

[0037] Traverse all enhanced image frames, determine the final prediction result corresponding to each enhanced image frame, determine all current image frames whose first probability prediction result is greater than or equal to the preset threshold as unenhanced image frames, and determine the first probability prediction result corresponding to the unenhanced image frame as the final prediction result of the unenhanced image frame.

[0038] Based on the final prediction results of the enhanced image frames and the final prediction results of the unenhanced image frames, the target prediction result for each current image frame in all current capsule endoscopy images is determined.

[0039] Secondly, a WCE image scene classification device based on reinforcement learning is provided, including:

[0040] The input unit is used to extract features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence of each current image frame, and input all current image frame feature sequences into a preset recurrent neural network model to obtain a first probability prediction result of each current image frame output by the preset recurrent neural network model.

[0041] The image processing unit is used to determine all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames. For any secondary classification image frame, the image processing unit performs image processing on the secondary classification image frame to obtain a processed image frame. The target action is selected using the Q-Learning algorithm. The image enhancement is performed on the processed image frame according to the target action to obtain an enhanced image frame.

[0042] An output unit is used to input the enhanced image frame into the preset recurrent neural network model to obtain a second probability prediction result for each enhanced image frame output by the preset recurrent neural network model.

[0043] A determining unit is configured to determine the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result.

[0044] The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0045] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the reinforcement learning-based WCE image scene classification method.

[0046] This invention reduces WCE (Wide-Cut Image Detection) reading time and workload by automatically classifying the stomach, small intestine, and large intestine, while accurately locating target scenes. This is significant for deep learning-based automatic image detection algorithms, helping to improve detection accuracy and efficiency. This invention focuses on classifying scenes such as the stomach, duodenum, jejunum, ileum, and large intestine in capsule endoscopy (CE) image sequences, proposing a framework combining deep learning and hybrid reinforcement learning for efficient and accurate scene classification in CE images. It uses hybrid reinforcement learning, combining Monte Carlo policy gradient (REINFORCE) and action value function (Q-Learning) methods to classify WCE image scenes, rather than using a single reinforcement learning method. This not only compensates for the shortcomings of single reinforcement learning methods and improves classification accuracy, but also automatically distinguishes between correctly classified images and images with uncertain classification results by evaluating the credibility of the classification results through probability thresholding during the automatic image enhancement correction test results stage. For the latter, image enhancement is performed to eliminate potential classification errors. Image enhancement adjusts the enhancement strategy by estimating action value, thereby ensuring optimized image quality. This invention not only optimizes the performance of the classification model using hybrid reinforcement learning, but also improves the ability to handle uncertain classification results using secondary prediction, thus having strong practical application value. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0048] Figure 1 This is one of the flowcharts of the WCE image scene classification method based on reinforcement learning provided by the present invention;

[0049] Figure 2 This is the second flowchart of the WCE image scene classification method based on reinforcement learning provided by the present invention;

[0050] Figure 3 This is a diagram of the training network structure provided by the present invention;

[0051] Figure 4 This is a test network structure diagram provided by the present invention;

[0052] Figure 5 This is a schematic diagram of the capsule endoscope image before deformation provided by the present invention;

[0053] Figure 6 This is a schematic diagram of the deformed capsule endoscope image provided by the present invention;

[0054] Figure 7 This is a flowchart of the Q-Learning algorithm provided by the present invention;

[0055] Figure 8 This is the third flowchart of the WCE image scene classification method based on reinforcement learning provided by the present invention;

[0056] Figure 9 This is a schematic diagram of the structure of the WCE image scene classification device based on reinforcement learning provided by the present invention;

[0057] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0059] Figure 1 This is one of the flowcharts illustrating the WCE image scene classification method based on reinforcement learning provided by the present invention. Specifically, it relates to scene classification processing of capsule endoscopy image data based on capsule endoscopy equipment. The WCE image scene classification method based on reinforcement learning includes:

[0060] Step 101: Use a preset convolutional neural network model to extract features from each current image frame in all current capsule endoscopy images to obtain the feature sequence of each current image frame. Input all current image frame feature sequences into the preset recurrent neural network model to obtain the first probability prediction result of each current image frame output by the preset recurrent neural network model.

[0061] Step 102: Determine all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames. For any secondary classification image frame, perform image processing on the secondary classification image frame to obtain a processed image frame. Use the Q-Learning algorithm to select the target action. Perform image enhancement on the processed image frame according to the target action to obtain an enhanced image frame.

[0062] Step 103: Input the enhanced image frame into the preset recurrent neural network model to obtain the second probability prediction result of each enhanced image frame output by the preset recurrent neural network model;

[0063] Step 104: Determine the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result;

[0064] The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0065] In step 101, before performing feature extraction on each current image frame in all current capsule endoscopy images using a preset convolutional neural network model, the method further includes:

[0066] Acquire raw capsule endoscopy images using capsule endoscopy;

[0067] Each original capsule endoscopy image is stretched and compressed at the edges to obtain all current capsule endoscopy images.

[0068] In an optional embodiment, a capsule endoscopy image contains 50,000 to 80,000 frames, and a single original capsule endoscopy image is as follows: Figure 5 As shown, the image is surrounded by a distinct black background and circular boundaries. However, these areas do not contain useful information. When extracting features directly from the original image, the black background and edges cause visual contamination, affecting subsequent task completion. Therefore, this invention introduces a geometric transformation method based on spatial coordinate mapping. This method adjusts the spatial distribution of image pixels to highlight the features of the central region of the image, while compressing or removing redundant peripheral regions. The mathematical expression for the deformation is as follows:

[0069]

[0070] Where x and y are the positions of the image pixels in the original coordinate system; u and v are the new positions of the pixels in the deformed coordinate system. This formula stretches and compresses the image edges, increasing the information density of the central region while gradually weakening the information density of the edge regions. After deformation processing, as shown... Figure 6 As shown, redundant backgrounds and circular boundaries of WCE image frames were removed, while retaining the maximum amount of information.

[0071] In step 101, a pre-defined convolutional neural network (CNN) model is used to extract features from each current image frame in all current capsule endoscopy images, resulting in a feature sequence for each current image frame. Specifically, the deformed data, on a case-by-case basis, needs to undergo feature extraction using the pre-defined CNN model. The CNN extracts local spatial features from the input single-frame image, including texture, edge color distribution, and shape information. This information is crucial for scene classification of WCE image frames. Through convolution operations, the CNN can capture local features at different scales, thereby supporting subsequent temporal analysis and classification. The CNN used here is a ResNet50 network, loaded with weights pre-trained on the ImageNet dataset.

[0072] ResNet50 is a 50-layer residual network that performs exceptionally well on various computer vision tasks. ResNet50 introduces residual connections, learning "residual mappings" instead of direct mappings, making the network easier to optimize. Its moderate depth strikes a good balance between feature extraction capabilities and computational cost, making it ideal for real-time analysis scenarios such as capsule endoscopy videos. By loading weights pre-trained on the ImageNet dataset, model convergence time can be significantly reduced while avoiding the high computational cost of training from scratch.

[0073] After completing CNN feature extraction, the system stores the features of each WCE case in HDF5 format, and records the patient name, total number of frames, and label information for that case. The label set Y is expressed as follows, where C is the number of classification categories (the classification categories are stomach, duodenum, jejunum, ileum, and large intestine), and y... t Let represent the t-th frame in set Y, where T is the total number of frames.

[0074] Y = {y t |y t ∈{0,1,…,C},t=1,…T};

[0075] This structured storage method lays the foundation for subsequent analysis and classification tasks.

[0076] Optionally, the step of extracting features from all sample image frames using a preset convolutional neural network model to obtain a sample image frame feature sequence includes:

[0077] The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample extracted features.

[0078] For each sample image frame, it is stored in HDF5 format, and the patient name, total number of frames, and sample label results of the case corresponding to the extracted features of the sample are recorded.

[0079] The obtained feature sequences are input into the BiRNN network according to the cases. The WCE image sequence of each case is represented as a sample image frame feature sequence, which is V={X1,X2,…,X…} T}, where X t Let V represent the feature vector of frame t, where T is the total number of frames and V represents the case.

[0080] Optionally, the default recurrent neural network model, BiRNN, captures the temporal dependencies between frames through a recurrent structure:

[0081] h t =f(W h ·h t-1 +W x ·x t +b h );

[0082] Among them, h t W represents the hidden state of frame t. h W x Let b be the weight matrix. h Here, f(·) is the bias term, and f(·) is the activation function. A bidirectional recurrent neural network combines forward and backward time information, as shown in the following equation, by concatenating the forward hidden states. and the state of hiding behind Capture contextual information.

[0083]

[0084] The classification probability of each frame is calculated through the Softmax layer, as shown in the following formula, where P(x t W is the classification probability distribution of frame t (classification categories are stomach, duodenum, jejunum, ileum, and large intestine). i and b o These are the weights and biases of the output layer, respectively. It is the hidden state of the t-th frame under bidirectional bi.

[0085]

[0086] This invention supports various recurrent neural network structures, such as bidirectional long short-term memory networks (BiLSTM), bidirectional gated recurrent units (BiGRU), and Transformer encoders.

[0087] Optionally, before inputting all current image frame feature sequences into a preset recurrent neural network model, the method further includes:

[0088] The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample image frame feature sequences. The sample image frame feature sequences are then input into the initial recurrent neural network model to obtain sample probability prediction results.

[0089] Sample actions are sampled using a multinomial distribution to obtain the sample results corresponding to the sample actions. The current reward value is obtained by using the sample label results and the sample results corresponding to the sample actions. The current state is determined based on the current reward value and the reward value before sampling. The first Q value is continuously updated using the current state.

[0090] The expected cumulative return is determined by using the action and reward value corresponding to the maximum first Q value. The gradient of the expected cumulative return is calculated, and the model parameters corresponding to the initial recurrent neural network model are updated using a preset formula to obtain the target model parameters. The preset recurrent neural network model is then determined based on the target model parameters.

[0091] In this optional embodiment, the present invention constructs a hybrid reinforcement learning framework, optimizing the parameters of the BiRNN network. Reinforcement learning (RL) is an important branch of machine learning, which learns optimal policies to maximize long-term rewards through agent interaction with the environment. There are various methods for reinforcement learning; this invention combines the advantages of Monte Carlo policy gradient (REINFORCE) and action-value function (Q-Learning) algorithms, while simultaneously optimizing the policy distribution to improve the agent's exploration ability and sample efficiency.

[0092] In each round of training, calculations are performed case-by-case, and the training process is the same for each case. First, the probability distribution P(x) of each frame in each case is obtained from the BiRNN output by the activation function. t Action 'a' is sampled from a probability distribution. The strategy used for action sampling is a multinomial distribution, as shown in the following formula:

[0093] a t ~CategoricalP(x t ));

[0094] Sampling is performed on the probability distribution of each frame to obtain the predicted category of that frame, resulting in a sequence A with predicted categories, as shown in the following formula: a t This represents the sampled data in frame t:

[0095] A={a t |a t∈ {0, 1, ..., C}, et = 1, ..., T};

[0096] The reward value is obtained by comparing the sampled set A with the label Y. As shown in the following formula, Vi is the i-th case, and is the reward for the i-th case.

[0097]

[0098] Construct the first Q-value storage table, initialize the first Q-value to 0, set the vertical axis of the Q-table as the mean of the sampled action probabilities, and the horizontal axis as the state. Set the state to two, and the current state is determined after sampling. Let R0 be the initial reward value when the probability of the RNN output is not sampled, and the reward after sampling is R1. The new state is determined according to whether R1 is less than, equal to, or greater than R0 respectively. When R1 ≥ R0, it is state s1, and when R1 < R0, it is s2. a n is the n-th action. The initial Q-table is as follows:

[0099] Table 1 Initial Q-table

[0100] Status / Action <![CDATA[a1]]> <![CDATA[a2]]> a… <![CDATA[a n ]]> <![CDATA[s1(R1≥R0)]]> 0 0 … 0 <![CDATA[s2(R1<R0)]]> 0 0 … 0

[0101] The update of the first Q-value in Table 1 is based on the Bellman equation, and the equation formula is as follows:

[0102]

[0103] where α is the learning rate, θ is the discount factor, s is the state, a is the action, R is the reward value, is the maximum first Q-value of the next state. After cycling N times, select the action corresponding to the maximum first Q-value in the Q-table and its corresponding reward value. The flowchart of the Q-Learning algorithm is as Figure 7 shown.

[0104] The goal of the above embodiment is to maximize the expected cumulative return of the policy:

[0105]

[0106] where p θ (a 1:T ) represents the probability distribution of the action corresponding to the maximum first Q-value after sampling. T is the total number of frames of the case, and the calculation of R can refer to the above calculation formula, π θ is defined by the network trained based on hybrid reinforcement learning. To maximize J(θ), the gradient of the policy parameter θ needs to be calculated. The formula for the policy gradient is:

[0107]

[0108] Among them, a t The action taken at time t, h t It is a hidden state from the RNN.

[0109] At the same time, the Cross-Entropy Loss (L) function is introduced. CE As one of the loss terms in model optimization, it measures the difference between the predicted distribution and the true distribution. The formula is as follows:

[0110]

[0111] Where T is the total number of frames, C is the number of classification categories, and y t It is the label of frame t, a t,i The action sampled in frame t is of category i.

[0112] In addition, an l2 regularization term was added to the weight parameter θ to avoid overfitting, as shown in the following formula:

[0113]

[0114] Where, θ i,j This represents the weight value of a certain layer or connection in the model, and the indices i and j represent the positions in the weight matrix.

[0115] Combining the above formulas, the step of updating the model parameters corresponding to the initial recurrent neural network model using a preset formula to obtain the target model parameters includes:

[0116]

[0117] Where α is the learning rate, β1 and β2 are the hyperparameters of the balancing weights, and L CE Let L be the cross-entropy loss function. weight As the regularization term, J is the expected cumulative return mentioned above (the formula is...). This indicates that (-J+β1L) CE +β2L weight The gradient is obtained by taking the global derivative of the parameter θ in the equation.

[0118] By continuously updating θ using the above method, a preset recurrent neural network model based on reinforcement learning can be trained using Adam as the optimizer. This model is used for the prediction of WCE sequence frames, so that in step 101, all current image frame feature sequences can be input into the preset recurrent neural network model to obtain the first probability prediction result of each current image frame output by the preset recurrent neural network model.

[0119] In step 102, the first probability prediction result is the predicted probability of different scene classifications, and the scene classifications include stomach, duodenum, jejunum, ileum and large intestine;

[0120] The step of determining all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames includes:

[0121] For any current image frame, if the predicted probability of each scene classification in the first probability prediction result is less than the preset threshold, the current image frame is determined to be a secondary classification image frame.

[0122] Iterate through all current image frames until all secondary classification image frames are determined.

[0123] Optionally, this invention performs preliminary screening of the input test data, extracts its features, and determines whether it meets the requirements. First, the input test case image sequence V = {X1, X2, ..., X...} is entered. T}, X t Let X represent the image at frame t. For each frame image X... t The input is fed into the first-stage CNN network to extract features, and the extracted features are represented as follows:

[0124]

[0125] Among them, F i Let represent the feature vector corresponding to the i-th frame image, with dimension d. The extracted features are input into the model trained in the first stage to predict the result and obtain the classification probability distribution for each frame. The model outputs the probability values ​​for each category of each frame image. A probability threshold λ is set, where λ is 0.6.

[0126] If the predicted probability of a certain category is greater than or equal to the threshold, the image is considered correctly classified. If the probability is lower than the threshold, the image is considered incorrectly classified and requires further processing. Based on this determination, the image is divided into two categories, and the image D is classified accordingly. pass Image D (above the threshold) and uncertain classification opt (Below the threshold).

[0127] Calculate the feature standard deviation M for uncertain frame images i The formula is shown below:

[0128]

[0129] Where i represents the i-th frame, F i,j μ is the j-th component of the feature vector F in the i-th frame. i It is F i The mean of the terms is calculated using the following formula:

[0130]

[0131] Optionally, the image processing of the secondary classification image frame includes rotating and scaling the image;

[0132] The process of selecting a target action using the Q-Learning algorithm and then enhancing the processed image frame based on the target action to obtain an enhanced image frame includes:

[0133] For any processed image frame, feature extraction is performed on the processed image frame using a preset convolutional neural network model to obtain the processed image features, and the first standard deviation of the current processed image frame is calculated.

[0134] Calculate the second standard deviation of the secondary classification image frame corresponding to the current processed image frame, and determine the target reward value based on the first standard deviation and the second standard deviation;

[0135] The second Q value is updated according to the target reward value, and the action corresponding to the maximum second Q value is used to process the secondary classification image frame to obtain the enhanced image frame.

[0136] Traverse all processed image frames to obtain all enhanced image frames.

[0137] After initially identifying image frames with uncertain predictions, a multi-view feature transformation approach is employed to improve their quality. This is combined with the Q-Learning algorithm of reinforcement learning to find the optimal action for image augmentation. The operations performed on these image frames are image rotation and image scaling. In this section, i represents the i-th frame. The rotated image is as follows:

[0138]

[0139] Where R(X) i Let ,θ) represent the transformation function that rotates the i-th frame image X by an angle θ. The scaled image is:

[0140]

[0141] Where S(X) i (,s) represents the transformation function with respect to the scaling factor s of the i-th frame image X. The generated transformation data set is X of each transformed image i 变换 The input is fed into the first-stage CNN to extract the corresponding feature vector F. i 变换 Then, the feature standard deviation M is calculated for the uncertain frame image. i The formula is used to calculate its characteristic standard deviation M. i 变换 By combining reinforcement learning, the optimal action 'a' for each input image is found.i The final transformation is selected for secondary prediction. Two states are designed: state S1 occurs when the standard deviation M1 of the transformed features is greater than or equal to the standard deviation M of the untransformed features; otherwise, state S2 occurs. The reward r is also designed based on the standard deviation of the features, as detailed below:

[0142]

[0143] When M1 is greater than M, the reward value is increased by one; when M1 is equal to M, the reward value is zero; when M1 is less than M, the reward value is decreased by one.

[0144] The Q-table of Q-Learning iterates N = a × m, where a is the number of actions (3) and m is a constant. After each iteration, the Q-table contains states and actions represented as Q(s,a). The Q-Learning update rule follows the Bellman equation. The action corresponding to the second largest Q value is selected from the Q-table. Generate optimized image The optimized dataset is denoted as

[0145] In step 103, the optimized dataset from step 102... The input will be fed into a pre-trained convolutional neural network (CNN) for feature extraction. The extracted features will be further input into a preset recurrent neural network model that has been trained in step 101 above. The second probability prediction result of each enhanced image frame will be output by the preset recurrent neural network model.

[0146] In step 104, determining the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result includes:

[0147] For any enhanced image frame, if the probability prediction result is greater than the preset threshold, the second probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame; otherwise, the first probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame.

[0148] Traverse all enhanced image frames, determine the final prediction result corresponding to each enhanced image frame, determine all current image frames whose first probability prediction result is greater than or equal to the preset threshold as unenhanced image frames, and determine the first probability prediction result corresponding to the unenhanced image frame as the final prediction result of the unenhanced image frame.

[0149] Based on the final prediction results of the enhanced image frames and the final prediction results of the unenhanced image frames, the target prediction result for each current image frame in all current capsule endoscopy images is determined.

[0150] Optionally, if the predicted probability of the enhanced image frame exceeds a set threshold, the classification result is considered sufficiently reliable, and this result is used as the final result for that frame; otherwise, the first prediction result is used as the final result. Finally, the prediction result of the optimized data is compared with the data D determined in step 101. pass The results are merged to form the final classification result.

[0151] This invention reduces WCE (Wide-Cut Image Detection) reading time and workload by automatically classifying the stomach, small intestine, and large intestine, while accurately locating target scenes. This is significant for deep learning-based automatic image detection algorithms, helping to improve detection accuracy and efficiency. This invention focuses on classifying scenes such as the stomach, duodenum, jejunum, ileum, and large intestine in capsule endoscopy (CE) image sequences, proposing a framework combining deep learning and hybrid reinforcement learning for efficient and accurate scene classification in CE images. It uses hybrid reinforcement learning, combining Monte Carlo policy gradient (REINFORCE) and action value function (Q-Learning) methods to classify WCE image scenes, rather than using a single reinforcement learning method. This not only compensates for the shortcomings of single reinforcement learning methods and improves classification accuracy, but also automatically distinguishes between correctly classified images and images with uncertain classification results by evaluating the credibility of the classification results through probability thresholding during the automatic image enhancement correction test results stage. For the latter, image enhancement is performed to eliminate potential classification errors. Image enhancement adjusts the enhancement strategy by estimating action value, thereby ensuring optimized image quality. This invention not only optimizes the performance of the classification model using hybrid reinforcement learning, but also improves the ability to handle uncertain classification results using secondary prediction, thus having strong practical application value.

[0152] Figure 2This is the second flowchart of the reinforcement learning-based WCE image scene classification method provided by this invention. For capsule endoscopy image frames, in the first stage of training the reinforcement learning-based WCE image scene classification network, deformation preprocessing is performed on all case WCE image frames. A CNN network extracts WCE image frame features, and a BiRNN network predicts the probability of each frame. The training model parameters are optimized by combining the Q-Learning algorithm and policy gradient. In the second stage, automatically adjusting image enhancement and correcting test results, the trained model is used to test each image frame and determine the reliability of the classification result. For images with uncertain classification, the Q-Learning algorithm is used to select the optimal action, and the optimal action is used for image enhancement. The enhanced image is then fed back into the model for prediction. The first and second prediction results are combined to obtain the final result, which is then output as the final classification result.

[0153] Figure 3 This is a training network structure diagram provided by the present invention. The training method employs an innovative end-to-end framework that organically combines feature extraction networks and temporal networks from deep learning with reinforcement learning methods. This fully leverages the potential of spatiotemporal features and policy optimization, providing an effective solution for WCE image frame classification tasks. The training network structure is as follows: Figure 3 As shown. First, the input image sequence V is processed by a Convolutional Neural Network (CNN). i The image is processed to extract features. Then, a Bidirectional Recurrent Neural Network (BiRNN) is used to process the temporal dependencies between frames, capturing dynamic information and temporal features in the video data and outputting the predicted probability p for each frame. t The reinforcement learning module selects actions and calculates corresponding rewards for frame categories based on a multinomial distribution, and uses the Q-Learning algorithm to calculate the Q-value function to select the optimal action. The selected optimal action is used to update the policy and loss function. By jointly optimizing the objective functions of supervised learning and reinforcement learning, the overall scene classification performance of WCE image frames is improved.

[0154] Figure 4 This is the test network structure diagram provided by the present invention. When loading the trained model for classification prediction, in cases of inaccurate classification, image augmentation is required for the misclassified image frames to improve the final classification performance. The process is as follows: Figure 4As shown, the trained model first predicts the test image sequence and, based on a preset probability threshold, classifies the prediction results into images with a definite classification (i.e., images with a predicted probability exceeding the threshold) and images with an uncertain classification (i.e., images with a predicted probability below the threshold). For image frames with uncertain classification, the Q-Learning method is used to select the optimal image enhancement strategy, which includes rotating the image by 90°, rotating the image by 180°, and scaling the image to 75%. Through multiple rounds of Q-value calculation, the action corresponding to the maximum Q-value is selected, and the image is enhanced based on this action. The enhanced image is then input into the model for secondary classification to obtain the second round of prediction results. If the predicted probability of the enhanced image exceeds the set threshold, the enhanced image classification is considered reliable enough and can be used as the final classification result for that frame; otherwise, the original classification result is retained. The final classification result is generated by combining the prediction results of the first and second classifications, thereby further optimizing the overall classification performance while ensuring the accuracy of the original test image classification results.

[0155] Figure 8 This is the third flowchart of the reinforcement learning-based WCE image scene classification method provided by this invention. The wireless capsule endoscopy image scene classification method uses a combination of deep learning and reinforcement learning to achieve scene classification of wireless capsule endoscopy images. The specific process is as follows: Figure 8 As shown, the specific process is as follows:

[0156] Step 1: Input capsule endoscopy image sequence frames;

[0157] Step 2: Deformation processing of capsule endoscopy images;

[0158] Step 3: For each case, the image frames are used to extract features using a pre-trained CNN network, and the features are saved in HDF5 format, including relevant case information such as patient name, total number of frames, and classification label.

[0159] Step 4: Input the feature sequence of the image frames of the case into a bidirectional recurrent neural network (BiRNN), and the BiRNN network outputs the predicted probability distribution of each frame;

[0160] Step 5: The classification probability distribution of each frame output by BiRNN uses a multinomial distribution for action sampling, and designs reasonable rewards and states. The Q-Learning algorithm is used to select the reward and sampled action with the maximum Q value in N rounds of sampling to calculate the policy gradient and cross-entropy loss function. Finally, the model is trained by optimizing the model parameters θ.

[0161] Step 6: Input the test image sequence into the model trained in Step 5, perform prediction, and separate the image set with a confirmed classification and the image set with an uncertain classification;

[0162] Step 7: For the image set with uncertain classification, use the Q-Learning algorithm to select the optimal action. After performing image enhancement on the selected optimal action, input the image into the trained model to predict the result again. Compare the prediction result with a threshold. If the prediction probability exceeds the set threshold, use the second prediction result as the final classification result for that frame; otherwise, retain the first classification result. This yields the final prediction result for the image with uncertain classification.

[0163] Step 8: Combine the classification results from Step 6 with the prediction results from Step 7 to obtain the final result.

[0164] Figure 9 This is a schematic diagram of the WCE image scene classification device based on reinforcement learning provided by the present invention. The WCE image scene classification device based on reinforcement learning includes an input unit 1. The input unit 1 is used to extract features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence of each current image frame. The feature sequences of all current image frames are input to a preset recurrent neural network model to obtain a first probability prediction result of each current image frame output by the preset recurrent neural network model. The working principle of the input unit 1 can be referred to the aforementioned step 101, and will not be repeated here.

[0165] The reinforcement learning-based WCE image scene classification device further includes an image processing unit 2. The image processing unit 2 is used to determine all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames. For any secondary classification image frame, image processing is performed on the secondary classification image frame to obtain a processed image frame. The target action is selected using the Q-Learning algorithm, and image enhancement is performed on the processed image frame according to the target action to obtain an enhanced image frame. The working principle of the determination unit 2 can be referred to the aforementioned step 102, and will not be repeated here.

[0166] The reinforcement learning-based WCE image scene classification device further includes an output unit 3, which is used to input the enhanced image frame into the preset recurrent neural network model to obtain the second probability prediction result of each enhanced image frame output by the preset recurrent neural network model. The working principle of the output unit 3 can be referred to the aforementioned step 103, and will not be repeated here.

[0167] The reinforcement learning-based WCE image scene classification device further includes a determination unit 4, which is used to determine the target prediction result of each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result. The working principle of the determination unit 4 can be referred to the aforementioned step 104, and will not be repeated here.

[0168] The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0169] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. For example... Figure 10As shown, the electronic device may include a processor 110, a communication interface 120, a memory 130, and a communication bus 140, wherein the processor 110, the communication interface 120, and the memory 130 communicate with each other via the communication bus 140. The processor 110 can call logical instructions in the memory 130 to execute a reinforcement learning-based WCE image scene classification method. This method includes: extracting features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence for each current image frame; inputting all current image frame feature sequences to a preset recurrent neural network model to obtain a first probability prediction result for each current image frame output by the preset recurrent neural network model; identifying all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames; for any secondary classification image frame, performing image processing on the secondary classification image frame to obtain a processed image frame; selecting a target action using a Q-Learning algorithm; and classifying the processed image frame according to the target action. Image enhancement is performed on each frame to obtain an enhanced image frame. The enhanced image frame is then input into the preset recurrent neural network model to obtain a second probability prediction result for each enhanced image frame output by the preset recurrent neural network model. Based on the first probability prediction result and the second probability prediction result, a target prediction result is determined for each current image frame in all current capsule endoscopy images. The preset recurrent neural network model is determined by inputting the feature sequence of sample image frames into an initial recurrent neural network model to obtain sample probability prediction results, sampling sample actions using a multinomial distribution, continuously updating the first Q value using sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0170] Furthermore, the logical instructions in the aforementioned memory 130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0171] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a WCE image scene classification method based on reinforcement learning provided by the above methods. The method includes: extracting features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence for each current image frame; inputting all current image frame feature sequences into a preset recurrent neural network model to obtain a first probability prediction result for each current image frame output by the preset recurrent neural network model; determining all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames; and for any secondary classification image frame, performing image processing on the secondary classification image frame to obtain a processed image frame, and then using... The Q-Learning algorithm selects the target action, and performs image enhancement on the processed image frame based on the target action to obtain the enhanced image frame. The enhanced image frame is input to the preset recurrent neural network model to obtain the second probability prediction result of each enhanced image frame output by the preset recurrent neural network model. The target prediction result of each current image frame in all current capsule endoscopy images is determined based on the first probability prediction result and the second probability prediction result. The preset recurrent neural network model is determined by inputting the feature sequence of sample image frames to the initial recurrent neural network model to obtain the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0172] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the reinforcement learning-based WCE image scene classification method provided by the above methods. This method includes: extracting features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence for each current image frame; inputting all current image frame feature sequences into a preset recurrent neural network model to obtain a first probability prediction result for each current image frame output by the preset recurrent neural network model; determining all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames; for any secondary classification image frame, performing image processing on the secondary classification image frame to obtain a processed image frame; and selecting the target using a Q-Learning algorithm. The action involves enhancing the processed image frame according to the target action to obtain an enhanced image frame; inputting the enhanced image frame into the preset recurrent neural network model to obtain a second probability prediction result for each enhanced image frame output by the preset recurrent neural network model; determining the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result; the preset recurrent neural network model is determined by inputting the feature sequence of sample image frames into an initial recurrent neural network model to obtain sample probability prediction results, sampling sample actions using a multinomial distribution, continuously updating the first Q value using sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model.

[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0174] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A WCE image scene classification method based on reinforcement learning, characterized in that, include: The features of each current image frame in all current capsule endoscopy images are extracted using a pre-defined convolutional neural network model to obtain the feature sequence of each current image frame. The feature sequences of all current image frames are then input into a pre-defined recurrent neural network model to obtain the first probability prediction result of each current image frame output by the pre-defined recurrent neural network model. All current image frames whose first probability prediction result is less than a preset threshold are identified as secondary classification image frames. For any secondary classification image frame, image processing is performed on the secondary classification image frame to obtain a processed image frame. The target action is selected using the Q-Learning algorithm. Based on the target action, the processed image frame is enhanced to obtain an enhanced image frame. The enhanced image frame is input into the preset recurrent neural network model to obtain the second probability prediction result of each enhanced image frame output by the preset recurrent neural network model; Based on the first probability prediction result and the second probability prediction result, determine the target prediction result for each current image frame in all current capsule endoscopy images; The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model. The step of determining the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result includes: For any enhanced image frame, if the probability prediction result is greater than the preset threshold, the second probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame; otherwise, the first probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame. Traverse all enhanced image frames, determine the final prediction result corresponding to each enhanced image frame, determine all current image frames whose first probability prediction result is greater than or equal to the preset threshold as unenhanced image frames, and determine the first probability prediction result corresponding to the unenhanced image frame as the final prediction result of the unenhanced image frame. Based on the final prediction results of the enhanced image frames and the final prediction results of the unenhanced image frames, the target prediction result for each current image frame in all current capsule endoscopy images is determined.

2. The WCE image scene classification method based on reinforcement learning according to claim 1, characterized in that, Before performing feature extraction on each current image frame in all current capsule endoscopy images using a pre-defined convolutional neural network model, the method further includes: Acquire raw capsule endoscopy images using capsule endoscopy; Each original capsule endoscopy image is stretched and compressed at the edges to obtain all current capsule endoscopy images.

3. The WCE image scene classification method based on reinforcement learning according to claim 1, characterized in that, Before inputting all current image frame feature sequences into a preset recurrent neural network model, the method further includes: The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample image frame feature sequences. The sample image frame feature sequences are then input into the initial recurrent neural network model to obtain sample probability prediction results. Sample actions are sampled using a multinomial distribution to obtain the sample results corresponding to the sample actions. The current reward value is obtained by using the sample label results and the sample results corresponding to the sample actions. The current state is determined based on the current reward value and the reward value before sampling. The first Q value is continuously updated using the current state. The expected cumulative return is determined by using the action and reward value corresponding to the maximum first Q value. The gradient of the expected cumulative return is calculated, and the model parameters corresponding to the initial recurrent neural network model are updated using a preset formula to obtain the target model parameters. The preset recurrent neural network model is then determined based on the target model parameters.

4. The WCE image scene classification method based on reinforcement learning according to claim 3, characterized in that, The step of extracting features from all sample image frames using a pre-defined convolutional neural network model to obtain a sample image frame feature sequence includes: The pre-defined convolutional neural network model is used to extract features from all sample image frames to obtain sample extracted features. For each sample image frame, it is stored in HDF5 format, and the patient name, total number of frames, and sample label results of the case corresponding to the extracted features of the sample are recorded. Among them, the feature sequence of the sample image frame is ,in, Let V represent the feature vector of frame t, where T is the total number of frames and V represents the case.

5. The WCE image scene classification method based on reinforcement learning according to claim 3, characterized in that, The step of updating the model parameters corresponding to the initial recurrent neural network model using a preset formula to obtain the target model parameters includes: ; Where α is the learning rate. and These are hyperparameters for balancing the weights. Let cross-entropy be the loss function. Here, J is the regularization term, and J is the expected cumulative return. Indicates to The gradient is obtained by taking the global derivative of the intermediate parameter θ.

6. The WCE image scene classification method based on reinforcement learning according to claim 1, characterized in that, The first probability prediction result is the predicted probability of different scene categories, and the scene categories include stomach, duodenum, jejunum, ileum and large intestine; The step of determining all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames includes: For any current image frame, if the predicted probability of each scene classification in the first probability prediction result is less than the preset threshold, the current image frame is determined to be a secondary classification image frame. Iterate through all current image frames until all secondary classification image frames are determined.

7. The WCE image scene classification method based on reinforcement learning according to claim 1, characterized in that, The image processing of the secondary classification image frame includes rotating the image and scaling the image; The process of selecting a target action using the Q-Learning algorithm and then enhancing the processed image frame based on the target action to obtain an enhanced image frame includes: For any processed image frame, feature extraction is performed on the processed image frame using a preset convolutional neural network model to obtain the processed image features, and the first standard deviation of the current processed image frame is calculated. Calculate the second standard deviation of the secondary classification image frame corresponding to the current processed image frame, and determine the target reward value based on the first standard deviation and the second standard deviation; The second Q value is updated according to the target reward value, and the action corresponding to the maximum second Q value is used to process the secondary classification image frame to obtain the enhanced image frame. Traverse all processed image frames to obtain all enhanced image frames.

8. A WCE image scene classification device based on reinforcement learning, characterized in that, include: The input unit is used to extract features from each current image frame in all current capsule endoscopy images using a preset convolutional neural network model to obtain a feature sequence of each current image frame, and input all current image frame feature sequences into a preset recurrent neural network model to obtain a first probability prediction result of each current image frame output by the preset recurrent neural network model. The image processing unit is used to determine all current image frames whose first probability prediction result is less than a preset threshold as secondary classification image frames. For any secondary classification image frame, the image processing unit performs image processing on the secondary classification image frame to obtain a processed image frame. The target action is selected using the Q-Learning algorithm. The image enhancement is performed on the processed image frame according to the target action to obtain an enhanced image frame. An output unit is used to input the enhanced image frame into the preset recurrent neural network model to obtain a second probability prediction result for each enhanced image frame output by the preset recurrent neural network model. A determining unit is configured to determine the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result. The preset recurrent neural network model is determined by inputting the feature sequence of the input sample image frame to the initial recurrent neural network model, obtaining the sample probability prediction result, sampling sample actions using a multinomial distribution, continuously updating the first Q value using the sample actions and states, obtaining the maximum expected cumulative reward and cross-entropy loss using the sample action and reward value corresponding to the maximum first Q value, and updating the model parameters corresponding to the initial recurrent neural network model. The step of determining the target prediction result for each current image frame in all current capsule endoscopy images based on the first probability prediction result and the second probability prediction result includes: For any enhanced image frame, if the probability prediction result is greater than the preset threshold, the second probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame; otherwise, the first probability prediction result corresponding to the enhanced image frame is determined as the final prediction result of the enhanced image frame. Traverse all enhanced image frames, determine the final prediction result corresponding to each enhanced image frame, determine all current image frames whose first probability prediction result is greater than or equal to the preset threshold as unenhanced image frames, and determine the first probability prediction result corresponding to the unenhanced image frame as the final prediction result of the unenhanced image frame. Based on the final prediction results of the enhanced image frames and the final prediction results of the unenhanced image frames, the target prediction result for each current image frame in all current capsule endoscopy images is determined.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the WCE image scene classification method based on reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Model training method, related equipment, storage medium and computer product

    CN116994019A

  • Threshold estimation sample enhancement method for hyperspectral remote sensing image deep learning classification

    CN118470469A