Method and apparatus for processing cystoscopic image, and computer device, readable storage medium and program product
Patent Information
- Application Number
- PCT/CN2026/085982
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026085982_01102026_PF_FP_ABST
Abstract
Description
Cystoscopic image processing methods, apparatus, computer equipment, readable storage media, and program products
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese patent application No. 2025103662721, filed on March 25, 2025, entitled "Cystoscopy Image Processing Method, Apparatus, Computer Equipment, Readable Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of image processing technology, and in particular to a cystoscopy image processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology
[0004] Bladder cancer is a common malignant tumor of the urinary system, and its early diagnosis usually relies on endoscopic examination. However, traditional methods of bladder cancer diagnosis often depend on the doctor's experience in observing endoscopic images, which has the problems of high subjectivity and high false negative rate.
[0005] With the development of related technologies, artificial intelligence and image processing technologies are gradually being applied to the auxiliary diagnosis of bladder cancer. However, the medical image analysis models used in these technologies are usually computationally complex, making it difficult to perform real-time detection in embedded devices within disposable endoscopes. Summary of the Invention
[0006] Therefore, it is necessary to provide a cystoscopy image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product to address the aforementioned technical problems.
[0007] In a first aspect, this application provides a cystoscopy image processing method, including:
[0008] Acquire cystoscopy images;
[0009] The target objects in the cystoscopy image are detected using a target detection model, and the target object detection results of the cystoscopy image are obtained.
[0010] The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. The first features are fused to obtain second features. The second features are residually concatenated with the input features to obtain output features.
[0011] In one embodiment, the step of using a target detection model to detect target objects in the cystoscopy image and obtaining a target object detection result for the cystoscopy image includes: extracting the brightness layer of the cystoscopy image; subtracting the brightness layer from the cystoscopy image to obtain a detail layer of the cystoscopy image; performing weighted fusion on the brightness layer and the detail layer to obtain a processed cystoscopy image; and using the target detection model to detect target objects in the processed cystoscopy image to obtain a target object detection result.
[0012] In one embodiment, extracting the brightness layer of the cystoscopy image includes: using the cystoscopy image as a guide image, performing guided filtering on the cystoscopy image to obtain the brightness layer.
[0013] In one embodiment, the target detection model is trained through the following steps: acquiring a first image sample containing the target object and a second image sample not containing the target object; segmenting the first image sample to obtain a target segmentation image corresponding to the target object; fusing the target segmentation image into the second image sample to obtain a third image sample; and training the target detection model to be trained using a sample set including the third image sample to obtain the target detection model.
[0014] In one embodiment, fusing the target segmented image into the second image sample to obtain a third image sample includes: performing edge blurring processing on the target segmented image to obtain a first target image; performing color migration processing on the first target image according to the second image sample to obtain a second target image; and inserting the second target image into the second image sample to obtain the third image sample.
[0015] In one embodiment, inserting the second target image into the second image sample to obtain the third image sample includes: scaling the second target image according to random size parameters to obtain a scaled second target image; determining the target insertion position in the second image sample according to random position parameters; and inserting the scaled second target image into the target insertion position to obtain the third image sample.
[0016] Secondly, this application also provides a cystoscopy image processing device, comprising:
[0017] The image acquisition module is used to acquire cystoscopy images;
[0018] The target detection module is used to detect target objects in the cystoscopy image using a target detection model, and obtain the target object detection result of the cystoscopy image;
[0019] The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module uses multiple convolution kernels of different scales to perform channel-wise convolution on the input features to obtain the first feature output by each convolution kernel. The first features are fused to obtain the second feature. The second feature is residually connected with the input features to obtain the output feature.
[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0021] Acquire cystoscopy images;
[0022] The target objects in the cystoscopy image are detected using a target detection model, and the target object detection results of the cystoscopy image are obtained.
[0023] The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. The first features are fused to obtain second features. The second features are residually concatenated with the input features to obtain output features.
[0024] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0025] Acquire cystoscopy images;
[0026] The target objects in the cystoscopy image are detected using a target detection model, and the target object detection results of the cystoscopy image are obtained.
[0027] The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. The first features are fused to obtain second features. The second features are residually concatenated with the input features to obtain output features.
[0028] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0029] Acquire cystoscopy images;
[0030] The target objects in the cystoscopy image are detected using a target detection model, and the target object detection results of the cystoscopy image are obtained.
[0031] The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. The first features are fused to obtain second features. The second features are residually concatenated with the input features to obtain output features.
[0032] The aforementioned cystoscopy image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product first acquire a cystoscopy image, then use a target detection model to detect target objects in the cystoscopy image, and obtain the target object detection result of the cystoscopy image; wherein, the target detection model is a model built based on the YOLOv11 network, the target detection model uses a multi-scale convolution module to replace the c3k2 module in the YOLOv11 network, and the output module of the target detection model includes a micro-target detection head; the multi-scale convolution module uses multiple convolution kernels of different scales to perform channel-wise convolution on the input features to obtain the first feature output by each convolution kernel, fuses the first features to obtain the second feature, and performs residual connection between the second feature and the input features to obtain the output feature. This method utilizes a target detection model built on a YOLOv11 network to detect targets in cystoscopy images. It replaces the C3K2 module in the YOLOv11 network with a multi-scale convolutional module, enabling parallel channel-wise convolution processing of input features using multiple convolutional kernels of different scales. Each kernel can extract features from different receptive fields, effectively capturing features of different sizes in the image while maintaining computational efficiency. By fusing the first features output from each convolutional kernel and then performing residual concatenation with the input features, output features containing multi-scale information and avoiding information loss are obtained. Therefore, the multi-scale convolutional module effectively reduces the computational cost and parameter count of the target detection model, improving the accuracy and efficiency of target object detection in cystoscopy images. Furthermore, introducing a micro-target detection head enhances the model's ability to detect small targets, improving the accuracy of detecting small lesions in cystoscopy images. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 is a flowchart illustrating a cystoscopy image processing method in one embodiment;
[0035] Figure 2 is a flowchart illustrating the processing of input features by the multi-scale convolution module in one embodiment;
[0036] Figure 3 is a schematic diagram of the process of using a target detection model to detect target objects in a cystoscopy image in one embodiment;
[0037] Figure 4 is a schematic diagram comparing the cystoscopy image before and after processing in one embodiment;
[0038] Figure 5 is a flowchart illustrating the training steps of an object detection model in one embodiment;
[0039] Figure 6 is a schematic diagram of the process of obtaining the third image sample in one embodiment;
[0040] Figure 7 is a schematic diagram of the process for obtaining the third image sample in another embodiment;
[0041] Figure 8 is a structural block diagram of a cystoscope image processing device in one embodiment;
[0042] Figure 9 is an internal structure diagram of a computer device in one embodiment. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0044] In one embodiment, as shown in Figure 1, a cystoscopy image processing method is provided. This embodiment illustrates the method applied to a terminal, but it is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. Exemplarily, the terminal can be an embedded device of a disposable endoscope. In this embodiment, the method includes the following steps:
[0045] Step S101: Obtain cystoscopy images.
[0046] Step S102: Target objects in the cystoscopy image are detected using a target detection model to obtain the target object detection result of the cystoscopy image. The target detection model is a model built based on the YOLOv11 network. The target detection model replaces the c3k2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head; the multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain the first feature corresponding to each scale, fuses the first features to obtain the second feature, and performs a residual connection between the second feature and the input features to obtain the output feature.
[0047] The cystoscopy images can be images acquired using a cystoscope (hereinafter referred to as "cystoscope"). For example, in step S101, cystoscopy images can be acquired using a disposable cystoscope with the patient's authorization.
[0048] In step S102, a target detection model can be used to detect target objects in the cystoscopy image to obtain corresponding target object detection results. The target object can be a tumor lesion site of bladder cancer. The target object detection results can indicate whether the cystoscopy image contains a target object, or they can indicate the location information of the target object in the cystoscopy image. For example, when the cystoscopy image contains a lesion site, the target object detection result output by the target detection model can include the location information of the lesion site in the cystoscopy image, such as the coordinate information of the target bounding box of the lesion site; while when the cystoscopy image does not contain a lesion site, the target object detection result can include information indicating that no target object was detected. Optionally, in step S102, the cystoscopy image can be directly input into the target detection model for target object detection, or the cystoscopy image can be preprocessed before being input into the target detection model.
[0049] Specifically, the target detection model used in step S102 is a model built based on the YOLOv11 network.
[0050] In traditional convolutional processing, the convolutional kernel needs to operate on each channel of the input features, often involving a large amount of computation, especially when the number of channels is large, the computational load increases dramatically. Furthermore, traditional convolutional processing typically uses a single-size convolutional kernel, thus only capturing features of a fixed size and failing to comprehensively capture features of objects of different sizes and scales in the image. Therefore, the object detection model in this application replaces the c3k2 module in the YOLOv11 network with a multi-scale convolutional module.
[0051] For example, the multi-scale convolution module in this application can process the input features as shown in Figure 2. The multi-scale convolution module can first perform point convolution processing on the input features (which can be of size H×W×C, where H is the height, W is the width, and C is the number of channels) to expand the number of channels of the input features (as shown in Figure 2, this could be expanding the number of channels of the input features from C to 2C). Then, it sequentially uses the batch normalization (BN) function and the ReLU6 activation function to process the expanded input features, obtaining the processed input features. The multi-scale convolution module includes multiple convolution kernels corresponding to different scales (as shown in Figure 2, these kernels can include various sizes such as p*p, q*q, ..., s*s, where p, q, ..., s can represent the width and height of the corresponding convolution kernels, respectively). Each scale convolution kernel performs channel-wise convolution processing on the processed input features and outputs the corresponding first feature. The first features corresponding to each scale are processed sequentially by batch normalization (BN) and ReLU6 activation function. Then, the first features corresponding to each scale can be fused to obtain the second feature (not shown in the figure) by sequentially performing channel concatenation and channel shuffle. The second feature is then connected to the processed input features through a first residual connection process (e.g., element-wise addition) to obtain the third feature. The number of channels is then adjusted through point convolution processing, and the adjusted third feature is connected to the input features through a second residual connection process (e.g., element-wise addition) to obtain the output feature.
[0052] Traditional YOLOv11 networks include large, medium, and small target detection heads, used to detect targets of different sizes in images. However, because bladder cancer often presents with small lesions, the target detection model in this application, designed for the auxiliary diagnosis of bladder cancer, adds a micro-target detection head to the output module of the YOLOv11 network. For example, the micro-target detection head can sample the input image with a smaller stride than the small target detection head in a traditional YOLOv11 network, and detect targets based on the obtained feature map. For instance, the sampling stride of the micro-target detection head can be 4, so for an input image of size 640*640, the micro-target detection head can sample it to obtain a feature map of size 160*160, and then detect targets based on this feature map. Optionally, the target detection model in this application can also remove the large target detection head from the output module. Thus, the output module can combine the outputs of the micro-target detection head and other target detection heads to obtain the target detection results of the targets in the cystoscopy image.
[0053] In the aforementioned cystoscopy image processing method, a target detection model based on a YOLOv11 network is used to detect targets in cystoscopy images. The multi-scale convolution module replaces the C3K2 module in the YOLOv11 network, enabling parallel channel-wise convolution processing of input features using multiple convolution kernels of different scales. Each convolution kernel can extract features from different receptive fields, thus effectively capturing features of different sizes in the image while maintaining computational efficiency. By fusing the first features output from each convolution kernel and then performing residual connections with the input features, output features containing multi-scale information and avoiding information loss can be obtained. Therefore, using the multi-scale convolution module effectively reduces the computational load and parameter count of the target detection model, and improves the detection accuracy and efficiency of target objects in cystoscopy images. Furthermore, introducing a micro-target detection head into the model enhances its ability to detect small targets, which is beneficial for improving the detection accuracy of small lesions in cystoscopy images.
[0054] In target object detection using cystoscopy images, bladder cancer lesions are often flat or non-protruding, making them easily overlooked under white light. Therefore, this application preprocesses the cystoscopy images before inputting them into the target detection model to improve the model's accuracy in detecting target objects.
[0055] In an exemplary embodiment, as shown in Figure 3, the target object detection model is used to detect the target object in the cystoscopy image, and the target object detection result of the cystoscopy image is obtained, which may include:
[0056] Step S301: Extract the brightness layer of the cystoscopy image.
[0057] Specifically, for cystoscopy images, edge-preserving filtering can be used to extract the brightness layer of the cystoscopy image while preserving edge information.
[0058] In one exemplary embodiment, extracting the brightness layer of a cystoscopy image may include: using the cystoscopy image as a guide image, performing guided filtering on the cystoscopy image to obtain the brightness layer.
[0059] Specifically, this application uses a guided filtering method to process cystoscopy images to extract the luminance layer. For example, assuming the cystoscopy image is an RGB three-channel image, then for each channel image I... c (x,y) can be used as a guide map for guided filtering to obtain the brightness layer L corresponding to each channel image. c (x,y)=fguidfilter(I c(x,y)). In the formula, c represents the image channel (e.g., R, G, B), and f_uidfilter is the guided filter function.
[0060] Step S302: Subtract the brightness layer from the cystoscopy image to obtain the detail layer of the cystoscopy image.
[0061] Specifically, based on the brightness layer extracted in step S301, it can be multiplied by a stretching factor to calculate the stretched brightness layer. Then, the stretched brightness layer can be subtracted from the cystoscopy image to obtain the detail layer. For example, this process can be shown in the following formula: D c (x,y)=I c (x,y)-β*L c (x,y)
[0062] In the formula, c represents the image channel (e.g., R, G, B), and D... c (x,y) represents the detail layer corresponding to channel c, I c (x,y) is the channel image corresponding to channel c, L c (x,y) represents the brightness layer corresponding to channel c, β is the stretching factor, and β*L c (x,y) represents the stretched brightness layer. The stretching coefficient β can be a preset value determined experimentally.
[0063] Step S303: Weighted fusion of the brightness layer and detail layer is performed to obtain the processed cystoscopy image.
[0064] In this step, based on the brightness and detail layers of the cystoscopy image, a weighted fusion can be performed to enhance the cystoscopy image. For example, the weighted fusion process can be illustrated as follows: E c (x,y)=α*D c (x,y)+β*L c (x,y)
[0065] In the formula, c represents the image channel (e.g., R, G, B), E c (x,y) corresponds to the fused image of channel c, D c (x,y) represents the detail layer corresponding to channel c, α is the gain coefficient, and L c (x,y) represents the luminance layer corresponding to channel c, and β is the stretching coefficient. The gain coefficient α can be a preset value determined experimentally.
[0066] After obtaining the fused images corresponding to each image channel, they can be combined to obtain a processed cystoscopy image. For example, please refer to Figure 4, where part (a) of Figure 4 is the original cystoscopy image and part (b) of Figure 4 is the processed cystoscopy image. The details in the processed cystoscopy image are significantly enhanced.
[0067] Step S304: Use the target detection model to detect the target objects in the processed cystoscopy image and obtain the target object detection results.
[0068] In this step, the processed cystoscopy image can be input into the target detection model to obtain the target object detection result of the cystoscopy image output by the target detection model.
[0069] In this embodiment, the brightness layer and detail layer of the cystoscopy image are extracted by using the edge-preserving filtering method, and then the brightness layer and detail layer are weighted and fused to obtain the processed cystoscopy image. This can enhance the detail contrast in the cystoscopy image, which is beneficial to highlighting the position of the target object in the image and improving the accuracy of subsequent target detection using the target detection model.
[0070] In related technologies, cystoscopy images containing bladder cancer lesions are typically used to train detection models. However, due to the scarcity of clinical data on bladder cancer, it is usually difficult to obtain a large number of cystoscopy images containing bladder cancer lesions, which can easily affect the training effect of the model. Furthermore, cystoscopy images are easily affected by various factors such as lighting, angle, and mucus, thus placing high demands on the model's generalization ability. Therefore, this application expands the training samples of the target detection model to increase the amount and diversity of training samples, thereby improving the detection accuracy of the target detection model.
[0071] In an exemplary embodiment, as shown in Figure 5, the object detection model can be trained through the following steps:
[0072] Step S501: Obtain a first image sample containing the target object and a second image sample not containing the target object.
[0073] The first image sample can be a cystoscopy image containing the target object (i.e., the tumor lesion site of bladder cancer), and the second image sample can be a cystoscopy image not containing the target object. For example, in this step, multiple first image samples and multiple second image samples can be acquired separately with the patient's authorization. The multiple second image samples can include cystoscopy images acquired under different lighting, angles, mucus, and other factors.
[0074] Step S502: Segment the first image sample to obtain the target segmentation image corresponding to the target object.
[0075] Specifically, the first image sample can be segmented to extract a target segmented image containing the target object. For example, the target object can be pre-segmented and labeled in the first image sample to obtain a mask region corresponding to the target object, and then the mask region can be used to segment the target segmented image corresponding to the target object from the first image sample.
[0076] Step S503: The target segmented image is fused into the second image sample to obtain the third image sample.
[0077] Specifically, by fusing the target segmented image into the second image sample, a third image sample containing the target object can be obtained. For example, the third image sample can be obtained by inserting the target segmented image into the second image sample.
[0078] Optionally, the same target segmentation image can be fused into multiple different second image samples to obtain multiple third image samples, or different target segmentation images can be fused into the same second image sample to obtain multiple third image samples.
[0079] Step S504: Train the target detection model to be trained using a sample set including the third image sample to obtain the target detection model.
[0080] After obtaining the third image sample, it can be added to the sample set used for model training. The target detection model to be trained is then trained using samples from this sample set to obtain the target detection model. For example, the sample set may include multiple third image samples, or it may also include other image samples.
[0081] In this embodiment, for training the object detection model, the target segmentation image corresponding to the target object is segmented from the first image sample containing the target object, and then fused into the second image sample that does not contain the target object. This synthesis method can obtain a third image sample containing the target object, which can effectively increase the number of samples used for model training and construct a diverse dataset. Then, the target detection model to be trained is trained using the sample set containing the third image sample, which can enhance the generalization and accuracy of the target detection model and effectively solve the problems of scarce and insufficient training data.
[0082] In an exemplary embodiment, as shown in FIG6, fusing the target segmented image into a second image sample to obtain a third image sample may include:
[0083] Step S601: Perform edge blurring processing on the target segmentation image to obtain the first target image.
[0084] In this step, the target segmentation image can be blurred using methods such as Gaussian filtering to obtain the first target image.
[0085] Step S602: Based on the second image sample, perform color migration processing on the first target image to obtain the second target image.
[0086] In this step, the second image sample can be used as a reference image to perform color transfer processing on the first target image in order to obtain a second target image whose color matches that of the second image sample.
[0087] For example, in this step, the Reinhard color transfer algorithm can be used to perform color transfer processing on the first target image. Specifically, the second image sample and the first target image can first be converted to the Lab color space, and the mean μ of the second image sample S can be calculated. S and standard deviation σ S And calculate the mean μ of the first target image T. T and standard deviation σ T Then, for each pixel of the first target image T... The adjusted pixel values can be obtained by performing a linear transformation using the following formula:
[0088] In the formula, The adjusted pixel value for pixel p. The standard deviation of the second image sample in channel L. The standard deviation of the first target image in channel L is given. Let p be the pixel value of the first target image in channel L. Let be the mean value of the first target image in channel L. The mean value of the second image sample in channel L; The standard deviation of the second image sample in channel a. Let be the standard deviation of the first target image in channel a. Let p be the pixel value of the first target image in channel a. Let be the mean value of the first target image in channel a. This represents the mean value of the second image sample in channel a. The standard deviation of the second image sample in channel b. The standard deviation of the first target image in channel b. Let p be the pixel value of the first target image in channel b. Let be the mean value of the first target image in channel b. This represents the mean value of the second image sample in channel b.
[0089] After adjusting all pixels of the first target image using the above method, the adjusted first target image can be converted to RGB space to obtain the second target image.
[0090] Step S603: Insert the second target image into the second image sample to obtain the third image sample.
[0091] In this step, a third image sample containing the target object is obtained by inserting the second target image into the second image sample. The insertion position of the second target image within the second image sample can be specified using a position parameter, and the size of the second target image within the second image sample can be specified using a size parameter.
[0092] It is understood that in some implementations, the first target image can be inserted into the second image sample first, and then the first target image can be color-shifted to obtain a third image sample containing the second target image. For example, referring to Figure 7, based on a first image sample containing the target object (corresponding to the first color) and a second image sample not containing the target object (corresponding to the second color), a target segmentation image can be first segmented from the first image sample, and its edges can be blurred to obtain the first target image. Then, the first target image is inserted as a layer into the second image sample, and the first target image is color-shifted according to the second image sample to obtain a third image sample containing the second target image.
[0093] In this embodiment, by performing edge blurring processing on the target sample image and color transfer processing on the target sample image based on the second image sample, the target object inserted into the second image sample can transition naturally with the second image sample as the background, and the color tone of the target object can be matched with that of the second image sample. This makes the synthesized third image sample closer to the real cystoscopy image. Using this third image sample for model training can increase the accuracy of the target detection model and improve the accuracy of target detection in cystoscopy images.
[0094] In an exemplary embodiment, inserting a second target image into a second image sample to obtain a third image sample may include: scaling the second target image according to random size parameters to obtain a scaled second target image; determining the target insertion position in the second image sample according to random position parameters; and inserting the scaled second target image into the target insertion position to obtain a third image sample.
[0095] Specifically, when inserting the second target image into the second image sample, random size parameters and random position parameters can be obtained first. The random size parameter indicates the target insertion size of the second target image in the second image sample, and the random position parameter indicates the target insertion position of the second target image in the second image sample. Both are randomly generated parameters. Based on the random size parameter, the second target image can be scaled to make its size conform to the target insertion size. Then, the scaled second target image can be inserted into the target insertion position indicated by the random position parameter to obtain the synthesized third image sample.
[0096] In this embodiment, by randomly setting the size and position of the second target image inserted into the second image sample, the data diversity of the third image sample can be improved, which is beneficial to enhancing the generalization and accuracy of the target detection model.
[0097] In one embodiment, a cystoscopy image processing method is provided, specifically including the following steps:
[0098] Step S1: Obtain cystoscopy images.
[0099] Step S2: Perform image preprocessing and enhancement on the cystoscopy image.
[0100] The process involves first using the cystoscopy image itself as a guide image, extracting the brightness layer of the cystoscopy image using a guided filtering method, then multiplying the brightness layer by a stretching factor to obtain a stretched brightness layer, subtracting the stretched brightness layer from the cystoscopy image to obtain a detail layer, multiplying the detail layer by a gain factor, and then adding it to the stretched brightness layer. The processed cystoscopy image is obtained through weighted fusion.
[0101] Step S3: Use the target detection model to detect the target objects in the processed cystoscopy image and obtain the target object detection results.
[0102] In this embodiment, the target detection model is based on the YOLOv11 network. This model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolutional module, and its output module includes a micro-target detection head. The multi-scale convolutional module in the target detection model uses convolutional kernels corresponding to different scales to perform channel-wise convolution on the input features to obtain first features corresponding to each scale. These first features are then fused to obtain second features, and finally, the second features are residually concatenated with the input features to obtain the output features.
[0103] The target detection model in this embodiment can be trained through the following steps:
[0104] Step S4: Obtain a first image sample containing the target object and a second image sample not containing the target object.
[0105] The number of second image samples can be multiple, and these multiple second image samples can include cystoscopy images acquired under different lighting, angles, mucus and other factors.
[0106] Step S5: Segment the first image sample to obtain the target segmentation image corresponding to the target object.
[0107] Specifically, a mask region can be obtained by segmenting and labeling the first image sample, and then the first image sample can be segmented using the mask region to obtain the target segmented image.
[0108] Step S6: Fuse the target segmented image into the second image sample to obtain the third image sample.
[0109] For example, in this step, the target segmented image can be inserted into the second image sample first, and Gaussian filtering can be used to blur the edges of the target segmented image. Then, the Reinhard color transfer algorithm can be used to perform color transfer processing on the target segmented image based on the second image sample so that the hue of the target segmented image matches that of the second image sample.
[0110] For example, in this step, the target segmentation image can first be processed by edge blurring to obtain a first target image, then the first target image can be processed by color shifting according to the second image sample to obtain a second target image, and then the second target image can be inserted into the second image sample to obtain a third image sample.
[0111] In this case, whether the target segmentation image is inserted into the second image sample first, or the target segmentation image is processed first and then the second target image is inserted into the second image sample, the target insertion size and target insertion position can be set randomly, thereby improving the data diversity of the third image sample.
[0112] Step S7: Train the target detection model to be trained using a sample set including the third image sample to obtain the target detection model.
[0113] In this embodiment, by synthesizing data, a third image sample is synthesized using a first image sample containing the target object and second image samples collected under different scenarios. This constructs a diverse dataset, providing more usable data for model training, enhancing the model's generalization ability, and effectively solving the problems of data scarcity and insufficient diversity. Furthermore, by introducing depthwise separable convolutions and multi-scale convolution kernels into the YOLOv11 network and combining them, the computational load is reduced while enhancing the model's feature extraction capabilities. For the application scenario of cystoscopy image detection, the large target detection head in the YOLOv11 network is removed and a micro-target detection head is added, enhancing the target detection model's ability to detect small targets. Simultaneously, before inputting the cystoscopy image into the target detection model, guided filtering is used to extract and enhance details in the image's RGB channels, effectively enhancing the image detail contrast of the cystoscopy image, highlighting the location of the target object, and further improving the target object detection accuracy.
[0114] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0115] Based on the same inventive concept, this application also provides a cystoscopy image processing apparatus for implementing the cystoscopy image processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the cystoscopy image processing apparatus provided below can be found in the limitations of the cystoscopy image processing method described above, and will not be repeated here.
[0116] In an exemplary embodiment, as shown in FIG8, a cystoscopy image processing apparatus 800 is provided, comprising:
[0117] Image acquisition module 801 is used to acquire cystoscopy images.
[0118] The target detection module 802 is used to detect target objects in cystoscopy images using a target detection model, and obtain the target object detection results of the cystoscopy images.
[0119] The object detection model is based on the YOLOv11 network. The object detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the object detection model includes a micro-object detection head and a multi-scale convolution module. The multi-scale convolution module uses multiple convolution kernels of different scales to perform channel-wise convolution on the input features to obtain the first feature output by each convolution kernel. The first features are fused to obtain the second feature. The second feature is residually connected with the input features to obtain the output feature.
[0120] In an exemplary embodiment, the target detection module 802 is configured to: extract the brightness layer of the cystoscopy image; subtract the brightness layer from the cystoscopy image to obtain the detail layer of the cystoscopy image; perform weighted fusion of the brightness layer and the detail layer to obtain a processed cystoscopy image; and use the target detection model to detect target objects in the processed cystoscopy image to obtain the target object detection result.
[0121] In one exemplary embodiment, the brightness extraction module is used to perform guided filtering on the cystoscopy image as a guide image to obtain a brightness layer.
[0122] In an exemplary embodiment, the object detection model is trained through the following steps: acquiring a first image sample containing the target object and a second image sample not containing the target object; segmenting the first image sample to obtain a target segmentation image corresponding to the target object; fusing the target segmentation image into the second image sample to obtain a third image sample; and training the object detection model to be trained using the sample set including the third image sample to obtain the object detection model.
[0123] In an exemplary embodiment, fusing a target segmented image into a second image sample to obtain a third image sample includes: performing edge blurring processing on the target segmented image to obtain a first target image; performing color migration processing on the first target image based on the second image sample to obtain a second target image; and inserting the second target image into the second image sample to obtain a third image sample.
[0124] In an exemplary embodiment, inserting a second target image into a second image sample to obtain a third image sample includes: scaling the second target image according to random size parameters to obtain a scaled second target image; determining the target insertion position in the second image sample according to random position parameters; and inserting the scaled second target image into the target insertion position to obtain a third image sample.
[0125] Each module in the aforementioned cystoscopy image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0126] In an exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram is shown in Figure 9. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a cystoscopy image processing method.
[0127] Those skilled in the art will understand that the structure shown in Figure 9 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.
[0128] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0129] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0130] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0134] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A cystoscope image processing method, characterized by, The method includes: Acquire cystoscopy images; The target objects in the cystoscopy image are detected using a target detection model, and the target object detection results of the cystoscopy image are obtained. The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs channel-wise convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. The first features are fused to obtain second features. The second features are residually concatenated with the input features to obtain output features.
2. The method of claim 1, wherein, The step of using a target detection model to detect target objects in the cystoscopy image and obtaining the target object detection result of the cystoscopy image includes: Extract the brightness layer from the cystoscopy image; The brightness layer is subtracted from the cystoscopy image to obtain the detail layer of the cystoscopy image; The brightness layer and the detail layer are weighted and fused to obtain the processed cystoscopy image; The target detection model is used to detect target objects in the processed cystoscopy image to obtain the target object detection result.
3. The method of claim 2, wherein, The step of extracting the brightness layer of the cystoscopy image includes: Using the cystoscopy image as a guide image, the cystoscopy image is subjected to guided filtering to obtain the brightness layer.
4. The method of claim 1, wherein, The target detection model is trained through the following steps: Obtain a first image sample containing the target object and a second image sample not containing the target object; The first image sample is segmented to obtain a target segmentation image corresponding to the target object; The target segmented image is fused into the second image sample to obtain the third image sample; The target detection model is trained using a sample set including the third image sample to obtain the target detection model.
5. The method of claim 4, wherein, The step of fusing the target segmented image into the second image sample to obtain the third image sample includes: The target segmentation image is subjected to edge blurring processing to obtain the first target image; Based on the second image sample, the first target image is subjected to color migration processing to obtain the second target image; The second target image is inserted into the second image sample to obtain the third image sample.
6. The method of claim 5, wherein, The step of inserting the second target image into the second image sample to obtain the third image sample includes: The second target image is scaled according to random size parameters to obtain a scaled second target image; The target insertion position in the second image sample is determined based on random position parameters, and the scaled second target image is inserted into the target insertion position to obtain the third image sample.
7. A cystoscope image processing apparatus characterized by comprising: The device includes: The image acquisition module is used to acquire cystoscopy images; The target detection module is used to detect target objects in the cystoscopy image using a target detection model, and obtain the target object detection result of the cystoscopy image; The target detection model is a model built on the YOLOv11 network. The target detection model replaces the C3K2 module in the YOLOv11 network with a multi-scale convolution module. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module uses multiple convolution kernels of different scales to perform channel-wise convolution on the input features to obtain the first feature output by each convolution kernel. The first features are fused to obtain the second feature. The second feature is residually connected with the input features to obtain the output feature.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.