Cystoscope image processing method and device, computer equipment, readable storage medium and program product

By introducing a multi-scale convolution module and micro-objective detection head in the yolov11 network, combining image preprocessing and training sample expansion, the subjectivity and computational complexity of traditional bladder cancer diagnosis is solved, and efficient and accurate bladder cancer assisted diagnosis is achieved.

CN120259824APending Publication Date: 2025-07-04GUANGZHOU RED PINE MEDICAL INSTR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510366272.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional bladder cancer diagnosis relies on doctor experience, has strong subjectivity, high missed detection rate, and the existing image processing technology has high computational complexity, making it difficult to detect in endoscopic devices in real time.

Method used

The object detection model based on the yolov11 network is adopted, and the c3k2 module is replaced by a multi-scale convolution module, and a micro-object detection head is introduced, combining image preprocessing and training sample expansion technology to improve detection accuracy and efficiency.

Benefits of technology

While reducing the calculation amount, the detection accuracy and efficiency of the target object in the cystoscopic image are improved, especially the detection ability of small lesions, and the generalization and accuracy of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259824A_ABST
    Figure CN120259824A_ABST
Patent Text Reader

Abstract

The invention relates to a cystoscope image processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a cystoscope image; detecting a target object in the cystoscope image by using a target detection model to obtain a target object detection result of the cystoscope image; wherein the target detection model is a model constructed based on a yolov11 network, the target detection model uses a multi-scale convolution module to replace a c3k2 module in the yolov11 network, and an output module of the target detection model comprises a micro-target detection head; and the multi-scale convolution module performs channel-by-channel convolution on the input features by using a plurality of convolution kernels with different scales to obtain first features output by the convolution kernels, fuses the first features to obtain second features, and performs residual connection on the second features and the input features to obtain output features. By adopting the method, the detection precision and efficiency of the target object in the cystoscope image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method and apparatus for processing cystoscope images, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Bladder cancer is a common malignant tumor in the urinary system, and its early diagnosis usually relies on endoscopic detection. However, the traditional method for diagnosing bladder cancer usually depends on doctors' experience to observe endoscopic images, which has problems of strong subjectivity and high missed detection rate.

[0003] With the development of related technologies, artificial intelligence and image processing technologies have gradually been applied to the auxiliary diagnosis of bladder cancer. However, the medical image analysis models used in related technologies usually have a high computational complexity and are difficult to perform real-time detection in embedded devices of disposable endoscopes. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and apparatus for processing cystoscope images, a computer device, a computer-readable storage medium, and a computer program product for the above technical problems.

[0005] In a first aspect, this application provides a method for processing cystoscope images, including:

[0006] Obtain a cystoscope image;

[0007] Use a target detection model to detect target objects in the cystoscope image, and obtain a target object detection result of the cystoscope image;

[0008] Wherein, the target detection model is a model constructed based on the yolov11 network, the target detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network, and the output module of the target detection model includes a micro target detection head; the multi-scale convolution module performs per-channel convolution on the input feature using convolution kernels corresponding to different scales respectively to obtain first features corresponding to each scale, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0009] In one embodiment, detecting a target object in the cystoscope image by using the target detection model to obtain the target object detection result of the cystoscope image includes: extracting the luminance layer of the cystoscope image; subtracting the luminance layer from the cystoscope image to obtain the detail layer of the cystoscope image; performing weighted fusion on the luminance layer and the detail layer to obtain a processed cystoscope image; and detecting the target object in the processed cystoscope image by using the target detection model to obtain the target object detection result.

[0010] In one embodiment, extracting the luminance layer of the cystoscope image includes: using the cystoscope image as a guidance map and performing guided filtering on the cystoscope image to obtain the luminance layer.

[0011] In one embodiment, the target detection model is trained through the following steps: obtaining a first image sample containing a target object and a second image sample not containing the target object; segmenting the first image sample to obtain a target segmentation image corresponding to the target object; fusing the target segmentation image into the second image sample to obtain a third image sample; and training a target detection model to be trained by using a sample set including the third image sample to obtain the target detection model.

[0012] In one embodiment, fusing the target segmentation image into the second image sample to obtain a third image sample includes: performing edge blurring on the target segmentation image to obtain a first target image; performing color transfer on the first target image according to the second image sample to obtain a second target image; and inserting the second target image into the second image sample to obtain the third image sample.

[0013] In one embodiment, inserting the second target image into the second image sample to obtain the third image sample includes: scaling the second target image according to a random size parameter to obtain a scaled second target image; determining a target insertion position in the second image sample according to a random position parameter, and inserting the scaled second target image into the target insertion position to obtain the third image sample.

[0014] In a second aspect, the present application further provides a cystoscope image processing device, including:

[0015] an image acquisition module, configured to acquire a cystoscope image;

[0016] a target detection module, configured to detect a target object in the cystoscope image by using a target detection model to obtain the target object detection result of the cystoscope image;

[0017] Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head. The multi-scale convolution module performs per-channel convolution on the input feature using convolution kernels of multiple different scales respectively to obtain first features output by each convolution kernel, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0018] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0019] Obtain a cystoscope image;

[0020] Use the object detection model to detect the target object in the cystoscope image to obtain the target object detection result of the cystoscope image;

[0021] Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head. The multi-scale convolution module performs per-channel convolution on the input feature using convolution kernels corresponding to different scales respectively to obtain first features corresponding to each scale, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0022] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0023] Obtain a cystoscope image;

[0024] Use the object detection model to detect the target object in the cystoscope image to obtain the target object detection result of the cystoscope image;

[0025] Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head. The multi-scale convolution module performs per-channel convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0026] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor, implements the following steps:

[0027] Obtain a cystoscope image;

[0028] Use the object detection model to detect the target object in the cystoscope image to obtain the target object detection result of the cystoscope image;

[0029] Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head. The multi-scale convolution module performs per-channel convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0030] The above cystoscope image processing method, device, computer device, computer-readable storage medium, and computer program product first obtain cystoscope images, and then use a target detection model to detect target objects in the cystoscope images to obtain the target object detection results of the cystoscope images. Among them, the target detection model is a model constructed based on the yolov11 network. The target detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the target detection model includes a micro-target detection head. The multi-scale convolution module performs per-channel convolution on the input features using multiple convolution kernels of different scales respectively to obtain the first features output by each convolution kernel, fuses the first features to obtain the second feature, and performs a residual connection between the second feature and the input feature to obtain the output feature. This method uses a target detection model constructed based on the yolov11 network to perform target detection on cystoscope images. Among them, using a multi-scale convolution module to replace the c3k2 module in the yolov11 network can perform parallel per-channel convolution processing on the input features using multiple convolution kernels of different scales. Each convolution kernel can extract features from different receptive field ranges, so that different-sized features in the image can be effectively captured while maintaining computational efficiency. By fusing the first features output by each convolution kernel and then performing a residual connection with the input feature, an output feature containing multi-scale information and avoiding information loss can be obtained. Therefore, using the multi-scale convolution module can effectively reduce the computational amount and the number of parameters of the target detection model, and can improve the detection accuracy and efficiency of target objects in cystoscope images. By introducing a micro-target detection head into the model, the detection ability of the target detection model for small targets can be enhanced, which is beneficial to improving the detection accuracy of small lesion sites in cystoscope images. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a flowchart showing the process of the cystoscope image processing method in an embodiment;

[0033] Figure 2 It is a flowchart showing the process of the multi-scale convolution module processing the input features in an embodiment;

[0034] Figure 3 It is a flowchart showing the process of using the target detection model to detect target objects in cystoscope images in an embodiment;

[0035] Figure 4 Schematic diagram for comparison before and after processing of cystoscope image in one embodiment;

[0036] Figure 5 Schematic flowchart of training steps of target detection model in one embodiment;

[0037] Figure 6 Schematic flowchart of obtaining the third image sample in one embodiment;

[0038] Figure 7 Schematic flowchart of obtaining the third image sample in another embodiment;

[0039] Figure 8 Block diagram of the structure of a cystoscope image processing device in one embodiment;

[0040] Figure 9 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0042] In one embodiment, as Figure 1 shown, a cystoscope image processing method is provided. In this embodiment, it is exemplified that this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Exemplarily, the terminal can be an embedded device of a disposable endoscope. In this embodiment, the method includes the following steps:

[0043] Step S101, obtain a cystoscope image.

[0044] Step S102, use a target detection model to detect target objects in the cystoscope image to obtain a target object detection result of the cystoscope image. Among them, the target detection model is a model constructed based on the yolov11 network. The target detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the target detection model includes a micro target detection head; the multi-scale convolution module performs per-channel convolution on the input features using convolution kernels corresponding to different scales to obtain first features corresponding to each scale, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

[0045] Among them, the cystoscope image can be an image collected by using a cystoscope (hereinafter simply referred to as "cystoscope"). Exemplarily, in step S101, with the patient's authorization, a cystoscope image can be collected by using a disposable cystoscope.

[0046] Among them, in step S102, a target detection model can be used to detect target objects in the cystoscope image to obtain corresponding target object detection results. Among them, the target object can be the tumor lesion site of bladder cancer, and the target object detection result can be used to indicate whether the cystoscope image contains the target object, or can also indicate the position information of the target object in the cystoscope image. Exemplarily, when the cystoscope image contains a lesion site, the target object detection result output by the target detection model can include the position information of the lesion site in the cystoscope image, for example, it can include the coordinate information of the target box of the lesion site; when the cystoscope image does not contain a lesion site, the target object detection result can contain information indicating that no target object is detected. Optionally, in step S102, the cystoscope image can be directly input into the target detection model for target object detection, or the cystoscope image can be preprocessed first and then input into the target detection model.

[0047] Specifically, the target detection model used in step S102 is a model constructed based on the yolov11 network.

[0048] Among them, in traditional convolution processing, the convolution kernel needs to operate on each channel of the input feature, which often involves a large amount of calculation. Especially when the number of channels is large, the amount of calculation increases sharply. At the same time, in traditional convolution processing, a convolution kernel of a single size is usually used for convolution processing, so that only features of a fixed size can be captured, and the features of objects of different sizes and scales in the image cannot be comprehensively captured. Based on this, the target detection model in this application replaces the c3k2 module in the yolov11 network with a multi-scale convolution module.

[0049] Exemplarily, the processing process of the multi-scale convolution module in this application for the input feature can be as Figure 2 shown. Among them, the multi-scale convolution module can first perform point convolution processing on the input feature (the size can be H×W×C, where H is the height, W is the width, and C is the number of channels) to expand the number of channels of the input feature (as Figure 2 shown, it can be to expand the number of channels of the input feature from C to 2C), and then sequentially use the batch normalization function (Batch Normalization, BN) and the ReLU6 activation function to process the expanded input feature to obtain the processed input feature. Among them, the multi-scale convolution module includes convolution kernels corresponding to multiple different scales (as Figure 2 shown, it can include sizes respectively , , ……, a variety of convolutional kernels, where , , ……, can respectively represent the width and height of the corresponding convolutional kernel), where the convolutional kernels of each scale perform per-channel convolutional processing on the processed input features and output the corresponding first features. After the first features corresponding to each scale are successively processed by a batch normalization function (Batch Normalization, BN) and a ReLU6 activation function, the first features corresponding to each scale can be fused to obtain a second feature (not shown in the figure) by means of successive channel concatenation (concat) and channel shuffle processing. Then, the second feature and the processed input features are subjected to a first residual connection process (for example, element-wise addition processing) to obtain a third feature. Then, the number of channels is adjusted by point convolution processing, and then the adjusted third feature and the input features are subjected to a second residual connection process (for example, element-wise addition processing) to obtain an output feature.

[0050] Among them, the traditional yolov11 network includes a large-object detection head, a medium-object detection head, and a small-object detection head, which are respectively used to detect target objects of different sizes in an image. Since fine lesions are likely to exist in bladder cancer, a micro-object detection head is added to the output module of the target detection model in this application for the auxiliary diagnosis of bladder cancer. Exemplarily, the micro-object detection head can sample the input image with a stride smaller than that of the small-object detection head in the traditional yolov11 network and detect target objects based on the obtained feature map. Exemplarily, the sampling stride of the micro-object detection head can be 4. Then, for an input image of size , the micro-object detection head can sample it to obtain a feature map of size , and then detect target objects based on this feature map. Optionally, the large-object detection head in the output module can also be deleted in the target detection model of this application. Thus, the output module can combine the outputs of the micro-object detection head and other target detection heads respectively to obtain the target detection result of the target object in the cystoscope image.

[0051] In the above cystoscope image processing method, a target detection model based on the yolov11 network is used to perform target detection on cystoscope images. Among them, by replacing the c3k2 module in the yolov11 network with a multi-scale convolution module, it is possible to perform parallel per-channel convolution processing on the input features using multiple convolution kernels of different scales. Each convolution kernel can extract features from different receptive field ranges, so that it is possible to effectively capture features of different sizes in the image while maintaining computational efficiency. By fusing the first features output by each convolution kernel and then performing a residual connection with the input features, it is possible to obtain output features that contain multi-scale information and avoid information loss. Thus, by using the multi-scale convolution module, the computational amount and the number of parameters of the target detection model can be effectively reduced, and the detection accuracy and efficiency of target objects in cystoscope images can be improved. By introducing a micro target detection head into the model, the detection ability of the target detection model for small targets can be enhanced, which is beneficial to improving the detection accuracy of small lesion sites in cystoscope images.

[0052] In the detection of target objects in cystoscope images, since bladder cancer is prone to flat or non-prominent lesions, it is easily overlooked under white light. Based on this, before inputting the cystoscope image into the target detection model, this application first preprocesses it to improve the detection accuracy of the target detection model for target objects.

[0053] In an exemplary embodiment, as Figure 3 shown, using the target detection model to detect target objects in the cystoscope image, and obtaining the target object detection result of the cystoscope image, may include:

[0054] Step S301, extracting the luminance layer of the cystoscope image.

[0055] Among them, for the cystoscope image, an edge-preserving filtering method can be used to extract the luminance layer of the cystoscope image while retaining edge information.

[0056] In an exemplary embodiment, extracting the luminance layer of the cystoscope image may include: using the cystoscope image as a guidance map, performing guided filtering on the cystoscope image to obtain the luminance layer.

[0057] Specifically, in this application, a guided filtering method can be used to process the cystoscope image to extract the luminance layer. Exemplarily, assuming that the cystoscope image is an RGB three-channel image, then for each channel image , its own self can be used as the guidance map for guided filtering to obtain the luminance layer corresponding to each channel image . Where c represents the image channel (such as R, G, B), is the guided filtering function.

[0058] Step S302: Subtract the luminance layer from the cystoscope image to obtain the detail layer of the cystoscope image.

[0059] Among them, according to the luminance layer extracted in step S301, it can be multiplied by the stretching coefficient to calculate the stretched luminance layer, and then the stretched luminance layer can be subtracted from the cystoscope image to obtain the detail layer. Exemplarily, this process can be shown as the following formula:

[0060]

[0061] In the formula, c represents the image channel (for example, it can be R, G, B), is the detail layer corresponding to channel c, is the channel image corresponding to channel c, is the luminance layer corresponding to channel c, is the stretching coefficient, is the stretched luminance layer. Among them, the stretching coefficient can take a preset value determined by experiments.

[0062] Step S303: Perform weighted fusion on the luminance layer and the detail layer to obtain the processed cystoscope image.

[0063] Among them, based on the luminance layer and the detail layer of the cystoscope image, in this step, the two can be weighted and fused to enhance the cystoscope image. Exemplarily, the weighted fusion process can be shown as the following formula:

[0064]

[0065] In the formula, c represents the image channel (for example, it can be R, G, B), is the fused image corresponding to channel c, is the detail layer corresponding to channel c, is the gain coefficient, is the luminance layer corresponding to channel c, is the stretching coefficient. Among them, the gain coefficient can take a preset value determined by experiments.

[0066] Among them, after obtaining the fused images corresponding to each image channel, they can be combined to obtain the processed cystoscope image. Exemplarily, please refer to Figure 4 , where Figure 4 part (a) in is the original cystoscope image, Figure 4 part (b) in is the processed cystoscope image, and the details in the processed cystoscope image are significantly enhanced.

[0067] Step S304: Use the object detection model to detect the target object in the processed cystoscope image to obtain the target object detection result.

[0068] In this step, the processed cystoscope image can be input into the object detection model to obtain the target object detection result of the cystoscope image output by the object detection model.

[0069] In this embodiment, by using the edge-preserving filtering method to extract the luminance layer and detail layer of the cystoscope image, and then performing weighted fusion on the luminance layer and detail layer to obtain the processed cystoscope image, the detail contrast in the cystoscope image can be enhanced, which is beneficial to highlighting the position of the target object in the image and improving the accuracy of subsequent object detection using the object detection model.

[0070] Among them, in the related art, the detection model is usually trained using cystoscope images containing bladder cancer lesion sites. However, due to the scarcity of clinical data on bladder cancer, cystoscope images containing bladder cancer lesion sites are usually difficult to obtain in large quantities, which is likely to affect the training effect of the model. At the same time, cystoscope images are easily interfered by various factors such as lighting, angle, and mucus, which places high requirements on the generalization ability of the model. Based on this, this application expands the training samples of the object detection model to increase the data volume and diversity of the training samples, thereby improving the detection accuracy of the object detection model.

[0071] In an exemplary embodiment, as Figure 5 shown, the object detection model can be trained through the following steps:

[0072] Step S501: Obtain the first image sample containing the target object and the second image sample not containing the target object.

[0073] Among them, the first image sample can be a cystoscope image containing the target object (i.e., the tumor lesion site of bladder cancer), and the second image sample can be a cystoscope image not containing the target object. Exemplarily, in this step, multiple first image samples and multiple second image samples can be obtained respectively under the authorization of the patient. Among them, the multiple second image samples can include cystoscope images collected under the influence of different lighting, angles, mucus and other factors.

[0074] Step S502: Segment the first image sample to obtain the target segmentation image corresponding to the target object.

[0075] Among them, for the first image sample, it can be segmented to extract the target segmentation image containing the target object. Exemplarily, the target object can be segmented and labeled in the first image sample in advance to obtain the mask region corresponding to the target object, and then the target segmentation image corresponding to the target object can be segmented from the first image sample using the mask region.

[0076] Step S503: Fuse the target segmented image into the second image sample to obtain a third image sample.

[0077] Among them, by fusing the target segmented image into the second image sample, a third image sample containing the target object can be obtained. Exemplarily, the third image sample can be obtained by inserting the target segmented image into the second image sample.

[0078] Optionally, the same target segmented image can be respectively fused into multiple different second image samples to obtain multiple third image samples, or different target segmented images can be respectively fused into the same second image sample to obtain multiple third image samples.

[0079] Step S504: Use the sample set including the third image sample to train the target detection model to be trained, and obtain the target detection model.

[0080] Among them, after obtaining the third image sample, it can be added to the sample set for model training, and the samples in the sample set are used to train the target detection model to be trained to obtain the target detection model. Exemplarily, the sample set can include multiple third image samples, or can also include other image samples.

[0081] In this embodiment, for the training of the target detection model, the target segmented image corresponding to the target object is segmented from the first image sample containing the target object, and then it is fused into the second image sample without the target object. A third image sample containing the target object can be obtained by synthesis, which can effectively increase the number of samples for model training, and can construct a diverse data set. Then, the sample set including the third image sample is used to train the target detection model to be trained, which can enhance the generalization and accuracy of the target detection model and effectively solve the problems of scarce training data and insufficient diversity.

[0082] In an exemplary embodiment, as Figure 6 shown, fusing the target segmented image into the second image sample to obtain a third image sample may include:

[0083] Step S601: Perform edge blurring processing on the target segmented image to obtain a first target image.

[0084] In this step, the target segmented image can be subjected to edge blurring processing, such as by Gaussian filtering, etc., to obtain the first target image.

[0085] Step S602: Perform color transfer processing on the first target image according to the second image sample to obtain a second target image.

[0086] In this step, the second image sample can be used as a reference image to perform color transfer processing on the first target image, so as to obtain a second target image whose color matches that of the second image sample.

[0087] Exemplarily, in this step, the Reinhard color transfer algorithm can be used to perform color transfer processing on the first target image. Among them, the second image sample and the first target image can be first converted to the Lab color space, and the mean and standard deviation of the second image sample can be calculated, and the mean and standard deviation of the first target image can be calculated. Then, for each pixel of the first target image , linear transformation processing can be performed through the following formula to obtain the adjusted pixel value: wherein,

[0088]

[0089] is the adjusted pixel value of the pixel , is the standard deviation of the second image sample in channel L, is the standard deviation of the first target image in channel L, is the pixel value of the pixel of the first target image in channel L, is the mean of the first target image in channel L, is the mean of the second image sample in channel L; is the standard deviation of the second image sample in channel a, is the standard deviation of the first target image in channel a, is the pixel value of the pixel of the first target image in channel a, is the mean of the first target image in channel a, is the mean of the second image sample in channel a; is the standard deviation of the second image sample in channel b, is the standard deviation of the first target image in channel b, is the pixel value of the pixel of the first target image in channel b, is the mean of the first target image in channel b, is the mean of the second image sample in channel b.

[0090] ​​After all pixels of the first target image are adjusted by the above method, the adjusted first target image can be converted to the RGB space to obtain a second target image.

[0091] Step S603: Insert the second target image into the second image sample to obtain a third image sample.

[0092] In this step, by inserting the second target image into the second image sample, a third image sample containing the target object can be obtained. Among them, the insertion position of the second target image in the second image sample can be specified using a position parameter, or the size of the second target image in the second image sample can be specified using a size parameter.

[0093] It can be understood that in some embodiments, the first target image can also be inserted into the second image sample first, and then color transfer processing is performed on the first target image to obtain a third image sample containing the second target image. Exemplarily, please refer to Figure 7 , based on the first image sample containing the target object and the second image sample not containing the target object, the target segmentation image can be first segmented from the first image sample, and its edges are blurred to obtain the first target image, and then the first target image is inserted into the second image sample as a layer, and then color transfer processing is performed on the first target image according to the second image sample to obtain a third image sample containing the second target image.

[0094] In this embodiment, by performing edge blurring processing on the target sample image and performing color transfer processing on the target sample image according to the second image sample, the target object inserted into the second image sample can be made to transition naturally with the second image sample as the background, and the tone of the target object can be made to match that of the second image sample, so that the synthesized third image sample is closer to the actually captured cystoscope image. Using this third image sample for model training can increase the accuracy of the target detection model and is beneficial to improving the accuracy of target detection for cystoscope images.

[0095] In an exemplary embodiment, inserting the second target image into the second image sample to obtain a third image sample may include: performing a scaling process on the second target image according to a random size parameter to obtain a scaled second target image; determining a target insertion position in the second image sample according to a random position parameter, and inserting the scaled second target image into the target insertion position to obtain a third image sample.

[0096] Specifically, when inserting the second target image into the second image sample, random size parameters and random position parameters can be obtained first. The random size parameters are used to indicate the target insertion size of the second target image in the second image sample, and the random position parameters are used to indicate the target insertion position of the second target image in the second image sample. Both are randomly generated parameters. Among them, according to the random size parameters, the second target image can be scaled so that the size of the second target image conforms to the target insertion size. Then, the scaled second target image can be inserted into the target insertion position indicated by the random position parameters to obtain the synthesized third image sample.

[0097] In this embodiment, by randomly setting the size and position of the second target image inserted into the second image sample, the data diversity of the third image sample can be improved, which is beneficial to enhancing the generalization and accuracy of the target detection model.

[0098] In one embodiment, a method for processing cystoscope images in a specific embodiment is provided, which specifically includes the following steps:

[0099] Step S1, obtain cystoscope images.

[0100] Step S2, perform image preprocessing enhancement on the cystoscope images.

[0101] Among them, the cystoscope image itself can be used as a guidance map first. The luminance layer of the cystoscope image is extracted by the guided filter method. Then, the luminance layer can be multiplied by a stretching coefficient to obtain a stretched luminance layer. Then, the stretched luminance layer is subtracted from the cystoscope image to obtain a detail layer. Then, the detail layer is multiplied by a gain coefficient and added to the stretched luminance layer, and the processed cystoscope image is obtained through weighted fusion.

[0102] Step S3, use the target detection model to detect the target objects in the processed cystoscope images to obtain the target object detection results.

[0103] Among them, the target detection model in this embodiment is a model constructed based on the yolov11 network. This target detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network, and its output module includes a micro target detection head. Among them, in the multi-scale convolution module of the target detection model, the input features can be convolved channel by channel using convolution kernels corresponding to different scales to obtain first features corresponding to each scale. Then, the first features are fused to obtain a second feature, and then the second feature is connected to the input feature by a residual connection to obtain an output feature.

[0104] Among them, the target detection model in this embodiment can be trained through the following steps:

[0105] Step S4, obtain a first image sample containing the target object and a second image sample not containing the target object.

[0106] Among them, the number of second image samples can be multiple, and the multiple second image samples can include cystoscope images collected under the influence of different factors such as illumination, angle, mucus, etc.

[0107] Step S6, segment the first image sample to obtain a target segmentation image corresponding to the target object.

[0108] Among them, the mask region can be obtained by segmenting and annotating the first image sample, and then the first image sample can be segmented using the mask region to obtain the target segmentation image.

[0109] Step S6, fuse the target segmentation image into the second image sample to obtain a third image sample.

[0110] Exemplarily, in this step, the target segmentation image can be first inserted into the second image sample, and the edges of the target segmentation image can be blurred using Gaussian filtering, and then the Reinhard color transfer algorithm can be used to perform color transfer processing on the target segmentation image based on the second image sample so that the hue of the target segmentation image matches that of the second image sample.

[0111] Exemplarily, in this step, the edges of the target segmentation image can be blurred first to obtain a first target image, then the first target image can be color transferred according to the second image sample to obtain a second target image, and then the second target image can be inserted into the second image sample to obtain a third image sample.

[0112] Among them, whether the target segmentation image is first inserted into the second image sample or the target segmentation image is first processed and then the second target image is inserted into the second image sample, the corresponding target insertion size and target insertion position can be randomly set, so as to improve the data diversity of the third image sample.

[0113] Step S7, use the sample set including the third image sample to train the target detection model to be trained to obtain the target detection model.

[0114] In this embodiment, by synthesizing data, a third image sample is synthesized using a first image sample containing the target object and second image samples collected under different scenarios, which can construct a diverse dataset, obtain more available data for model training, enhance the generalization of the model, and effectively solve the problems of data scarcity and insufficient diversity. Moreover, in this solution, by introducing depthwise separable convolution and multi-scale convolution kernels into the yolov11 network and combining them, while reducing the computational amount, the model's feature extraction ability can be enhanced. And for the application scenario of cystoscope image detection, the large object detection head in the yolov11 network is deleted and a micro-object detection head is added, which can enhance the detection ability of the target detection model for small objects. At the same time, before inputting the cystoscope image into the target detection model, guided filtering is first used to extract and enhance the details of the RGB channels of the image, which can effectively enhance the image detail contrast of the cystoscope image, highlight the position of the target object, and is conducive to further improving the detection accuracy of the target object.

[0115] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0116] Based on the same inventive concept, an embodiment of the present application also provides a cystoscope image processing device for implementing the above-mentioned cystoscope image processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following cystoscope image processing device can refer to the limitations on the cystoscope image processing method in the above text, and will not be repeated here.

[0117] In an exemplary embodiment, as Figure 8 shown, a cystoscope image processing device 800 is provided, including:

[0118] An image acquisition module 801, configured to acquire a cystoscope image.

[0119] A target detection module 802, configured to detect a target object in the cystoscope image using a target detection model to obtain a target object detection result of the cystoscope image.

[0120] Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head. The multi-scale convolution module performs per-channel convolution on the input feature using multiple convolution kernels with different scales respectively to obtain the first features output by each convolution kernel, fuses the first features to obtain the second feature, and performs a residual connection between the second feature and the input feature to obtain the output feature.

[0121] In an exemplary embodiment, the object detection module 802 is configured to: extract the luminance layer of the cystoscope image; subtract the luminance layer from the cystoscope image to obtain the detail layer of the cystoscope image; perform weighted fusion on the luminance layer and the detail layer to obtain the processed cystoscope image; use the object detection model to detect the target object in the processed cystoscope image to obtain the target object detection result.

[0122] In an exemplary embodiment, the luminance extraction module is configured to perform guided filtering on the cystoscope image with the cystoscope image as the guidance map to obtain the luminance layer.

[0123] In an exemplary embodiment, the object detection model is trained through the following steps: obtaining a first image sample containing the target object and a second image sample not containing the target object; segmenting the first image sample to obtain a target segmentation image corresponding to the target object; fusing the target segmentation image into the second image sample to obtain a third image sample; using the sample set including the third image sample to train the object detection model to be trained to obtain the object detection model.

[0124] In an exemplary embodiment, fusing the target segmentation image into the second image sample to obtain the third image sample includes: performing edge blurring on the target segmentation image to obtain a first target image; performing color transfer on the first target image according to the second image sample to obtain a second target image; inserting the second target image into the second image sample to obtain the third image sample.

[0125] In an exemplary embodiment, inserting the second target image into the second image sample to obtain the third image sample includes: scaling the second target image according to a random size parameter to obtain a scaled second target image; determining a target insertion position in the second image sample according to a random position parameter, and inserting the scaled second target image into the target insertion position to obtain the third image sample.

[0126] Each module in the above cystoscope image processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of a computer device in hardware form or independent thereof, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0127] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a cystoscope image processing method.

[0128] Those skilled in the art can understand that Figure 9 the structure shown in

[0129] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0130] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0131] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0135] The above embodiments only express several implementation manners of this application, and their descriptions are relatively specific and detailed. However, it should not be construed as a limitation to the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A method for processing images of a cystoscope, characterized in that, The method includes: Obtaining a cystoscope image; Using a target detection model to detect target objects in the cystoscope image, obtaining a target object detection result of the cystoscope image; Wherein, the target detection model is a model constructed based on the yolov11 network, the target detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network, and the output module of the target detection model includes a micro target detection head; the multi-scale convolution module performs per-channel convolution on the input features using convolution kernels corresponding to different scales respectively, obtains first features corresponding to each of the scales, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

2. The method according to claim 1, wherein The detecting of the target objects in the cystoscope image using the target detection model to obtain the target object detection result of the cystoscope image includes: Extracting the brightness layer of the cystoscope image; Subtracting the brightness layer from the cystoscope image to obtain the detail layer of the cystoscope image; Performing weighted fusion on the brightness layer and the detail layer to obtain a processed cystoscope image; Using the target detection model to detect target objects in the processed cystoscope image to obtain the target object detection result.

3. The method according to claim 2, characterized in that, The extracting of the brightness layer of the cystoscope image includes: Using the cystoscope image as a guidance map and performing guided filtering on the cystoscope image to obtain the brightness layer.

4. The method according to claim 1, wherein The target detection model is trained through the following steps: Obtaining a first image sample containing target objects and a second image sample not containing the target objects; Segmenting the first image sample to obtain a target segmentation image corresponding to the target objects; Fusing the target segmentation image into the second image sample to obtain a third image sample; Using a sample set including the third image sample to train a target detection model to be trained to obtain the target detection model.

5. The method according to claim 4, wherein The fusing of the target segmentation image into the second image sample to obtain a third image sample includes: Performing edge blurring on the target segmentation image to obtain a first target image; Performing color transfer on the first target image according to the second image sample to obtain a second target image; Inserting the second target image into the second image sample to obtain the third image sample.

6. The method according to claim 5, wherein The inserting of the second target image into the second image sample to obtain the third image sample includes: Performing scaling on the second target image according to random size parameters to obtain a scaled second target image; Determining a target insertion position in the second image sample according to random position parameters, and inserting the scaled second target image into the target insertion position to obtain the third image sample.

7. A cystoscope image processing device, characterized in that, The device includes: An image acquisition module for obtaining a cystoscope image; A target detection module for using a target detection model to detect target objects in the cystoscope image, obtaining a target object detection result of the cystoscope image; Among them, the object detection model is a model constructed based on the yolov11 network. The object detection model uses a multi-scale convolution module to replace the c3k2 module in the yolov11 network. The output module of the object detection model includes a micro-object detection head; the multi-scale convolution module respectively performs per-channel convolution on the input features using multiple convolution kernels of different scales to obtain first features output by each convolution kernel, fuses the first features to obtain a second feature, and performs a residual connection between the second feature and the input feature to obtain an output feature.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 6.