Defect detection method, electronic equipment and program product

By combining the interactive attention unit of visual encoder and text encoder, the problems of low efficiency and poor robustness of traditional medical balloon detection are solved, and efficient and accurate defect detection of medical balloons is achieved.

CN121982015APending Publication Date: 2026-05-05SUZHOU HUANQIU MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU HUANQIU MEDICAL TECHNOLOGY CO LTD
Filing Date
2026-03-24
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional medical balloon defect detection relies on manual observation and traditional image processing, which has low detection efficiency, poor robustness, and insufficient sensitivity to minute defects. It cannot accurately identify minute defects such as fine scratches, tiny fisheyes, and trace amounts of gel-like substances, making it difficult to meet the strict quality control standards and high-efficiency detection requirements of medical balloon production.

Method used

A defect detection model is adopted, including a visual encoder, a text encoder, and a defect recognition module. An interactive attention unit is set in the visual encoder. By acquiring the target medical balloon image and the detection prompt text, multimodal joint reasoning is performed by combining the visual features of the image and the semantic information of the text to identify whether there are defects in the balloon image and their types.

Benefits of technology

It improves the accuracy and robustness of defect detection in medical balloons, enables precise identification of minute defects, and meets the quality control requirements of large-scale production of medical balloons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982015A_ABST
    Figure CN121982015A_ABST
Patent Text Reader

Abstract

The invention discloses a defect detection method, electronic equipment and a program product. The method comprises the steps of obtaining a to-be-detected target medical balloon image and a detection prompt text; the target medical balloon image and the detection prompt text are input into a defect detection model, a defect detection result corresponding to the target medical balloon image is obtained, the defect detection model comprises a visual encoder, a text encoder and a defect recognition module, the visual encoder is provided with an interactive attention unit, and the text encoder is provided with an interactive attention unit; the defect detection result is used for indicating whether the medical balloon image has defects or not and the defect type of the medical balloon image when the defects exist. According to the technical scheme, the image of the to-be-detected target medical balloon is analyzed through the defect detection model, the problems of low detection efficiency, poor robustness and the like of traditional defect detection of the medical balloon are solved, and the defect detection accuracy of the medical balloon is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of defect detection technology, and in particular to a defect detection method, electronic device and program product. Background Technology

[0002] With the rapid development of medicine, balloons are used as key functional components in interventional devices. Their structure and performance directly affect the dilation effect, vascular / tissue safety, and the controllability of the surgical procedure. If the balloon has defects, it may lead to abnormal dilation pressure, balloon rupture, air leakage failure, uneven dilation, vascular damage, or intraoperative complications in clinical use.

[0003] Traditional defect detection in medical balloons relies primarily on manual observation and conventional image processing methods, which suffer from low detection efficiency, poor robustness, and insufficient sensitivity to minute defects. It cannot accurately identify minute defects such as fine scratches, tiny fisheyes, or trace amounts of gelatinous material, nor can it guarantee the consistency and stability of test results, making it difficult to meet the stringent quality control standards and high-efficiency testing requirements of medical balloon production. Summary of the Invention

[0004] This invention provides a defect detection method, electronic device, and program product to improve the accuracy of defect detection for medical balloons.

[0005] According to one aspect of the present invention, a defect detection method is provided, comprising: Acquire the target medical balloon image to be detected and the detection prompt text, wherein the detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image; The target medical balloon image and the detection prompt text are input into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image and the type of defect in the medical balloon image when a defect exists.

[0006] According to another aspect of the present invention, a defect detection device is provided, the device comprising: The data acquisition module is used to acquire the target medical balloon image to be detected and the detection prompt text. The detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. The defect detection result determination module is used to input the target medical balloon image and the detection prompt text into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image and the type of defect in the medical balloon image when a defect exists.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the defect detection method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the defect detection method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, which, when executed by a processor, implements the defect detection method as described in any of the embodiments of the present invention.

[0010] This invention acquires a target medical balloon image and detection prompt text. The detection prompt text indicates the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. By acquiring the target medical balloon image and detection prompt text, a complete data foundation and input conditions are provided for the subsequent inference process of the defect detection model, ensuring the execution of subsequent detection operations. The target medical balloon image and detection prompt text are input into the defect detection model to obtain a defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result indicates whether the medical balloon image has a defect, and if so, the type of defect. By jointly inputting the target medical balloon image and detection prompt text into the defect detection model, and combining image visual features with text prompt information to achieve multimodal joint inference, the accuracy and robustness of defect detection are effectively improved. In summary, the technical solution of this invention analyzes the image of the target medical balloon to be detected using a defect detection model, which solves the problems of low detection efficiency and poor robustness of traditional medical balloon defect detection, and improves the accuracy of defect detection for medical balloons.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a defect detection method according to an embodiment of the present invention; Figure 2 This is a flowchart of a defect detection method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the defect detection model in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the interactive attention unit in an embodiment of the present invention; Figure 5 This is a detailed structural diagram of the defect detection model in this embodiment of the invention; Figure 6 This is a schematic diagram of the structure of a defect detection device according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0017] Figure 1 This is a flowchart illustrating a defect detection method provided in an embodiment of the present invention. This embodiment is applicable to defect detection of medical balloon images. The method can be executed by the defect detection device in this embodiment, which can be implemented in software and / or hardware. This device can be integrated into electronic devices such as computer equipment, servers, mobile terminals, processors, or surgical robots. Figure 1 As shown, the method specifically includes the following steps: S110. Obtain the target medical balloon image to be detected and the detection prompt text, wherein the detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image.

[0018] In this embodiment of the invention, the target medical balloon image can be specifically understood as a medical balloon image that requires defect detection. The target medical balloon image can be acquired through industrial cameras or imaging equipment, covering various possible defects that may exist during the production and quality inspection processes of the medical balloon. The target medical balloon image provides input data for the subsequent defect detection model. The detection prompt text can be specifically understood as text information used to guide the defect detection model to perform a specified detection task. The detection prompt text clearly indicates the type of detection result that the defect detection model needs to detect in the target medical balloon image. Through the combination of text and image information, a clear detection guide is provided to the model, enabling the model to perform accurate detection according to the preset detection intent. The defect detection model can be specifically understood as a model used to detect defects in medical balloon images. The defect detection model can determine whether a medical balloon image has defects and accurately output the corresponding defect type when defects are present. The defect detection model can improve the accuracy of defect detection through multimodal cooperation of images and text.

[0019] Specifically, an image of the target medical balloon to be detected is acquired using a specific imaging device. Based on the image content of the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection results that the defect detection model needs to output for the target medical balloon image, detection prompt text is set. By acquiring the target medical balloon image and the detection prompt text, a complete data foundation and input conditions are provided for the subsequent inference process of the defect detection model, ensuring the execution of subsequent detection operations.

[0020] S120. Input the target medical balloon image and the detection prompt text into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image and the type of defect in the medical balloon image when a defect exists.

[0021] In this embodiment of the invention, the defect detection result can be specifically understood as the output of the defect detection model. The defect detection result is used to clearly indicate whether there are defects in the target medical balloon image. If defects are found, the corresponding defect type (bubble, fisheye, gel-like substance, scratch) is further precisely output. The defect detection result provides medical balloon quality inspectors with an intuitive and clear basis for judgment, enabling rapid identification of defects in medical balloons, replacing the subjectivity and lag of traditional manual judgment, and adapting to the quality control needs of large-scale medical balloon production.

[0022] The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder incorporates an interactive attention unit. Specifically, the visual encoder is responsible for extracting visual features from the target medical balloon image. It transforms image information into visual feature vectors that can be used for subsequent defect detection, providing a reliable visual feature foundation. The interactive attention unit enhances image feature extraction and improves defect perception capabilities. It performs spatial enhancement and cross-scale feature fusion processing on the feature map output by the visual encoder. The interactive attention unit significantly amplifies the response intensity of defects in the feature map, thereby improving the defect detection model's sensitivity to multi-scale defects and its robustness in boundary localization, providing more accurate and effective visual features for the subsequent defect recognition module.

[0023] The text encoder is a module used for feature extraction and semantic encoding of the detection prompt text. It transforms the input detection prompt text into text features, converting textual information into features recognizable by the defect detection model. By semantically encoding the detection prompt text, the text encoder forms multimodal feature matching with the image features output by the visual encoder, providing accurate detection guidance and semantic constraints for the subsequent defect recognition module. The defect recognition module is used for multimodal feature fusion and defect identification within the defect detection model. It fuses and infers the image features output by the visual encoder and the text features output by the text encoder. The defect recognition module further determines whether defects exist in the target medical balloon image, and if defects are present, further identifies specific defect types such as streaks, fisheyes, gel-like substances, and scratches, outputting the final defect detection result.

[0024] Specifically, the defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder incorporates an interactive attention unit. The target medical balloon image and the detection prompt text are input into the defect detection model to obtain a defect detection result corresponding to the target medical balloon image. The defect detection result includes an indication of whether a defect exists in the medical balloon image, and if so, the type of defect. By inputting both the target medical balloon image and the detection prompt text into the defect detection model, the model can fully utilize the visual features of the image and the semantic information of the text, making the defect detection results more accurate.

[0025] Optionally, the training process of the defect detection model includes: acquiring an original defect image dataset, the original defect image dataset including multiple sample medical balloon images, and the multiple sample medical balloon images including normal medical balloon images and defective medical balloon images, the defective medical balloon images including at least one defective region; performing data augmentation processing on at least some sample medical balloon images in the original defect image dataset to obtain expanded medical balloon images, adding the expanded medical balloon images to the original defect image dataset to obtain a target defect image dataset; training the model based on the target defect image dataset to obtain the defect detection model.

[0026] In this embodiment of the invention, the original defect image dataset is a sample dataset used to train the defect detection model. The original defect image dataset includes images of normal medical balloons and images of defective medical balloons. As the original data for training the defect detection model, the original defect image dataset provides the model with basic feature comparisons between defective and normal samples. The sample medical balloon images constitute the images in the original defect image dataset. These sample medical balloon images can be acquired using imaging equipment. They provide the model with realistic and rich visual features of the balloon surface, serving as the training sample basis for the defect detection model to learn defect features and distinguish between normal and defective samples. Normal medical balloon images are images of qualified medical balloons without any defects. Normal medical balloon images are images of qualified medical balloons whose surfaces are free of any defects such as streaks, fisheyes, gel-like substances, scratches, etc., and which meet quality standards. Normal medical balloon images are used to provide normal sample features during the defect detection model training process, enabling the defect detection model to learn and distinguish the feature differences between the surfaces of normal and defective balloons, thereby improving the model's discrimination accuracy and robustness. Defective medical balloon images are images of balloons containing at least one defective region, such as bubbly bands, fisheyes, gel-like substances, or scratches. These images provide defect feature learning samples for defect detection models, enabling the model to effectively learn the appearance, contour, and texture features of various defects during training, thus providing data support for subsequent defect identification, classification, and judgment. The defective region is the actual location of defects such as bubbly bands, fisheyes, gel-like substances, or scratches on the balloon surface. The defective region is the target area that the defect detection model needs to identify, locate, and judge during training and inference, providing crucial evidence for the model to learn defect features and achieve defect detection and classification.

[0027] The extended medical balloon images are obtained by augmenting sample medical balloon images from the original defect image dataset. Augmenting the original samples expands the number of training samples, enriches sample diversity, improves the generalization ability and robustness of the defect detection model, and avoids overfitting during training. The target defect image dataset is composed of the original defect image dataset and the extended medical balloon images. The target defect image dataset combines the realism of the original samples with the diversity of the extended samples, effectively improving the model's feature learning ability, generalization ability, and robustness, providing sufficient and high-quality data support for training a stable and accurate defect detection model.

[0028] Specifically, multiple images of normal and defective medical balloons are acquired to construct an original defective image dataset. Each defective medical balloon image contains at least one defective region. This original dataset, constructed by collecting both normal and defective medical balloon images, provides realistic and effective basic samples for training the defect detection model. Data augmentation processing is performed on at least a portion of the medical balloon images in the original defective image dataset to obtain expanded medical balloon images. These expanded medical balloon images are then added to the original defective image dataset to obtain the target defective image dataset. The model is trained on the target defective image dataset to obtain the defect detection model. Data augmentation and expansion of the original samples into the target dataset effectively increases the number and diversity of samples, significantly enhancing the generalization ability of the defect detection model.

[0029] Optionally, data augmentation processing is performed on at least a portion of the sample medical balloon images in the original defective image dataset to obtain extended medical balloon images, including: annotating at least one defective region in multiple defective medical balloon images to obtain multiple defect-annotated images; generating multiple mask images corresponding to multiple defective regions based on the multiple defect-annotated images, and using the minimum bounding rectangle of the mask images as a foreground patch; when the intersection-union ratio (IUU) of the foreground patch and the defective region in the image region corresponding to the defective medical balloon image is less than a first preset threshold, adding the foreground patch to the effective imaging region of at least a portion of the sample medical balloon images to obtain an enhanced defective image, wherein the effective imaging region is the region where the medical balloon is located in the image; determining the area ratio of the pixel area of ​​the defective region in the enhanced defective image to the pixel area of ​​the defective region in the mask image, and determining the enhanced defective image with the area ratio greater than or equal to a second preset threshold as an extended medical balloon image.

[0030] In this embodiment of the invention, the defect-annotated image is an image obtained by annotating the location and category of defective regions in a defective medical balloon image. The defect-annotated image records the coordinates, contour, and category (strips, fisheyes, gel-like substances, scratches, etc.) of the defective region, providing accurate defect location information for subsequent generation of mask images, foreground patches, and data augmentation processing. For example, the image annotation tool labelimg can be used for defect annotation. The mask image is a binary image generated based on the defect-annotated image to distinguish between defective and background regions. By clearly identifying defective and non-defective regions in the image through different pixel values, it can accurately locate the contour, location, and extent of defects, providing precise region segmentation information for subsequent foreground patch extraction and data augmentation operations.

[0031] The foreground patch is an image captured by cropping the smallest bounding rectangle of the defect region in the mask image. The foreground patch contains the complete defect and serves as the foreground unit used in data augmentation to copy and add to generate expanded medical balloon images. The foreground patch can expand sample diversity through migration and pasting without altering the essential characteristics of the original defect. The first preset threshold is a value used to determine whether the intersection-union ratio (IU) between the foreground patch and the original defect region meets the requirements. This first preset threshold mainly controls the foreground patch from excessively overlapping with the original defect region when added to the sample medical balloon image, ensuring a reasonable defect distribution and effective samples in the generated enhanced defect image, thereby improving the reliability of data augmentation and the stability of model training. For example, the first preset threshold can be 0.2. The effective imaging region is the area actually occupied by the medical balloon in the image. The effective imaging region excludes pure background and invalid non-balloon regions, representing the effective image range truly used for defect detection and data augmentation. The effective imaging region ensures that the foreground patch is only pasted within the real area of ​​the balloon surface, avoiding the generation of invalid defects in non-balloon background areas, ensuring the realism of the enhanced image and the effectiveness of training. The enhanced defect image is the image obtained by adding a foreground patch to the effective imaging area of ​​the sample medical balloon image as required. The enhanced defect image is an intermediate sample image generated through data augmentation, used to enrich the distribution location and combination form of defects. The second preset threshold is a pre-set area ratio threshold used to filter the enhanced defect images. By comparing the pixel area ratio of the defect region in the enhanced defect image and the mask image, the effectiveness of the generated enhanced defect image and the completeness and clarity of the defects are determined, ensuring the quality and reliability of the expanded medical balloon image and avoiding a decrease in model training effect due to defect deformation or incompleteness. For example, the second preset threshold can be set to 0.5.

[0032] Specifically, at least one defective region in multiple defective medical balloon images is labeled to obtain multiple defect-labeled images. By labeling defective regions and associating them with defect categories, a reliable data foundation is provided for subsequent generation of mask images and extraction of foreground patches, ensuring the integrity and accuracy of the defect target during data augmentation. Multiple mask images corresponding to multiple defective regions are generated based on the multiple defect-labeled images, and the minimum bounding rectangle of the mask images is used as the foreground patch. When the intersection-union ratio (IU) of the foreground patch and the defective region in the image region corresponding to the defective medical balloon image is less than a first preset threshold, the foreground patch is added to the effective imaging region of at least some sample medical balloon images to obtain an enhanced defective image. By controlling the IU of the foreground patch and the original defective region within the first preset threshold and adding the foreground patch to the effective imaging region of the sample medical balloon images, the number of defects and their distribution positions can be flexibly increased and changed without compromising the original defect realism, thereby enriching the diversity of defect scenes in the training samples. The ratio of the pixel area of ​​the defective region in the enhanced defective image to the pixel area of ​​the defective region in the mask image is determined. Enhanced defect images with an area ratio less than a second preset threshold are discarded, while enhanced defect images with an area ratio greater than or equal to the second preset threshold are identified as expanded medical balloon images. By discarding enhanced defect images with an area ratio less than the second preset threshold, invalid samples with severely deformed, occluded, incomplete, or missing defects can be effectively filtered out, retaining only valid samples with complete, clear defects that conform to the actual situation.

[0033] Optionally, when adding a foreground patch to sample medical balloon images, the probability of adding the patch to each sample medical balloon image and the maximum number of times each sample medical balloon image can be added can be set. For example, the probability of adding the patch to each sample medical balloon image can be set to 0.3, and the maximum number of times each sample medical balloon image can be added can be set to 3. By setting the probability of adding the foreground patch to each sample medical balloon image and the maximum number of times it can be added, the number of enhanced samples generated and the enhancement intensity can be reasonably controlled while ensuring the data augmentation effect, avoiding problems such as abnormal defect distribution and sample distortion caused by over-enhancing the same sample.

[0034] Optionally, after obtaining the defect detection results, the false negative rate and false positive rate can be used to evaluate the defect detection results, and the calculation formula is as follows: Quantitatively evaluating defect detection results using false negative and false positive rates allows for an objective and accurate measurement of the detection performance and reliability of the defect detection model. Evaluating the defect detection results provides a quantitative basis for subsequent model optimization, ensuring the practicality and effectiveness of the solution.

[0035] The technical solution of this embodiment acquires a target medical balloon image and detection prompt text. The detection prompt text is used to indicate the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. By acquiring the target medical balloon image and detection prompt text, a complete data foundation and input conditions are provided for the subsequent inference process of the defect detection model, ensuring the execution of subsequent detection operations. The target medical balloon image and detection prompt text are input into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image, and if so, the type of defect in the medical balloon image. By inputting the target medical balloon image and detection prompt text together into the defect detection model, and combining image visual features and text prompt information to achieve multimodal joint inference, the accuracy and robustness of defect detection are effectively improved. In summary, the technical solution of this invention analyzes the image of the target medical balloon to be detected using a defect detection model, which solves the problems of low detection efficiency and poor robustness of traditional medical balloon defect detection, and improves the accuracy of defect detection for medical balloons.

[0036] Figure 2 This is a flowchart of a defect detection method provided in Embodiment 2 of the present invention. Based on any optional technical solution of the present invention, the technical solution of this embodiment can further refine the structure of the defect detection model and the defect detection processing flow. Optionally, the defect detection model includes a visual encoder, a text encoder, and a defect recognition module, wherein the visual encoder is provided with an interactive attention unit. A schematic diagram of the defect detection model is shown below. Figure 3 As shown. Detailed implementation methods can be found in the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2 As shown, the method may specifically include: S210. Obtain the target medical balloon image to be detected and the detection prompt text, wherein the detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image.

[0037] S220. Input the target medical balloon image into a visual encoder to obtain a target feature map, and input the detection prompt text into a text encoder to obtain text features.

[0038] In this embodiment of the invention, the visual encoder can be specifically understood as a module for extracting features from image input. The visual encoder encodes the target medical balloon image into a target feature map. The features extracted by the visual encoder support subsequent feature fusion and defect detection. The target feature map can be specifically understood as a high-dimensional feature representation output by the visual encoder after feature extraction and encoding of the target medical balloon image. The target feature map retains key visual information related to defect detection in the target medical balloon image, discards irrelevant redundant pixel information, and presents high-dimensional semantic features in matrix form, which can be recognized and utilized by subsequent modules, providing core visual feature support for fusion with text features and completion of the defect detection task. The text features can be specifically understood as features obtained by the text encoder after semantically encoding the detection prompt text. The text features can be aligned and fused with the target feature map in the same feature space, thereby realizing medical balloon defect detection based on text prompt guidance and improving the targeting and accuracy of detection.

[0039] Specifically, the target medical balloon image is input into a visual encoder to obtain a target feature map. The detection prompt text is input into a text encoder to obtain text features. The target feature map and text features provide a reliable feature basis for subsequent defect detection, improving the targeting, flexibility, and accuracy of defect detection.

[0040] Optionally, the visual encoder includes multiple serially connected feature extraction modules; at least two of the feature extraction modules include serially connected feature extraction units and interactive attention units, and the interactive attention units in at least two of the feature extraction modules correspond to different feature transformation scales.

[0041] In this embodiment of the invention, the feature extraction module can be specifically understood as a processing module within the visual encoder. The feature extraction module is used for feature transformation and information extraction of medical balloon images. The feature extraction module can process input features at multiple scales and levels, progressively extracting features from the balloon image. At least two feature extraction modules include a serially connected feature extraction unit and an interactive attention unit. The feature extraction unit is the core unit within the feature extraction module. Its core function is to perform preliminary abstraction, filtering, and information extraction on the input target medical balloon image, capturing basic visual features such as underlying textures and edge contours, and transforming them into feature representations that can be further processed. The interactive attention unit is a unit within the feature extraction module that is serially connected to the feature extraction unit. Its core function is to perform targeted enhancement and multi-scale interaction on the basic features output by the feature extraction unit, focusing on and highlighting key features related to defects in the medical balloon, and suppressing irrelevant redundant information. The interactive attention units in at least two feature extraction modules correspond to different feature transformation scales. The feature extraction unit and the interactive attention unit are connected in series. The feature extraction unit first extracts and performs preliminary transformation of basic features, and then passes the output features to the interactive attention unit for multi-scale feature interaction and enhancement. This provides basic support for the visual encoder to extract deep semantic features and defect-related features of medical balloon images layer by layer, ensuring the accuracy of subsequent feature fusion and defect detection.

[0042] Optionally, the step of inputting the target medical balloon image into a visual encoder to obtain a target feature map includes: inputting the target medical balloon image into the feature extraction module, performing feature extraction on the target medical balloon image through the feature extraction unit of the feature extraction module to obtain an original feature map; if the feature extraction module includes an interactive attention unit, performing attention enhancement processing on the original feature map based on the interactive attention unit to obtain a target enhanced feature map, and inputting the target enhanced feature map into the next feature extraction module; if the feature extraction module does not include an interactive attention unit, inputting the original feature map into the next feature extraction module; and using the feature map output from the last feature extraction module as the target feature map.

[0043] In this embodiment of the invention, the original feature map can be specifically understood as the feature map obtained after the feature extraction unit performs preliminary feature extraction on the target medical balloon image. The original feature map has not been processed by the interactive attention unit. The original feature map is used to provide input for the subsequent interactive attention unit and is the prerequisite and foundation for achieving attention enhancement. The target enhanced feature map can be specifically understood as the feature map obtained by the interactive attention unit processing the original feature map. Based on the original feature map, the target enhanced feature map highlights the key features related to the medical balloon defect, weakens irrelevant background and redundant information, and has stronger representation ability and discriminative power, which can provide more accurate and effective feature support for the subsequent feature extraction module or final defect detection.

[0044] Specifically, the target medical balloon image is input into the feature extraction module. The feature extraction unit of this module extracts features from the target medical balloon image to obtain an original feature map. If the feature extraction module includes an interactive attention unit, the original feature map is input into the interactive attention unit, which performs attention enhancement processing on the original feature map to obtain a target enhanced feature map. The target enhanced feature map is then input into the next feature extraction module. If the feature extraction module does not include an interactive attention unit, the original feature map is input into the next feature extraction module. The feature map output from the last feature extraction module is used as the target feature map. By using a combination of feature extraction and interactive attention units to process the target medical balloon image, layer-by-layer attention enhancement of defect-related features can be performed on top of basic feature extraction, highlighting key features of balloon surface defects and suppressing irrelevant background information.

[0045] Optionally, the interactive attention unit includes a local feature subunit and a semantic feature subunit. The step of performing attention enhancement processing on the original feature map based on the interactive attention unit to obtain a target enhanced feature map includes: processing the original feature map based on the local feature subunit to obtain a first distribution map and a first channel vector; processing the original feature map based on the semantic feature subunit to obtain a second distribution map and a second channel vector; performing matrix multiplication on the first distribution map and the second channel vector to obtain a first spatial attention weight, and performing matrix multiplication on the second distribution map and the first channel vector to obtain a second spatial attention weight; obtaining an attention weight based on the first spatial attention weight and the second spatial attention weight; and performing a matrix dot product of the attention weight and the original feature map to obtain the target enhanced feature map.

[0046] In this embodiment of the invention, Figure 4This is a schematic diagram of the interactive attention unit. The local feature subunit is one of the components of the interactive attention unit. The core function of the local feature subunit is to extract local features, enhance spatial representation, and characterize features from the input original feature map, focusing on local details related to defects in the medical balloon image. The local feature subunit provides local feature support for subsequent feature interactions with the semantic feature subunit and the calculation of spatial attention weights. The semantic feature subunit is also one of the components of the interactive attention unit. The semantic feature subunit is used to extract and characterize high-level semantic features from the input original feature map, focusing on global semantic information related to defects in the medical balloon image. The semantic feature subunit provides more comprehensive feature support for subsequent defect detection.

[0047] The first distribution map can be understood as the output distribution map after the local feature subunit extracts local spatial features from the original feature map. The first distribution map is used for matrix multiplication with the second channel vector to construct the first spatial attention weights, achieving attention weighting at the local feature level. The first channel vector can be understood as representing the importance and response intensity of each feature channel in the local space. The first channel vector is subsequently multiplied with the second distribution map, providing the channel-dimensional feature basis for calculating the second spatial attention weights. The second distribution map can be understood as the output distribution map after the semantic feature subunit extracts features from the original feature map. The second distribution map is used for matrix multiplication with the first channel vector to construct the second spatial attention weights, achieving attention weighting at the semantic level. The second channel vector can be understood as the feature vector obtained by the semantic feature subunit aggregating global semantic features at the channel level from the original feature map, used to represent the importance and response intensity of each feature channel at the global semantic level.

[0048] The first spatial attention weight is a weight matrix obtained by performing matrix multiplication between the first distribution map and the second channel vector. The first spatial attention weight quantifies the importance of each local spatial location in the medical balloon image in terms of local detail and global semantic dimensions. The core function of the first spatial attention weight is to work in conjunction with the second spatial attention weight to jointly constitute the attention weight. The second spatial attention weight is a weight matrix obtained by performing matrix multiplication between the second distribution map and the first channel vector. The second spatial attention weight quantifies the importance of each global semantic spatial location in the medical balloon image in terms of global semantic and local detail dimensions. The core function of the second spatial attention weight is to work in conjunction with the first spatial attention weight to jointly constitute the attention weight. The attention weight is the final weight matrix obtained by fusing the first and second spatial attention weights. The attention weight is a comprehensive quantification result of the importance of defect-related features in the medical balloon image. The attention weight retains the precise weighting of the first spatial attention weight for local texture and edges of defects, while also incorporating the enhancement of the second spatial attention weight for global morphology and semantically related regions of defects. The role of attention weights is to perform matrix multiplication with the original feature map to achieve precise enhancement of the features of the key defect region and effective suppression of irrelevant background information, thereby generating a target enhanced feature map with stronger representation capabilities.

[0049] Specifically, the local feature subunit processes the original feature map to obtain a first distribution map and a first channel vector. The semantic feature subunit processes the original feature map to obtain a second distribution map and a second channel vector. The first distribution map and the second channel vector are then multiplied by a matrix to obtain the first spatial attention weight. The specific formula is shown below: in, The first spatial attention weights; This is the first distribution map; This is the second channel vector; This represents matrix multiplication. The second distribution map is multiplied by the first channel vector to obtain the second spatial attention weights. The specific formula is as follows: in, For second-space attention weights; This is the second distribution map; This is the first channel vector; This represents matrix multiplication. The attention weights are obtained based on the first and second spatial attention weights. The specific formula is shown below: in, Attention weights; This represents the Sigmoid function. The attention weights are multiplied by the original feature map using a matrix dot product to obtain the target augmented feature map. The specific formula is as follows: in, This is the original feature map; Enhance the feature map for the target; This involves matrix dot product. By combining cross-matrix operations and weighted fusion of local and semantic features, the local details and global semantic features of medical balloon defects can be enhanced simultaneously, improving the accuracy and robustness of defect detection.

[0050] Optionally, the local feature subunit includes a first pooling layer, a first convolutional layer, a normalization layer, a second pooling layer, and a first probability distribution layer; the processing of the original feature map based on the local feature subunit to obtain a first distribution map and a first channel vector includes: performing global average pooling on the original feature map in the horizontal direction based on the first pooling layer to obtain a first description vector; and performing global average pooling on the original feature map in the vertical direction based on the first pooling layer to obtain a second description vector; and performing convolution processing on the first description vector and the second description vector respectively based on the first convolutional layer, and performing activation functions on the convolutions respectively. The first description vector and the second description vector after product are processed to obtain a first attention weight vector and a second attention weight vector; the first attention weight vector and the second attention weight vector matrix are multiplied and then matrix-dot multiplied with the original feature map to obtain a spatial augmented feature map; the spatial augmented feature map is normalized based on the normalization layer, and the normalized spatial augmented feature map is processed by the first probability distribution layer to obtain the first distribution map; and the normalized spatial augmented feature map is pooled based on the second pooling layer to obtain the first channel vector.

[0051] In this embodiment of the invention, the first pooling layer is a feature dimensionality reduction layer within the local feature subunit used to perform global average pooling on the original feature map. The first pooling layer extracts global descriptive information of the feature map in the horizontal and vertical directions, providing a foundation for subsequent spatial attention weight calculation. The first description vector is the feature vector obtained by the first pooling layer performing global average pooling on the original feature map in the horizontal direction. The first description vector represents the global statistical information of the original feature map in the horizontal dimension and is one of the basic inputs for subsequent convolution processing and attention weight vector generation. The second description vector is the feature vector obtained by the first pooling layer performing global average pooling on the original feature map in the vertical direction. The second description vector represents the global information of the feature map in the vertical dimension and is one of the basic inputs for subsequent convolution processing and attention weight vector generation.

[0052] The first convolutional layer is a convolutional layer within the local feature subunit used to perform convolution operations on the first and second description vectors. This first convolutional layer is used to further refine features and generate corresponding attention weight vectors. For example, the convolutional kernel of the first convolutional layer can be... The activation function is a function that performs a non-linear transformation on the features output by the first convolutional layer to enhance the model's expressive power and output an effective attention weight vector. For example, this activation function can be the Sigmoid activation function. The first attention weight vector is the vector obtained by convolving the first description vector through the first convolutional layer and then applying the activation function. The first attention weight vector focuses on the feature response distribution in the horizontal direction of the original feature map, enabling adaptive learning and quantification of the importance of features at different positions in the horizontal direction. The first attention weight vector and the second attention weight vector work together to perform matrix multiplication, providing horizontal dimension weight constraints for the subsequent generation of spatially enhanced feature maps, thereby improving the accuracy and robustness of local feature extraction. The second attention weight vector is the vector obtained by convolving the second description vector through the first convolutional layer and then applying the activation function. The second attention weight vector mainly represents the importance of each feature position in the vertical direction of the original feature map. The second attention weight vector and the first attention weight vector work together to complete the weighted constraints on the spatial dimension of the feature map, providing vertical weight support for the generation of spatially enhanced feature maps, further improving the precision and reliability of local feature extraction.

[0053] The spatial augmented feature map is obtained by performing a matrix multiplication of the first attention weight vector and the second attention weight vector, followed by a matrix dot product with the original feature map. The spatial augmented feature map is the core intermediate result of the local feature sub-unit, providing a high-quality, high-discrimination feature foundation for the subsequent generation of the first distribution map and the first channel vector. The normalization layer is a network layer that performs a normalization operation on the spatial augmented feature map. The normalization layer adjusts the feature value distribution of the spatial augmented feature map to a stable range, eliminating interference caused by differences in numerical scales between different feature dimensions. The normalization layer effectively alleviates the gradient vanishing or gradient exploding problems that occur in local feature sub-units after multi-layer operations, accelerating the feature learning process and improving feature robustness. The first probability distribution layer is a network layer that performs probabilistic mapping and feature transformation on the normalized spatial augmented feature map. The first probability distribution layer transforms the absolute values ​​of the spatial augmented feature map into a relative probability distribution, accurately quantifying the relative importance of the medical balloon defect region and the background region. For example, the first probability distribution layer can be a Softmax function. The second pooling layer is a network layer that performs global pooling on the normalized spatial augmented feature map. By performing global average pooling or global max pooling on the spatial augmented feature map, the second pooling layer globally aggregates and compresses the high-dimensional spatial features along the channel dimension, obtaining the first channel vector. The second pooling layer can reduce the feature dimensionality while preserving key feature information, reducing subsequent computation and making the extracted local features more globally representative and stable.

[0054] Specifically, the original feature map is subjected to global average pooling in the horizontal direction based on the first pooling layer to obtain the first description vector. The specific formula is shown below: in, This is the first descriptive vector; This is the original feature map; These represent the number of channels, height, and width of the feature map, respectively; and the feature value at a single location in the original feature map. Let be the three-dimensional coordinates of a single feature value in the feature map, and b, i, and j correspond to the channel, row, and column, respectively. Based on the first pooling layer, global average pooling is performed on the original feature map in the vertical direction to obtain the second description vector. The specific formula is shown below: in, The first description vector is the second description vector. The first convolutional layer convolves the first and second description vectors, and then processes the convolved first and second description vectors using activation functions to obtain the first attention weight vector. Second attention weight vector The spatially enhanced feature map is obtained by multiplying the first attention weight vector and the second attention weight vector matrix together and then performing a matrix dot product with the original feature map. The specific formula is shown below: in, For spatial enhancement feature maps. This represents the Sigmoid function. Represents matrix multiplication. This represents matrix dot product. The spatial augmented feature map is normalized using a normalization layer. The normalized spatial augmented feature map is then processed using a first probability distribution layer to obtain a first distribution map. For example, the first probability distribution layer can be a Softmax function, with the specific formula shown below: The normalized spatial enhancement feature map is pooled using the second pooling layer to obtain the first channel vector. The specific formula is as follows: By employing horizontal and vertical bidirectional pooling, attention-weighted and normalized probabilistic processing, spatial features of defects in medical balloons can be accurately extracted and enhanced, thereby improving feature stability and detection robustness.

[0055] Optionally, the semantic feature subunit includes a second convolutional layer, a second probability distribution layer, and a third pooling layer; the step of processing the original feature map based on the semantic feature subunit to obtain a second distribution map and a second channel vector includes: performing convolution processing on the original feature map based on the second convolutional layer to obtain a convolution result; inputting the convolution result into the second probability distribution layer to obtain the second distribution map; and inputting the convolution result into the third pooling layer to obtain the second channel vector.

[0056] In this embodiment of the invention, the second convolutional layer is a network layer in the semantic feature subunit used to perform convolution operations on the original feature map. The second convolutional layer extracts global semantic information from the feature map through local receptive fields and weight sharing, outputting the convolution result used to subsequently generate the second distribution map and the second channel vector, providing a foundation for semantic feature extraction. The second probability distribution layer is a network layer in the semantic feature subunit used to probabilistically normalize the convolution result output by the second convolutional layer. For example, the second probability distribution layer can be a Softmax function. The second probability distribution layer provides stable and interpretable semantic spatial distribution features for subsequent cross-attention calculations with local features. The third pooling layer is a network layer in the semantic feature subunit that performs global average pooling on the convolution result output by the second convolutional layer. The third pooling layer globally aggregates and compresses the convolution result in the spatial dimension to obtain a fixed-dimensional second channel vector, reducing the feature dimension while preserving global semantic information, providing a compact and stable channel feature representation for subsequent cross-weighting of local features and semantic features.

[0057] Specifically, the original feature map is convolved using the second convolutional layer to obtain the convolution result. For example, if the convolution kernel of the second convolutional layer is... The formula for convolution processing the original feature map is shown below: in, This is the result of the convolution; Indicates the kernel size as The convolution operation is performed; the convolution result is input into the second probability distribution layer to obtain the second distribution map. For example, the second probability distribution layer can be a Softmax function, with the specific formula shown below: The convolution result is input into the third pooling layer to obtain the second channel vector. The specific formula is shown below: By extracting semantic features through the second convolutional layer, normalizing the probability distribution of the second probability distribution layer, and aggregating channel information through global pooling, the semantic expressive power and stability of the features are effectively enhanced.

[0058] S230. Input the text features and the target feature map into the defect recognition module to obtain the defect detection result corresponding to the target medical balloon image.

[0059] Specifically, the text features and target feature map are input into the defect recognition module to obtain the defect detection result corresponding to the target medical balloon image. The defect recognition module analyzes the fused features and outputs whether the target medical balloon image has a defect and the type of defect. The specific structure diagram of the defect detection model is shown below. Figure 5As shown, the defect identification module can include a detection unit, a mask template matching unit, a memory attention unit, a tracking unit, and a memory bank unit. The detection unit receives image features, text features, and geometric cues, segments all defects matching the cues, and outputs the defect mask, bounding box, and confidence score. Geometric cues can be understood as regions of interest specified by the user through points or bounding boxes. The memory attention unit receives the target feature map of the current image and queries historical features from the memory bank, achieving information fusion through an attention mechanism to generate enhanced features. When performing defect detection on multiple images in a time sequence, the tracking unit can generate defect masks for subsequent images based on the enhanced features generated by the memory attention unit. The mask template matching unit performs identity matching between the defect masks detected by the detection unit and the defect masks from the tracking unit, outputting the defect detection result. The memory bank unit acts as a storage queue, continuously storing previously encountered defects. The defect identification module achieves accurate detection of defects in medical balloons through precise localization by the detection unit, historical information provided by the memory unit, consistency maintenance by the tracking unit, and identity conflict resolution by the matching unit.

[0060] By acquiring a target medical balloon image and detection prompt text, the detection prompt text is used to indicate the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. The target medical balloon image is input into a visual encoder to obtain a target feature map, and the detection prompt text is input into a text encoder to obtain text features. By adding an interactive attention unit to the visual encoder, the target feature map extracted by the visual encoder can be accurately captured, effectively suppressing background noise interference. The text features and the target feature map are input into a defect recognition module to obtain a defect detection result corresponding to the target medical balloon image. By inputting the text features and the target feature map together into the defect recognition module, the fusion and joint discrimination of multimodal features are realized, improving the accuracy, robustness, and generalization ability of defect detection. In summary, the technical solution of this embodiment of the invention, by inputting the text features and the target feature map together into the defect recognition module for analysis, solves the problems of low detection efficiency and poor robustness of traditional medical balloon defect detection, and improves the accuracy of medical balloon defect detection.

[0061] Figure 6 This is a schematic diagram of a defect detection device provided in an embodiment of the present invention. This embodiment is applicable to defect detection of medical balloon images. The device can be implemented using software and / or hardware, and can be integrated into any device that provides defect detection functionality, such as… Figure 6 As shown, the defect detection device specifically includes: a data acquisition module 310 and a defect detection result determination module 320.

[0062] The data acquisition module 310 is used to acquire the target medical balloon image to be detected and the detection prompt text. The detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. The defect detection result determination module 320 is used to input the target medical balloon image and the detection prompt text into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder and a defect recognition module. The visual encoder is provided with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image and the type of defect in the medical balloon image when a defect exists.

[0063] By acquiring the target medical balloon image and detection prompt text, the detection prompt text is used to indicate the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image. Acquiring the target medical balloon image and detection prompt text provides a complete data foundation and input conditions for the subsequent inference process of the defect detection model, ensuring the execution of subsequent detection operations. The target medical balloon image and detection prompt text are input into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether the medical balloon image has a defect, and if a defect exists, the type of defect in the medical balloon image. By jointly inputting the target medical balloon image and detection prompt text into the defect detection model, combining image visual features and text prompt information to achieve multimodal joint inference, the accuracy and robustness of defect detection are effectively improved. In summary, the technical solution of this invention analyzes the image of the target medical balloon to be detected using a defect detection model, which solves the problems of low detection efficiency and poor robustness of traditional medical balloon defect detection, and improves the accuracy of defect detection for medical balloons.

[0064] Based on the above embodiments, optionally, the defect detection result determination module 320 is used to: input the target medical balloon image into a visual encoder to obtain a target feature map, and input the detection prompt text into a text encoder to obtain text features; input the text features and the target feature map into a defect recognition module to obtain a defect detection result corresponding to the target medical balloon image.

[0065] Based on the above embodiments, optionally, the visual encoder includes a plurality of serially connected feature extraction modules; at least two of the feature extraction modules include serially connected feature extraction units and interactive attention units, and the interactive attention units in at least two of the feature extraction modules correspond to different feature transformation scales.

[0066] Based on the above embodiments, optionally, the defect detection result determination module 320 is configured to: input the target medical balloon image into the feature extraction module; extract features from the target medical balloon image through the feature extraction unit of the feature extraction module to obtain an original feature map; if the feature extraction module includes the interactive attention unit, perform attention enhancement processing on the original feature map based on the interactive attention unit to obtain a target enhanced feature map, and input the target enhanced feature map into the next feature extraction module; if the feature extraction module does not include the interactive attention unit, input the original feature map into the next feature extraction module; and use the feature map output from the last feature extraction module as the target feature map.

[0067] Based on the above embodiments, optionally, the interactive attention unit includes a local feature subunit and a semantic feature subunit.

[0068] Based on the above embodiments, optionally, the defect detection result determination module 320 is configured to: process the original feature map based on the local feature subunit to obtain a first distribution map and a first channel vector; process the original feature map based on the semantic feature subunit to obtain a second distribution map and a second channel vector; perform matrix multiplication on the first distribution map and the second channel vector to obtain a first spatial attention weight, and perform matrix multiplication on the second distribution map and the first channel vector to obtain a second spatial attention weight; obtain an attention weight based on the first spatial attention weight and the second spatial attention weight; and perform matrix dot product on the attention weight and the original feature map to obtain the target enhancement feature map.

[0069] Based on the above embodiments, optionally, the local feature subunit includes a first pooling layer, a first convolutional layer, a normalization layer, a second pooling layer, and a first probability distribution layer.

[0070] Based on the above embodiments, optionally, the defect detection result determination module 320 is further configured to: perform global average pooling on the original feature map in the horizontal direction based on the first pooling layer to obtain a first description vector; and perform global average pooling on the original feature map in the vertical direction based on the first pooling layer to obtain a second description vector; perform convolution processing on the first description vector and the second description vector respectively based on the first convolution layer, and process the convolved first description vector and the second description vector respectively based on the activation function to obtain a first attention weight vector and a second attention weight vector; multiply the first attention weight vector and the second attention weight vector matrix and then perform matrix dot product with the original feature map to obtain a spatial augmented feature map; perform normalization processing on the spatial augmented feature map based on the normalization layer, and process the normalized spatial augmented feature map based on the first probability distribution layer to obtain a first distribution map; and perform pooling processing on the normalized spatial augmented feature map based on the second pooling layer to obtain the first channel vector.

[0071] Based on the above embodiments, optionally, the semantic feature subunit includes a second convolutional layer, a second probability distribution layer, and a third pooling layer.

[0072] Based on the above embodiments, optionally, the defect detection result determination module 320 is further configured to: perform convolution processing on the original feature map based on the second convolution layer to obtain a convolution result; input the convolution result into the second probability distribution layer to obtain the second distribution map; and input the convolution result into the third pooling layer to obtain the second channel vector.

[0073] Optionally, based on the above embodiments, the device further includes a model training module, configured to: acquire an original defect image dataset, the original defect image dataset including multiple sample medical balloon images, wherein the multiple sample medical balloon images include normal medical balloon images and defective medical balloon images, and the defective medical balloon images include at least one defective region; perform data augmentation processing on at least some of the sample medical balloon images in the original defect image dataset to obtain extended medical balloon images; add the extended medical balloon images to the original defect image dataset to obtain a target defect image dataset; and train the model based on the target defect image dataset to obtain the defect detection model.

[0074] Optionally, based on the above embodiments, the model training module is further configured to: annotate at least one defective region in multiple defective medical balloon images to obtain multiple defect-annotated images; generate multiple mask images corresponding to the multiple defective regions based on the multiple defect-annotated images, and use the minimum bounding rectangle of the mask images as a foreground patch; when the intersection-union ratio (IUU) of the foreground patch and the defective region in the image region corresponding to the defective medical balloon image is less than a first preset threshold, add the foreground patch to the effective imaging region of at least some of the sample medical balloon images to obtain an enhanced defective image, wherein the effective imaging region is the region where the medical balloon is located in the image; determine the area ratio of the pixel area of ​​the defective region in the enhanced defective image to the pixel area of ​​the defective region in the mask image, and determine the enhanced defective image with the area ratio greater than or equal to a second preset threshold as an extended medical balloon image.

[0075] The above-described products can perform the methods provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for performing the methods.

[0076] Figure 7 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0077] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0078] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0079] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as defect detection methods.

[0080] In some embodiments, the defect detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the defect detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the defect detection method by any other suitable means (e.g., by means of firmware).

[0081] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0082] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0083] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0085] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0086] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0087] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0088] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the defect detection method according to any embodiment of the invention.

[0089] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0090] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A defect detection method, characterized in that, include: Acquire the target medical balloon image to be detected and the detection prompt text, wherein the detection prompt text is used to prompt the image content in the target medical balloon image that needs to be detected by the defect detection model and / or the defect detection result that the defect detection model needs to output for the target medical balloon image; The target medical balloon image and the detection prompt text are input into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image. The defect detection model includes a visual encoder, a text encoder, and a defect recognition module. The visual encoder is equipped with an interactive attention unit. The defect detection result is used to indicate whether there is a defect in the medical balloon image and the type of defect in the medical balloon image when a defect exists.

2. The method according to claim 1, characterized in that, The step of inputting the target medical balloon image and detection prompt text into the defect detection model to obtain the defect detection result corresponding to the target medical balloon image includes: The target medical balloon image is input into a visual encoder to obtain a target feature map, and the detection prompt text is input into a text encoder to obtain text features; The text features and the target feature map are input into the defect recognition module to obtain the defect detection result corresponding to the target medical balloon image.

3. The method according to claim 2, characterized in that, The visual encoder includes multiple serially connected feature extraction modules; at least two of the feature extraction modules include serially connected feature extraction units and interactive attention units, and the interactive attention units in at least two of the feature extraction modules correspond to different feature transformation scales; The step of inputting the target medical balloon image into a visual encoder to obtain a target feature map includes: The target medical balloon image is input into the feature extraction module, and the feature extraction unit of the feature extraction module performs feature extraction on the target medical balloon image to obtain the original feature map; When the feature extraction module includes the interactive attention unit, the original feature map is subjected to attention enhancement processing based on the interactive attention unit to obtain a target enhanced feature map, and the target enhanced feature map is input into the next feature extraction module; If the feature extraction module does not include the interactive attention unit, the original feature map is input into the next feature extraction module. The feature map output from the last feature extraction module is used as the target feature map.

4. The method according to claim 3, characterized in that, The interactive attention unit includes a local feature subunit and a semantic feature subunit; The step of performing attention enhancement processing on the original feature map based on the interactive attention unit to obtain the target enhanced feature map includes: The original feature map is processed based on the local feature sub-units to obtain a first distribution map and a first channel vector; The original feature map is processed based on the semantic feature subunit to obtain a second distribution map and a second channel vector; The first distribution map and the second channel vector are multiplied by a matrix to obtain the first spatial attention weight, and the second distribution map and the first channel vector are multiplied by a matrix to obtain the second spatial attention weight. The attention weights are obtained based on the first spatial attention weight and the second spatial attention weight; The attention weights are multiplied by the original feature map to obtain the target enhancement feature map.

5. The method according to claim 4, characterized in that, The local feature subunit includes a first pooling layer, a first convolutional layer, a normalization layer, a second pooling layer, and a first probability distribution layer; The process of processing the original feature map based on the local feature sub-units to obtain a first distribution map and a first channel vector includes: The original feature map is subjected to global average pooling in the horizontal direction based on the first pooling layer to obtain a first description vector, and the original feature map is subjected to global average pooling in the vertical direction based on the first pooling layer to obtain a second description vector. The first description vector and the second description vector are convolved based on the first convolutional layer, and the convolved first description vector and the second description vector are processed based on the activation function to obtain the first attention weight vector and the second attention weight vector. Multiply the first attention weight vector and the second attention weight vector matrix and then perform a matrix dot product with the original feature map to obtain a spatially enhanced feature map; The spatial augmented feature map is normalized based on the normalization layer, and the normalized spatial augmented feature map is processed based on the first probability distribution layer to obtain the first distribution map; and... The normalized spatial enhancement feature map is pooled based on the second pooling layer to obtain the first channel vector.

6. The method according to claim 4, characterized in that, The semantic feature subunit includes a second convolutional layer, a second probability distribution layer, and a third pooling layer; The process of processing the original feature map based on the semantic feature subunit to obtain a second distribution map and a second channel vector includes: The original feature map is convolved based on the second convolutional layer to obtain the convolution result; The convolution result is input into the second probability distribution layer to obtain the second distribution map, and the convolution result is input into the third pooling layer to obtain the second channel vector.

7. The method according to claim 1, characterized in that, The training process of the defect detection model includes: Obtain an original defect image dataset, which includes multiple sample medical balloon images, and the multiple sample medical balloon images include normal medical balloon images and defective medical balloon images, wherein the defective medical balloon images include at least one defective region; At least a portion of the medical balloon images in the original defect image dataset are subjected to data augmentation processing to obtain extended medical balloon images. The extended medical balloon images are then added to the original defect image dataset to obtain the target defect image dataset. The defect detection model is obtained by training the model based on the target defect image dataset.

8. The method according to claim 7, characterized in that, At least a portion of the medical balloon images in the original defective image dataset are subjected to data augmentation processing to obtain expanded medical balloon images, including: At least one defective region in multiple images of the defective medical balloon is labeled to obtain multiple defect-labeled images; Multiple mask images corresponding to multiple defect regions are generated based on multiple defect-annotated images, and the minimum bounding rectangle of the mask images is used as the foreground patch. When the intersection-union ratio of the foreground patch and the defective region in the image region corresponding to the defective medical balloon image is less than a first preset threshold, the foreground patch is added to the effective imaging region of at least a portion of the sample medical balloon images to obtain an enhanced defective image, wherein the effective imaging region is the region where the medical balloon in the image is located. The ratio of the pixel area of ​​the defect region in the enhanced defect image to the pixel area of ​​the defect region in the mask image is determined, and the enhanced defect image with the ratio greater than or equal to a second preset threshold is determined as an expanded medical balloon image.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the defect detection method according to any one of claims 1-8.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the defect detection method according to any one of claims 1-8.