A bronchoscope image recognition classification method, device, equipment and medium

By adding adversarial patches to bronchoscopic images and performing ablation training, and utilizing the coordinate attention and self-attention modules in the visual transformer, the misjudgment problem in bronchoscopic image recognition is solved, achieving higher recognition accuracy and robustness.

CN117218651BActive Publication Date: 2025-10-17SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311243649.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-10-17
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

Existing bronchoscopic image recognition technology is easily interfered with by other tissues in the diagnosis of peripheral lung lesions, leading to misjudgment. It also has low sampling efficiency and severe trauma. How to improve recognition accuracy is an urgent problem to be solved.

Method used

By obtaining historical images collected by bronchoscope, adding adversarial patches to form variants, performing preprocessing and ablation training, and using visual converters for training, a smooth classifier is obtained, which can resist the recognition interference of adversarial patches and improve classification accuracy.

Benefits of technology

The recognition accuracy of bronchoscopic images is improved, the anti-interference ability to adversarial patches is enhanced, and the robustness and accuracy of recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218651B_ABST
    Figure CN117218651B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and particularly relates to a bronchoscope image recognition and classification method, device, equipment and medium, the method comprising: acquiring N historical images collected by a bronchoscope; adding an adversarial patch to part of the N historical images to obtain a variant for each historical image in the part of the historical images, and other historical images except the part of the historical images are non-variants; preprocessing the non-variants and the variants to obtain initial images; performing ablation training on the initial images to obtain training data; training a visual converter based on the training data to obtain a smoothing classifier, the smoothing classifier is used for accurate classification of the historical images and can resist recognition interference caused by the adversarial part; obtaining a target image to be recognized; obtaining a recognition and classification result of the target image based on the target image and the smoothing classifier, and improving recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a bronchoscope image recognition and classification method, device, equipment and medium. BACKGROUND

[0002] Lung cancer is a malignant tumor, and early diagnosis and treatment is the key to improving the survival of lung cancer patients.

[0003] Although the introduction of bronchoscopic forceps biopsy, percutaneous biopsy and surgical lung biopsy can obtain specimens, the number of samples obtained is small, the efficiency is low, the trauma is large, and there is a risk of pneumothorax, so more empirical diagnosis is used in clinical practice.

[0004] This year, virtual navigation assisted ultrasound electronic bronchoscopy with guide sheath has been introduced into clinical practice. Compared with the previous method, although the sampling efficiency and diagnostic efficiency can be improved, peripheral lung lesions lack specific imaging manifestations and are easily disturbed by other tissues, leading to misjudgment of the system.

[0005] Therefore, how to effectively avoid these misjudgments and improve the accuracy of system recognition is a technical problem that needs to be solved at present. SUMMARY

[0006] In view of the above problems, the present application provides a bronchoscope image recognition and classification method, device, equipment and medium which can overcome the above problems or at least partially solve the above problems.

[0007] In a first aspect, the present application provides a bronchoscope image recognition and classification method, comprising:

[0008] Obtaining N historical images collected by a bronchoscope;

[0009] Adding an adversarial patch to part of the N historical images to obtain a variant for each historical image in the part of the historical images, and other historical images except the part of the historical images are non-variants, the adversarial patch is used to cause identification interference to the part of the historical images;

[0010] Preprocessing the non-variants and the variants respectively to obtain initial images;

[0011] Using ablation training to process the initial images to obtain training data;

[0012] Training a visual converter based on the training data to obtain a smooth classifier, the smooth classifier is used to accurately classify the historical images and resist identification interference caused by the adversarial patch, wherein the visual converter includes a coordinate attention module and a self-attention module;

[0013] obtaining a target image to be recognized;

[0014] obtaining a recognition classification result of the target image based on the target image and the smooth classifier.

[0015] Preferably, the pre-processing of the non-variant and the variant to obtain the initial image comprises:

[0016] performing standard pre-processing of gray value normalization, center clipping and resampling on the non-variant and the variant respectively to obtain the initial image.

[0017] Preferably, the standard pre-processing of gray value normalization, center clipping and resampling on the non-variant and the variant respectively to obtain the initial image comprises:

[0018] performing gray value normalization processing on the non-variant and the variant respectively to obtain a normalized picture, so that the pixel gray value of the normalized picture is distributed between 0 and 255;

[0019] clipping a picture of a preset length and a preset width around the center position from the center position of the normalized picture to obtain a clipped picture;

[0020] interpolating the clipped picture using the sampled points to obtain the initial image.

[0021] Preferably, the ablation training processing of the initial image to obtain the training data comprises:

[0022] cutting each initial image to obtain a plurality of row regions or a plurality of column regions;

[0023] labeling the plurality of row regions or the plurality of column regions respectively to obtain a label token of each row region or each column region, wherein a first total area of a row region or a column region where the adversarial patch exists is less than a second total area of a row region or a column region where the adversarial patch does not exist, and the label token is used as training data, and when training with the label token of any row region or column region, the label tokens of other row regions or column regions are shielded.

[0024] Preferably, the training of the visual converter based on the training data to obtain the smooth classifier, wherein the smooth classifier is used for accurate classification of the historical image and can resist recognition interference caused by the adversarial patch, and the visual converter comprises a coordinate attention module and a self-attention module, comprising:

[0025] inputting the training data into the classifier to obtain a first classification result for each historical image;

[0026] train the visual transformer based on the training data and the first classification result, wherein the coordinate attention module is configured to enhance features of the training data, the self-attention module is configured to extract the features to obtain a smooth classifier, the smooth classifier is configured to accurately classify the historical images and resist recognition interference caused by the adversarial patches, and an output of the coordinate attention module is connected to an input of the self-attention module.

[0027] Preferably, the smooth classifier is further configured to accurately classify the historical images and resist recognition interference caused by adversarial samples.

[0028] In a second aspect, the present application further provides a bronchoscope image recognition and classification device, comprising:

[0029] A first obtaining module is configured to obtain N historical images collected by a bronchoscope.

[0030] A first obtaining module is configured to obtain N historical images collected by a bronchoscope.

[0031] A second obtaining module is configured to pre-process the non-variant and the variant respectively to obtain initial images.

[0032] A third obtaining module is configured to perform ablation training processing on the initial images to obtain training data.

[0033] A fourth obtaining module is configured to train a visual transformer based on the training data to obtain a smooth classifier, wherein the smooth classifier is configured to accurately classify the historical images and resist recognition interference caused by the adversarial patches, and the visual transformer comprises a coordinate attention module and a self-attention module.

[0034] A second obtaining module is configured to obtain a target image to be recognized.

[0035] A fifth obtaining module is configured to obtain a recognition and classification result of the target image based on the target image and the smooth classifier.

[0036] In a third aspect, the present application further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to realize the method steps of the first aspect.

[0037] In a fourth aspect, the present application further provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method steps of the first aspect.

[0038] The one or more technical solutions in the embodiments of the present application have at least the following technical effects or advantages:

[0039] The present application provides a bronchoscope image recognition classification method, which comprises the following steps: acquiring N historical images collected by a bronchoscope; adding an adversarial patch to part of the N historical images to obtain a variant for each historical image in the part of the historical images, and other historical images except the part of the historical images are non-variants, the adversarial part is used to cause recognition interference to the part of the historical images; pre-processing the non-variants and the variants to obtain initial images; performing ablation training processing on the initial images to obtain training data; training a visual converter based on the training data to obtain a smooth classifier, which is used for accurate classification of the historical images and can resist the recognition interference caused by the adversarial part, wherein the visual converter comprises a coordinate attention module and a self-attention module; acquiring a target image to be recognized; obtaining a recognition classification result of the target image based on the target image and the smooth classifier; by actively adding an adversarial patch, increasing the number of interference samples, and processing the initial images by using an ablation training processing mode to obtain training data, the anti-interference capability to the adversarial patch can be improved in the training, and the visual converter used comprises a coordinate attention module and a self-attention module, and the recognition accuracy is improved by enhancing the features. BRIEF DESCRIPTION OF DRAWINGS

[0040] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. Furthermore, the same reference numerals in different drawings represent the same or similar components. In the drawings:

[0041] Figure 1 A step flowchart of the bronchoscope image recognition classification method in the embodiments of the present application is shown;

[0042] Figure 2 A structure diagram of the self-attention module in the embodiments of the present application is shown;

[0043] Figure 3 A structure diagram of the compression-excitation network in the embodiments of the present application is shown;

[0044] Figure 4 A structure diagram of the compression-excitation network adding position information in the embodiments of the present application is shown;

[0045] Figure 5 a schematic diagram showing the overall logical idea in the embodiment of the present application is shown;

[0046] Figure 6 a schematic diagram showing the recognition and classification device of the bronchoscope image in the embodiment of the present application is shown;

[0047] Figure 7 a structural schematic diagram of a computer device for implementing the recognition and classification method of the bronchoscope image in the embodiment of the present application is shown. DETAILED DESCRIPTION

[0048] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art.

[0049] Embodiment One

[0050] The embodiment of the present application provides a recognition and classification method of a bronchoscope image, as shown in the method comprises the following steps: Figure 1 as shown, comprising:

[0051] S101, acquiring N historical images collected by a bronchoscope;

[0052] S102, adding an adversarial patch to part of the historical images to obtain a variant for each historical image in the part of the historical images, and other historical images except the part of the historical images are non-variants, and the adversarial patch is used to cause identification interference to the part of the historical images;

[0053] S103, pre-processing the non-variants and the variants to obtain initial images;

[0054] S104, using ablation training processing on the initial images to obtain training data;

[0055] S105, training a visual converter based on the training data to obtain a smooth classifier, the smooth classifier is used for accurate classification of the historical images and can resist identification interference caused by the adversarial part, wherein the visual converter comprises a coordinate attention module and a self-attention module;

[0056] S106, acquiring a target image to be recognized;

[0057] S107, obtaining a recognition and classification result of the target image based on the target image and the smooth classifier.

[0058] First, S101, acquiring N historical images collected by a bronchoscope.

[0059] Specifically, the bronchoscope is inserted from the mouth or nose of the patient and reaches the bronchus to obtain image information, the number of historical image information samples is small, and the number of samples with adversarial patches is even smaller.

[0060] Therefore, it is necessary to manually increase samples with interference. Therefore, S102 is performed to increase adversarial patches to part of the historical images in N historical images to obtain variants for each historical image in the part of historical images, and other historical images except the part of historical images are non-variants, and the adversarial patches are used to cause identification interference to the part of historical images.

[0061] Specifically, for the part of historical images, part of the area of each historical image is cut out to obtain a blank area, and then a similar image with the same size as the blank area is used to patch the blank area, and the similar historical image can be distinguished by the naked eye or cannot be distinguished by the naked eye, thereby obtaining a variant of the original historical image. The adversarial patch can cause interference to identification. For other historical images except the part of historical images, they are non-variants.

[0062] After obtaining the variants and non-variants, S103 is performed to pre-process the non-variants and variants respectively to obtain initial images.

[0063] In an optional embodiment, the non-variants and variants are respectively subjected to standard preprocessing of gray value normalization, center cropping and resampling to obtain initial images. The preprocessing is a standard preprocessing process, and of course, there are other deformation processes, which are not limited herein.

[0064] Specifically, the non-variants and variants are respectively subjected to gray value normalization processing to obtain normalized pictures, so that the pixel gray value of the normalized picture is distributed between 0 and 255, thereby avoiding insufficient contrast of the picture and bringing interference to subsequent processing.

[0065] The normalized picture is cropped from the center position to obtain a picture with a preset length and a preset width around the center position to obtain a cropped picture. By cropping pictures with the same size, subsequent processing is facilitated.

[0066] The cropped picture is interpolated using the sampled points to obtain an initial image, which improves the sampling efficiency and does not affect subsequent processing.

[0067] Next, S104 is performed to process the initial images using ablation training to obtain training data.

[0068] Specifically, each initial image is cut to obtain a plurality of row regions or a plurality of column regions.

[0069] The plurality of row-oriented regions or the plurality of column-oriented regions are respectively marked to obtain a marking token of each row-oriented region or each column-oriented region, wherein a first area sum of the row-oriented region or the column-oriented region in which the adversarial part exists is less than a second area sum of the row-oriented region or the column-oriented region in which the adversarial patch does not exist, and the marking token is used as training data, and when training is performed on the marking token of any row-oriented region or column-oriented region, the marking tokens of other row-oriented regions or column-oriented regions are shielded.

[0070] In a specific embodiment, an initial image is divided into a plurality of regions, which can be row-oriented regions or column-oriented regions, and are respectively numbered to obtain a marking token of each row-oriented region or each column-oriented region. The column-oriented region is columnar segmentation, and a historical image is divided into a plurality of columnar small images. Since the range of the region in which the adversarial patch exists is small, and the range of the region in which the non-adversarial patch exists is large, when training is performed on any row-oriented region or column-oriented region as training data, other regions are shielded, and if the any row-oriented region is the region in which the non-adversarial patch exists, the region in which the adversarial part exists is shielded.

[0071] The marking method is adopted, and the marking tokens of other regions are shielded, so that the marking tokens are suitable for subsequent processing, the complexity of data is reduced, and the processing efficiency is improved.

[0072] Next, training is performed, that is, S105 is executed, the visual converter is trained based on the training data to obtain a smoothing classifier, and the smoothing classifier is used for accurate classification of the historical image and can resist recognition interference caused by the adversarial patch, wherein the visual converter includes a coordinate attention module and a self-attention module.

[0073] Specifically, the training data is input into the classifier to obtain a first classification result for each historical image.

[0074] The visual converter is trained based on the training data and the first classification result, wherein the coordinate attention module is used to enhance the features of the training data, the self-attention module is used to extract the features, a smoothing classifier is obtained, the smoothing classifier is used for accurate classification of the historical image and can resist recognition interference caused by the adversarial patch, and the output of the coordinate attention module is connected to the input of the self-attention module.

[0075] For the variant, the segmentation into multiple row regions or column regions is input into the classifier, and the first classification result can be determined according to the classification results of the multiple row regions or column regions. Since the range of the row region or the column region of the adversarial patch is small, the classification result of the comprehensive variant ignores the classification result corresponding to the adversarial part, so that the first classification result can ignore the classification result corresponding to the adversarial patch. Further improve the robustness of identification.

[0076] The visual converter used is described below:

[0077] In the present application, the visual converter is used instead of a convolutional network as the main body, mainly using a self-attention module that can ignore the masked region. The self-attention module, as shown in Figure 2 , creates three vectors for the input training data, including a query vector, a key vector, and an input element value vector. The self-attention generated from the query vector and the key vector can capture long-range dependencies and adaptability.

[0078] The present application mainly uses a self-attention module with channel attention mechanism, i.e. a compression-activation network, as shown in Figure 3 , which is divided into compression and activation parts. The purpose of the compression part is to compress the global spatial information and learn features in the channel dimension to form the importance of each channel. Finally, the activation part assigns different weights to each channel.

[0079] Using this self-attention module, the label tokens of the above-mentioned masked region can be ignored and not processed.

[0080] Since the compression-activation network only applies weights to the channels and ignores the position information, it only considers the encoding of the channel information and ignores the importance of the position information, but the position information is actually crucial for visual tasks that need to capture target structures. In order to alleviate the loss of position information caused by two-dimensional global pooling, the channel attention is decomposed into two parallel one-dimensional feature encodings, and the spatial coordinate information (vertical and horizontal directions) is aggregated into two independent direction perception feature attention maps, which capture long-range dependencies along this direction. Finally, the position information is saved and used to enhance the expression ability of the feature map. As shown in Figure 4 .

[0081] In order to improve the accuracy of identification, the visual converter further comprises a coordinate attention module (CA) located before the input of the self-attention module, which is used to enhance the feature expression ability.

[0082] For a visual classifier, nc (x) > max c′≠c n c′ (x) + 2D

[0083] Wherein, let D represent the number of adversarial patches that an adversarial patch can intersect at most, for the adversarial patch of the column region with the width of b, mxm adversarial patch, D = m + b + 1, and n c When (x) reaches such a threshold, the most frequent classification will be guaranteed not to change, that is, the adversarial patch destroys each ablation it intersects, n c (x) is the total number of samples whose classification is c. The smoothed classifier obtained in this way can accurately and reliably predict.

[0084] Although the present application has the effect of robust recognition of adversarial patches, which is beneficial to improve the accuracy of recognition, the present application is also applicable to adversarial samples, and can accurately classify images and resist the recognition interference caused by adversarial samples.

[0085] When the size of the adversarial patch meets the preset size, the accuracy rate of the same type of framework, such as mobile-net v2, and the framework with compression-excitation network enhancement robustness, increases by about 1% ~ 2%, about 74.6%, and the smaller the size of the adversarial patch, the higher the accuracy rate.

[0086] The above is obtained by training the smoothed classifier, and then S106 and S107 are executed to apply the smoothed classifier.

[0087] S106, obtaining a target image to be recognized, of course, the target image is the image collected by the bronchoscope, then inputting the target image into the smoothed classifier to obtain the recognition classification result of the target image.

[0088] Of course, during the processing, the target image still needs to be cut to obtain multiple row regions or column regions of the target image, and then be marked with tokens and input into the smoothed classifier to finally obtain the classification result.

[0089] The above processing method and smoothed classifier can improve the recognition accuracy.

[0090] As Figure 5 shown, the overall logical idea of the present application. By inputting N historical images, then adding adversarial patches to part of the N historical images, then preprocessing the variants and non-variants, then performing ablation training, tokenization and other operations, inputting the visual converter for training, thereby obtaining a smoothed classifier, the visual converter includes a coordinate attention module (CA) and a self-attention module, and finally applying the smoothed classifier.

[0091] The one or more technical solutions in the embodiments of the present application have at least the following technical effects or advantages:

[0092] The present application provides a bronchoscope image recognition classification method, which comprises the following steps: acquiring N historical images collected by a bronchoscope; adding an adversarial patch to part of the N historical images to obtain a variant for each historical image in the part of the historical images, and the other historical images except the part of the historical images are non-variants, and the adversarial part is used to cause recognition interference to the part of the historical images; pre-processing the non-variants and the variants to obtain initial images; using ablation training processing on the initial images to obtain training data; training a visual converter based on the training data to obtain a smooth classifier, which is used for accurate classification of the historical images and can resist the recognition interference caused by the adversarial part, wherein the visual converter comprises a coordinate attention module and a self-attention module; acquiring a target image to be recognized; obtaining a recognition classification result of the target image based on the target image and the smooth classifier; by actively adding an adversarial patch, increasing the number of interference samples, and using ablation training processing on the initial images to obtain training data, the anti-interference ability to the adversarial patch can be improved in the training, and the visual converter used comprises a coordinate attention module and a self-attention module, which improves the recognition accuracy by enhancing the features.

[0093] Embodiment two

[0094] Based on the same inventive concept, the present application also provides a bronchoscope image recognition classification device, as shown in Figure 6 , which comprises:

[0095] The first acquisition module 601 is configured to acquire N historical images collected by a bronchoscope.

[0096] The first obtaining module 602 is configured to add an adversarial patch to part of the N historical images to obtain a variant for each historical image in the part of the historical images, and the other historical images except the part of the historical images are non-variants, and the adversarial patch is used to cause recognition interference to the part of the historical images.

[0097] The second obtaining module 603 is configured to pre-process the non-variants and the variants respectively to obtain initial images.

[0098] The third obtaining module 604 is configured to use ablation training processing on the initial images to obtain training data.

[0099] The fourth obtaining module 605 is configured to train a visual transformer based on the training data to obtain a smooth classifier, the smooth classifier being configured to accurately classify the historical image and resist recognition interference caused by the adversarial patch.

[0100] The second obtaining module 606 is configured to obtain a target image to be recognized.

[0101] The fifth obtaining module 607 is configured to obtain a recognition classification result of the target image based on the target image and the smooth classifier.

[0102] In an optional implementation, the second obtaining module 603 is configured to:

[0103] The non-variant and the variant are respectively subjected to standard preprocessing of grayscale value normalization, center clipping and resampling to obtain initial images.

[0104] In an optional implementation, the second obtaining module 603 is specifically configured to:

[0105] The non-variant and the variant are respectively subjected to grayscale value normalization processing to obtain normalized pictures, so that pixel grayscale values of the normalized pictures are distributed between 0 and 255.

[0106] The normalized pictures are clipped from a center position to obtain clipping pictures with a preset length and a preset width around the center position.

[0107] The clipping pictures are interpolated by using sampled points to obtain the initial images.

[0108] In an optional implementation, the third obtaining module 604 is configured to:

[0109] Each initial image is cut to obtain a plurality of row regions or a plurality of column regions.

[0110] The plurality of row regions or the plurality of column regions are respectively labeled to obtain a label token of each row region or each column region, wherein a first total area of a row region or a column region in which the adversarial patch exists is less than a second total area of a row region or a column region in which the adversarial patch does not exist, and the label token is taken as training data, and when the label token of any row region or column region is trained, the label tokens of other row regions or column regions are shielded.

[0111] In an optional implementation, the fourth obtaining module 605 is configured to:

[0112] Inputting the training data into a classifier to obtain a first classification result for each historical image;

[0113] Based on the training data and the first classification result, the visual converter is trained, wherein the coordinate attention module is used to enhance the features of the training data, and the self-attention module is used to extract the features to obtain a smooth classifier. The smooth classifier is used to accurately classify the historical image and can resist the recognition interference caused by the adversarial patch, and the output of the coordinate attention module is connected to the input of the self-attention module.

[0114] In an optional embodiment, the smoothing classifier is further used to accurately classify the historical images and resist recognition interference caused by adversarial samples.

[0115] Example 3

[0116] Based on the same inventive concept, an embodiment of the present invention provides a computer device, such as Figure 7 As shown, it includes a memory 704, a processor 702 and a computer program stored in the memory 704 and executable on the processor 702. When the processor 702 executes the program, the steps of the above-mentioned bronchoscopic image recognition and classification method are implemented.

[0117] Among them, Figure 7 In the embodiment of the present invention, a bus architecture (represented by bus 700) is shown. Bus 700 may include any number of interconnected buses and bridges, and bus 700 links together various circuits including one or more processors represented by processor 702 and memory represented by memory 704. Bus 700 may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 706 provides an interface between bus 700 and receiver 701 and transmitter 703. Receiver 701 and transmitter 703 may be the same component, namely a transceiver, which provides a unit for communicating with various other devices over a transmission medium. Processor 702 is responsible for managing bus 700 and general processing, while memory 704 may be used to store data used by processor 702 when performing operations.

[0118] Example 4

[0119] Based on the same inventive concept, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned bronchoscopic image recognition and classification method.

[0120] The algorithms and displays presented herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with programs in accordance with the teachings herein, or it can prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description above. In addition, the present application is not intended to be limited to any particular programming language. It will be appreciated that there are many programming languages that can be used to implement the teachings herein, and any specific language can be chosen for use in this application.

[0121] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description.

[0122] Similarly, it is to be understood that the phraseology or terminology employed herein, and not otherwise specifically set forth in this specification, is for the purpose of description only and not of limitation. Rather, the disclosed aspects will be understood to apply to any suitable object having the same or similar structure or function.

[0123] Those skilled in the art will appreciate that the modules in the apparatuses in the embodiments can be adapted and placed in one or more apparatuses other than the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into more sub-modules or sub-units or sub-components. Any combination of all the features disclosed in the specification (including the accompanying claims, abstract and drawings), and any method or of the apparatus disclosed in the specification, can be taken, except that at least some of such features and / or processes or units are mutually exclusive, unless explicitly stated otherwise. Each feature disclosed in the specification (including the accompanying claims, abstract and drawings) can be replaced by alternative features serving the same, equivalent or similar purpose, unless explicitly stated otherwise.

[0124] Furthermore, those skilled in the art will recognize that, while certain embodiments described herein include certain features that are not included in other embodiments, combinations of features of the different embodiments are to be construed as being within the scope of the application and forming different embodiments. For example, in the DETAILED DESCRIPTION, any of the claimed embodiments can be used in any combination.

[0125] The various component embodiments of the application can be implemented in hardware, or as software modules running in one or more processors, or in combinations thereof. Those skilled in the art will appreciate that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of the bronchoscope image recognition classification apparatus, some or all of the components of the computer device according to the embodiments of the application. The application can also be implemented as a program (e.g., computer program and computer program product) for performing part or all of the methods described herein. Such program implementing the application can be stored on a computer readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier medium, or in any other form.

[0126] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps other than those listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In a unitary claim, several devices, apparatuses or means can be listed, comprising means for performing a certain function. The functions of the separate means can be carried out by one specific means performing the functions. Similarly, a device recited as meaning for performing a certain function can be interpreted to mean a specific means performing the function. The words 'first','second' and 'third' etc. do not imply any order but are used for naming purposes only.

Claims

1. A bronchoscopic image recognition and classification method, characterized in that: include: Obtain N historical images collected by bronchoscope; Adding adversarial patches to some of the N historical images to obtain variants of each of the historical images, with the other historical images other than the some historical images being non-variants, the adversarial patches being used to cause recognition interference to the some historical images. Adding adversarial patches to some of the N historical images includes: cutting out a portion of each of the historical images to obtain a blank area; and patching the blank area with a similar image of the same size as the blank area. Preprocessing the non-variant and the variant respectively to obtain initial images; The initial image is subjected to ablation training processing to obtain training data, including: Cut each initial image to obtain a plurality of row-oriented regions or a plurality of column-oriented regions; Marking the multiple row regions or the multiple column regions respectively to obtain a marking token for each row region or each column region, wherein a first sum of the areas of the row regions or column regions where the adversarial patch exists is smaller than a second sum of the areas of the row regions or column regions where the adversarial portion does not exist, and using the marking tokens as training data. When training with the marking tokens of any row region or column region, the marking tokens of other row regions or column regions are masked. Based on the training data, the visual converter is trained to obtain a smoothing classifier, wherein the smoothing classifier is used to accurately classify the historical image and resist recognition interference caused by the adversarial patch, wherein the visual converter includes a coordinate attention module and a self-attention module, including: Inputting the training data into a classifier to obtain a first classification result for each historical image; Based on the training data and the first classification result, the visual converter is trained, wherein the coordinate attention module is used to enhance features of the training data, and the self-attention module is used to extract the features to obtain a smoothing classifier, wherein the smoothing classifier is used to accurately classify the historical image and resist recognition interference caused by the adversarial patch, and the output of the coordinate attention module is connected to the input of the self-attention module; Acquire the target image to be identified; Based on the target image and the smoothing classifier, a recognition and classification result of the target image is obtained.

2. The method according to claim 1, wherein The preprocessing of the non-variant and the variant to obtain an initial image includes: Standard preprocessing of grayscale normalization, center cropping, and resampling is performed on the non-variant and the variant, respectively, to obtain an initial image.

3. The method according to claim 2, wherein The step of performing standard preprocessing of grayscale normalization, center cropping, and resampling on the non-variant and the variant to obtain an initial image includes: performing grayscale normalization processing on the non-variant and the variant respectively to obtain a normalized image, so that the pixel grayscale values ​​of the normalized image are distributed between 0 and 255; Starting from the center position of the normalized image, cropping the image with a preset length and preset width around the center position to obtain a cropped image; The cropped image is interpolated using the sampled points to obtain an initial image.

4. The method according to claim 1, wherein The smoothing classifier is also used to accurately classify the historical images and can resist recognition interference caused by adversarial samples.

5. A bronchoscopic image recognition and classification device, characterized in that: include: A first acquisition module is used to acquire N historical images collected by a bronchoscope; A first obtaining module is configured to add an adversarial patch to a portion of the N historical images to obtain a variant of each of the portion of the historical images, wherein the other historical images except the portion of the historical images are non-variants, and the adversarial patch is configured to cause recognition interference to the portion of the historical images. The first obtaining module is specifically configured to: Cutting out a portion of each historical image from the partial historical images to obtain a blank area; patching the blank area with a similar image having the same size as the blank area; A second obtaining module is used to preprocess the non-variant and the variant respectively to obtain an initial image; a third obtaining module for applying ablation training processing to the initial image to obtain training data, wherein the third obtaining module is used to cut each initial image to obtain a plurality of row regions or a plurality of column regions; respectively mark the plurality of row regions or the plurality of column regions to obtain a marking token for each row region or each column region, wherein a first sum of the areas of the row regions or the column regions where the adversarial patch exists is less than a second sum of the areas of the row regions or the column regions where the adversarial part does not exist, and use the marking tokens as training data. When training is performed with the marking tokens of any row region or column region, the marking tokens of other row regions or column regions are masked; A fourth obtaining module is used to train the visual converter based on the training data to obtain a smoothing classifier, wherein the smoothing classifier is used to accurately classify the historical image and resist the recognition interference caused by the adversarial patch, wherein the visual converter includes a coordinate attention module and a self-attention module, and the fourth obtaining module is used to, Inputting the training data into a classifier to obtain a first classification result for each historical image; Based on the training data and the first classification result, the visual converter is trained, wherein the coordinate attention module is used to enhance features of the training data, and the self-attention module is used to extract the features to obtain a smoothing classifier, wherein the smoothing classifier is used to accurately classify the historical image and resist recognition interference caused by the adversarial patch, and the output of the coordinate attention module is connected to the input of the self-attention module; A second acquisition module is used to acquire a target image to be identified; The fifth obtaining module is used to obtain the recognition and classification result of the target image based on the target image and the smoothing classifier.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method steps according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method steps according to any one of claims 1 to 4 are implemented.