Image detection methods, apparatus, systems and electronic devices

By acquiring and fusing features from white light and NBI images, and using a trained model for image quality detection, the problem of low accuracy in electronic laryngoscope image detection is solved, and an efficient and accurate image detection method is achieved.

CN116012325BActive Publication Date: 2026-05-26ZHEJIANG HEALNOC TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG HEALNOC TECH CO LTD
Filing Date
2022-12-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The accuracy of image detection using electronic laryngoscopes in existing technologies is low, and they mainly rely on manual image capture by the human eye, resulting in low efficiency and poor reliability.

Method used

By acquiring white light images and NBI images, color depth features and texture depth features are extracted respectively, and depth feature fusion is performed. Image quality detection is then performed using a well-trained target image detection model to generate target image detection results.

Benefits of technology

It achieves accurate and efficient image detection, and can select high-quality images from massive images, thus improving the accuracy and efficiency of image detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012325B_ABST
    Figure CN116012325B_ABST
Patent Text Reader

Abstract

This application relates to an image detection method, apparatus, system, and electronic device. The image detection method includes: acquiring a white light image and a NBI image of an image to be detected; performing feature extraction processing on the white light image and the NBI image respectively to obtain color depth features of the white light image and texture depth features of the NBI image; performing depth feature fusion processing based on the color depth features and the texture depth features to obtain fused depth features; performing image quality detection processing based on the fused depth features; and generating a target image detection result for the image to be detected. This application solves the problem of low accuracy in image detection and achieves an accurate and efficient image detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image detection technology, and in particular to image detection methods, apparatus, systems and electronic devices. Background Technology

[0002] Electronic laryngoscopy is a primary auxiliary tool used by otolaryngologists to diagnose ear, nose, and throat diseases. Analysis of electronic laryngoscopy images is a direct reference for doctors in assessing these conditions. During electronic laryngoscopy, a large number of video images are generated. Clear and high-quality images of key organs are crucial for technicians in writing examination reports and for doctors in making diagnoses. Therefore, selecting high-quality images from this massive dataset is of great significance. Currently, related technologies primarily rely on the human eye to manually capture images of different areas during electronic laryngoscopy. This method suffers from problems such as missed images, low efficiency, and poor reliability, resulting in low accuracy in image detection.

[0003] Currently, no effective solution has been proposed to address the low accuracy of image detection in related technologies. Summary of the Invention

[0004] This application provides an image detection method, apparatus, system, and electronic device to at least address the problem of low accuracy in image detection in related technologies.

[0005] In a first aspect, embodiments of this application provide an image detection method, the method comprising:

[0006] Acquire white light images and narrow band imaging (NBI) images of the image to be detected;

[0007] The white light image and the NBI image are respectively subjected to feature extraction processing to obtain the color depth features of the white light image and the texture depth features of the NBI image. Then, the depth features are fused based on the color depth features and the texture depth features to obtain the fused depth features.

[0008] Image quality detection processing is performed based on the fused depth features, and target image detection results are generated for the image to be detected.

[0009] In some embodiments, the method further includes:

[0010] Both the white light image and the NBI image are input into the fully trained target image detection model;

[0011] The white light image is processed by the first feature extraction network in the target image detection model to obtain the color depth feature, and the NBI image is processed by the second feature extraction network in the target image detection model to obtain the texture depth feature.

[0012] In some embodiments, before inputting both the white light image and the NBI image into the fully trained target image detection model, the method further includes:

[0013] Acquire training data; wherein the training data carries image classification and annotation information;

[0014] The training data is input into a preset initial image detection model for training to obtain the training detection results;

[0015] The loss function result is calculated based on the training detection results and the image classification annotation information, and the gradient of the loss function result is backpropagated to the initial image detection model for iterative training to generate the target image detection model.

[0016] In some embodiments, the image quality detection processing based on the fused depth features and the generation of target image detection results includes:

[0017] Temporal features are obtained by performing temporal feature extraction based on the fusion depth features, and image quality detection is performed based on the temporal features to obtain the target image detection result.

[0018] In some embodiments, the feature fusion process based on the color depth feature and the texture depth feature to obtain the fused depth feature includes:

[0019] Obtain the initial optimal image features;

[0020] The fused depth features are obtained by performing feature fusion processing based on the color depth features, the texture depth features, and the initial optimal image features.

[0021] In some embodiments, the method further includes:

[0022] When the target image detection result indicates that the image to be detected is the target optimal image, the temporal features are input into the optimal feature extraction network in the target image detection model, and the optimal feature extraction network is used to calculate the current optimal image features based on the initial optimal image features and the temporal features.

[0023] Feature fusion processing is performed based on the color depth feature, the texture depth feature, and the current optimal image feature to obtain a new fused depth feature; a new temporal feature is obtained based on the new fused depth feature, and a new target image detection result is generated based on the new temporal feature.

[0024] In some embodiments, the step of performing image quality detection processing based on the temporal features and generating target image detection results includes:

[0025] The temporal features are input into the classification network of the target image detection model to perform image classification, and the image classification result is obtained. Based on the image classification result, the optimal target image in the image to be detected is determined.

[0026] The target image detection result is generated based on the optimal target image.

[0027] Secondly, embodiments of this application provide an image detection device, the device comprising: an acquisition module, a feature module, and a generation module;

[0028] The acquisition module is used to acquire the white light image and NBI image of the image to be detected;

[0029] The feature module is used to perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth feature of the white light image and the texture depth feature of the NBI image, and to perform depth feature fusion processing based on the color depth feature and the texture depth feature to obtain the fused depth feature;

[0030] The generation module is used to perform image quality detection processing based on the fused depth features and generate target image detection results for the image to be detected.

[0031] Thirdly, embodiments of this application provide an image detection system, the system comprising: an image acquisition device and a main control device; wherein the main control device is connected to the image acquisition device;

[0032] The image acquisition device is used to acquire the image to be detected and send the image to be detected to the main control device;

[0033] The main control device is used to execute the image detection method as described in the first aspect above.

[0034] Fourthly, embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image detection method as described in the first aspect above.

[0035] Compared to related technologies, the image detection method, apparatus, system, and electronic device provided in this application acquire a white light image and an NBI image of the image to be detected; perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth features of the white light image and the texture depth features of the NBI image; perform depth feature fusion processing based on the color depth features and the texture depth features to obtain fused depth features; perform image quality detection processing based on the fused depth features, and generate a target image detection result for the image to be detected. This solves the problem of low accuracy in image detection and realizes an accurate and efficient image detection method.

[0036] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0037] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0038] Figure 1 This is an application environment diagram of an image detection method according to an embodiment of this application;

[0039] Figure 2 This is a flowchart of an image detection method according to an embodiment of this application;

[0040] Figure 3 This is a flowchart of another image detection method according to an embodiment of this application;

[0041] Figure 4 This is a flowchart of an image detection method according to a preferred embodiment of this application;

[0042] Figure 5 This is a structural block diagram of an image detection device according to an embodiment of this application;

[0043] Figure 6 This is a structural block diagram of an image detection system according to an embodiment of this application;

[0044] Figure 7 This is a structural diagram of the internal structure of a computer device according to an embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0046] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0047] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0048] The image detection method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal device 102 communicates with server device 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server device 104 or placed on a cloud or other network server. Server device 104 acquires the white light image and NBI image of the image to be detected, performs feature extraction processing on the white light image and NBI image respectively to obtain the color depth features of the white light image and the texture depth features of the NBI image, and performs depth feature fusion processing based on the color depth features and texture depth features to obtain fused depth features. Server device 104 performs image quality detection processing based on the fused depth features to generate a target image detection result for the image to be detected, and the terminal device 102 displays the generated target image detection result. Terminal device 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices, such as smartwatches, smart bracelets, and head-mounted devices. Server device 104 can be implemented using a standalone server or a server cluster composed of multiple servers.

[0049] This embodiment provides an image detection method. Figure 2 This is a flowchart of an image detection method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0050] Step S210: Obtain the white light image and NBI image of the image to be detected.

[0051] The images to be detected are consecutive frames extracted from multiple nasopharyngoscopy videos; the images to be detected include white light images and NBI images, that is, each nasopharyngoscopy video has corresponding white light and NBI images.

[0052] Step S220: Perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth features of the white light image and the texture depth features of the NBI image. Then, perform depth feature fusion processing based on the color depth features and the texture depth features to obtain the fused depth features.

[0053] The feature extraction process described above can be as follows: Features are extracted from the white light image and the NBI image using methods such as machine learning, pattern recognition, or image processing. This extracts salient features from the white light image to obtain color depth features, and extracts salient features from the NBI image to obtain texture depth features. The feature fusion process can be as follows: The color depth features and texture depth features are concatenated along the channel dimension using machine learning or image processing to obtain a fused depth feature. It should be noted that before performing feature extraction and fusion on the white light image and the NBI image, preprocessing such as image alignment can be performed to ensure that the locations of lesions in the white light image are the same as those in the NBI image, and that the image size and other information of the white light image are consistent with those of the NBI image, thereby improving the accuracy of subsequent image detection.

[0054] Step S230: Perform image quality detection processing based on the fused depth features, and generate target image detection results for the image to be detected.

[0055] Specifically, image quality detection processing can be performed on the aforementioned fused deep features using methods such as machine learning to obtain the target image detection results. For example, neural network models such as classification networks, discriminators, or generative adversarial networks can be used to detect the fused deep features. Based on the classification or detection results, it can be determined whether the image in the current frame belongs to the optimal image, and the target image detection results for the aforementioned image to be detected can be obtained.

[0056] Through steps S210 to S230, the color depth features of the white light image and the texture depth features of the NBI image are extracted respectively, and the features are fused to obtain the fused depth features. Finally, image quality detection is performed based on the fused depth features. This allows for the simultaneous combination of features from both the white light image and the NBI image. By utilizing the rich color of the white light image and the rich texture of the NBI image, accurate target image detection results can be obtained. This method can optimize the selection of images from nasopharyngoscopes, choosing high-quality images from massive image data. It solves the problem of low accuracy in image detection and realizes an accurate and efficient image detection method.

[0057] In some embodiments, the image detection method further includes the following steps: inputting both the white light image and the NBI image into a fully trained target image detection model; using a first feature extraction network in the target image detection model to perform feature extraction processing on the white light image to obtain the color depth feature; and using a second feature extraction network in the target image detection model to perform feature extraction processing on the NBI image to obtain the texture depth feature.

[0058] The aforementioned target image detection model includes two feature extraction networks. When the white light image and NBI image at the same time are acquired, they are input into the two feature extraction networks to extract the depth features of the images, as shown in Formulas 1 and 2 below.

[0059]

[0060] f b =φ b (x b ) Formula 2

[0061] In the above formula, x a x b Images are shown in white light and NBI format, φ a (x a ), φ b (x b ) are feature extraction networks for processing white light images and NBI images, respectively. a f b These represent the depth features corresponding to the white light image and the NBI image, respectively. Further, φ a (x a ) and φ b (x b Feature extraction networks can have the same network structure, for example, they can use the VGG network structure, and the output feature dimension is...

[0062] Through the above embodiments, by inputting white light images and NBI images into two feature extraction networks in the target image detection model, the efficiency and accuracy of feature extraction for white light images and NBI images are improved, thereby improving the efficiency and accuracy of the image detection method.

[0063] This embodiment also provides an image detection method. Figure 3 This is a flowchart of another image detection method according to an embodiment of this application, such as... Figure 3 As shown, the process includes Figure 2 Steps S210 to S220 shown herein also include the following steps:

[0064] Step S310: Temporal feature extraction is performed based on the fusion depth feature to obtain temporal features, and image quality detection is performed based on the temporal features to obtain the target image detection result.

[0065] After calculating the aforementioned fused deep features, further time-series feature extraction can be performed using deep learning network models or autoregressive algorithms to obtain time-series features that reflect changes over a period of time. Specifically, these time-series features characterize the sequence relationship between tn and t in the current feature map; t represents the current time, and n is a positive integer. The fused deep features are input into a fully trained target time-series feature extraction network for training, and the network outputs the aforementioned time-series features. This target time-series feature network can be a recurrent neural network (RNN), a long short-term memory (LSTM) network, or a gated recurrent unit (GRU) network, or any other neural network capable of time-series prediction. Taking an LSTM network as an example, the time-series feature extraction process is shown in Formula 3 below:

[0066] f β =LSTM(f α ) Formula 3

[0067] In the above formula, f α To fuse deep features, f β For time-series features, LSTM(f) α The above-mentioned target temporal feature extraction network is an LSTM network. Furthermore, the target temporal feature network can be connected to the aforementioned feature extraction networks and deployed in the same target image detection model to improve model training efficiency. After obtaining the temporal features through the above steps, these features can be detected using neural network models such as classification networks, discriminators, or generative adversarial networks. Based on the classification or detection results, it can be determined whether the image in the current frame belongs to the optimal image, and the target image detection result for the image to be detected can be obtained.

[0068] Through the above step S310, temporal features are obtained by extracting temporal features from the fused depth features. Finally, the target image detection result is obtained based on the temporal feature detection. Thus, by combining the spatiotemporal feature extraction of white light images and NBI images, the spatial and temporal characteristics of the data are comprehensively considered, which is conducive to further improving the accuracy of image detection.

[0069] In some embodiments, the feature fusion process based on the color depth feature and the texture depth feature to obtain the fused depth feature further includes the following steps: obtaining initial optimal image features; and performing feature fusion processing based on the color depth feature, the texture depth feature, and the initial optimal image features to obtain the fused depth feature. The initial optimal image features can be features extracted from a randomly generated image. Specifically, the feature fusion process is achieved by concatenating the color depth feature, the texture depth feature, and the initial optimal image features along the channel dimension, thereby obtaining the fused depth feature.

[0070] Furthermore, the above image detection method also includes the following steps:

[0071] Step S241: If the target image detection result indicates that the image to be detected is the target optimal image, the temporal feature is input into the optimal feature extraction network in the target image detection model, and the current optimal image feature is calculated based on the initial optimal image feature and the temporal feature using the optimal feature extraction network.

[0072] After detecting that the current frame image is the highest quality image through the above steps, the temporal features corresponding to the frame image are input into the optimal feature extraction network, as shown in Formula 4 below:

[0073] f c =EXT(f′) β ,f β ) Formula 4

[0074] In the above formula, f β Let f' be the previous optimal image feature, and let f be the initial optimal image feature. β ′=f β EXT(f′) β ,f β This represents the optimal feature extraction network, which consists of 1×1 convolutional networks. 1x1 The optimal image features are obtained by performing a 3x3 convolution.

[0075] Step S242: Perform feature fusion processing based on the color depth feature, the texture depth feature, and the current optimal image feature to obtain a new fused depth feature; obtain a new temporal feature based on the new fused depth feature, and generate a new target image detection result based on the new temporal feature.

[0076] The feature fusion process described above is shown in Formula 5:

[0077]

[0078] In the above formula, f c The features that constitute the current optimal image. The structure mainly consists of one feature concatenated cat(f) a ,f b ,f c ) and 2 convolutions conv 1x1 conv 3x3 Composition, to obtain f a f b f c Features are obtained through cat(f) a ,f b ,f c By splicing along the channel dimension, we obtain... Features are then processed by 1x1 convolutions (conv). 1x1 A new fused depth feature is obtained by performing a 3x3 convolution (conv3x3). Furthermore, the input of this feature fusion network can be connected to the outputs of the aforementioned feature extraction networks, and the output of this feature fusion network can be connected to the input of the aforementioned target temporal feature extraction network. All networks are deployed within the same target image detection model to improve model training efficiency. It should be noted that in this embodiment, temporal feature extraction and classification detection are performed on the newly obtained fused deep features until the loss function of the currently detected classifier is minimized or the number of iterations or duration is reached, ultimately yielding the aforementioned new target image detection result.

[0079] Through steps S241 to S242, new fused depth features are obtained by fusing the current best image features with the color depth features and texture depth features obtained through feature extraction. This dynamically updates the historical best image features and filters out duplicate best images, further improving the accuracy of the image detection method.

[0080] In some embodiments, the above-described image quality detection processing based on the temporal features and the generation of target image detection results further includes the following steps:

[0081] Step S231: The temporal features are input into the classification network of the target image detection model for image classification to obtain the image classification result. Based on the image classification result, the optimal target image in the image to be detected is determined. Specifically, the aforementioned temporal features are input into the fully trained classification network for image classification to determine whether the current frame image is the optimal image. This image classification process is shown in Formula 6 below:

[0082] y pred =γ(f α ) Formula 6

[0083] In the above formula, γ(f α ) is used to represent a classification network, which can consist of two fully connected layers; y pred The classification result is shown below. Furthermore, the input to this classification network can be connected to the output of the aforementioned target temporal feature extraction network to improve model training efficiency.

[0084] Step S232: Generate the target image detection result based on the optimal image of the target.

[0085] Specifically, if the current frame image is detected as the optimal image by the output of the above classification network, a first target image detection result indicating that the quality of the frame image is optimal is generated; if the current frame image is not detected as the optimal image by the output of the above classification network, a second target image detection result indicating that the quality of the frame image is substandard is generated.

[0086] Through steps S231 to S232, the trained classification network performs image classification and detection on the temporal features, thereby enabling the deep learning network to quickly and accurately determine whether the current frame image is the optimal image, thus effectively improving the accuracy and efficiency of image detection.

[0087] In some embodiments, before inputting both the white light image and the NBI image into the fully trained target image detection model, the image detection method further includes the following steps:

[0088] Step S201: Obtain training data; wherein, the training data carries image classification and annotation information.

[0089] The training data mentioned above can be obtained by processing pre-prepared historical images; these historical images can be multiple. Specifically, staff can pre-examine each historical image and determine whether it is a high-quality image based on constraints such as image clarity and whether the lesion is centered in the image. Each historical image is then labeled; for example, if the historical image is a high-quality image, an image classification label of 1 is added, otherwise a label of 0 is added. Then, the image classification label information is combined with the above-mentioned image classification label information to annotate each historical image, resulting in the training dataset.

[0090] Step S202: Input the training data into the preset initial image detection model for training to obtain the training detection result.

[0091] The initial image detection model can be a neural network model such as a Deep Residual Network (ResNet), VGG-19, DenseNet, or High-Resolution Network (HRNet). Specifically, this initial image detection model can consist of a two-way feature extraction network, a feature fusion network, an LSTM network, a classification network, and an optimal image feature extraction network. The training data is then input into this initial image detection model and trained sequentially through each of the aforementioned networks to obtain the training detection results.

[0092] Step S203: Calculate the loss function result based on the training detection result and the image classification annotation information, and backpropagate the gradient of the loss function result to the initial image detection model for iterative training to generate the target image detection model.

[0093] The above loss function result can be calculated using the Binary Cross Entropy Loss (BCE); the loss function is shown in Equation 7 below:

[0094] loss = BCE(y pred ,y ture ) Formula 7

[0095] In the above formula, y ture y is used to represent the image classification annotation information carried in the above training data. pred This is used to represent the training detection results. The gradient of the loss function result calculated by the above formula is then backpropagated to the above target image detection model for iterative training until the number of iterations or the iteration time is satisfied, or the model tends to converge, and finally the above fully trained target image detection model is obtained.

[0096] Through the above embodiments, the loss function result is calculated by training detection results and image classification annotation information, and the initial image detection model is optimized and trained according to the loss function result to obtain an optimized target image detection model, which can improve the accuracy of model processing and thus effectively improve the accuracy of image detection methods.

[0097] The embodiments of this application will be described and illustrated below through preferred embodiments. Figure 4 This is a flowchart of an image detection method according to a preferred embodiment of this application, such as... Figure 4 As shown, the process includes the following steps:

[0098] Step S401: Obtain the white light image and NBI image to be detected.

[0099] In step S402, the white light image and the NBI image are respectively input into the two feature extraction networks in the target image detection model to extract the color depth features of the white light image and the texture depth features of the NBI image.

[0100] Step S403: The color depth feature, texture depth feature and initial optimal image feature are fused using the above formula 5 to obtain the fused depth feature.

[0101] Step S404: Input the current fused depth features into the temporal extraction network in the target image detection model to obtain temporal features.

[0102] Step S405: Input the above temporal features into the classifier in the target image detection model, and determine whether the current frame is the optimal image based on the classification results.

[0103] Step S406: If the image is determined to be the optimal image through step S405 above, the corresponding temporal feature is input into the optimal feature extraction network in the target image detection model to obtain the current optimal image, and the value is returned to step S403 above to update and obtain new fused depth features. Then, the image is further processed by LSTM and classification network to generate new target image detection results, and finally the image with the best quality is selected from the massive number of images.

[0104] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0105] This embodiment also provides an image detection device for implementing the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0106] Figure 5 This is a structural block diagram of an image detection device according to an embodiment of this application, such as... Figure 5As shown, the device includes: an acquisition module 52, a feature module 54, and a generation module 56; the acquisition module 52 is used to acquire a white light image and an NBI image of the image to be detected; the feature module 54 is used to perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth feature of the white light image and the texture depth feature of the NBI image, and perform depth feature fusion processing based on the color depth feature and the texture depth feature to obtain fused depth features; the generation module 56 is used to perform image quality detection processing based on the fused depth features and generate a target image detection result for the image to be detected.

[0107] Through the above embodiments, the feature module 54 extracts the color depth features of the white light image and the texture depth features of the NBI image respectively, and performs feature fusion to obtain the fused depth features. Finally, the generation module 56 performs image quality detection based on the fused depth features. This allows for the simultaneous combination of white light image and NBI image features, utilizing the rich color of white light and the rich texture of NBI to obtain accurate target image detection results. This solves the problem of low image detection accuracy and realizes an accurate and efficient image detection device.

[0108] In some embodiments, the feature module 54 is further configured to input both the white light image and the NBI image into a fully trained target image detection model; the feature module 54 uses the first feature extraction network in the target image detection model to perform feature extraction processing on the white light image to obtain the color depth feature, and uses the second feature extraction network in the target image detection model to perform feature extraction processing on the NBI image to obtain the texture depth feature.

[0109] In some embodiments, the image detection device further includes a training module; the training module is used to acquire training data; wherein the training data carries image classification annotation information; the training module inputs the training data into a preset initial image detection model for training to obtain training detection results; the training module calculates a loss function result based on the training detection result and the image classification annotation information, and backpropagates the gradient of the loss function result to the initial image detection model for iterative training to generate the target image detection model.

[0110] In some embodiments, the generation module 56 is further configured to perform temporal feature extraction processing based on the fusion depth features to obtain temporal features, and perform image quality detection processing based on the temporal features to obtain the target image detection result.

[0111] In some embodiments, the feature module 54 is further configured to obtain initial optimal image features; the feature module 54 performs feature fusion processing based on the color depth feature, the texture depth feature and the initial optimal image features to obtain the fused depth feature.

[0112] In some embodiments, the image detection device further includes a loop module; the loop module is configured to, when the target image detection result indicates that the image to be detected is the target optimal image, input the temporal features into the optimal feature extraction network in the target image detection model, and use the optimal feature extraction network to calculate the current optimal image features based on the initial optimal image features and the temporal features; the loop module performs feature fusion processing based on the color depth features, the texture depth features and the current optimal image features to obtain new fused depth features; the loop module obtains new temporal features based on the new fused depth features, and generates a new target image detection result based on the new temporal features.

[0113] In some embodiments, the generation module 56 is further configured to input the temporal features into the classification network in the target image detection model to perform image classification, obtain image classification results, and determine the optimal target image in the image to be detected based on the image classification results; the generation module 56 generates the target image detection result based on the optimal target image.

[0114] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0115] This embodiment also provides an image detection system. Figure 6 This is a structural block diagram of an image detection system according to an embodiment of this application, such as... Figure 6As shown, the system includes an image acquisition device 62 and a main control device 64; wherein the main control device 64 is connected to the image acquisition device 62; the image acquisition device 62 is used to acquire an image to be detected and send the image to be detected to the main control device 64; the main control device 64 is used to execute the steps in any of the above method embodiments. The main control device 64 includes, but is not limited to, various microcontrollers, main control chips, computers, server devices, or other hardware devices used to control the above image detection process. Further, the main control device 64 can communicate with the image acquisition device 62 through a transmission device. The transmission device is used to receive or send data via a network. The network includes a wireless network provided by the platform's communication provider. In one embodiment, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station to communicate with the Internet. In one embodiment, the transmission device can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0116] Through the above embodiments, the main control device 64 extracts the color depth features of the white light image and the texture depth features of the NBI image, respectively, and performs feature fusion to obtain fused depth features. Finally, image quality detection is performed based on the fused depth features. This allows for the simultaneous combination of white light image and NBI image features, utilizing the rich color of white light and the rich texture of NBI to obtain accurate target image detection results. This solves the problem of low image detection accuracy and realizes an accurate and efficient image detection system.

[0117] In some embodiments, a computer device is provided, which may be a server. Figure 7 This is a structural diagram of the internal structure of a computer device according to an embodiment of this application, such as... Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores target image detection results. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the aforementioned image detection method.

[0118] Those skilled in the art will understand that Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0119] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0120] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0121] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0122] S1, acquire the white light image and NBI image of the image to be detected.

[0123] S2, perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth features of the white light image and the texture depth features of the NBI image, and perform depth feature fusion processing based on the color depth features and the texture depth features to obtain the fused depth features.

[0124] S3, based on the fused depth features, perform image quality detection processing and generate target image detection results for the image to be detected.

[0125] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0126] Furthermore, in conjunction with the image detection methods described in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the image detection methods described in the above embodiments.

[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0128] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0129] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An image detection method characterized by, The method includes: Acquire the white light image and NBI image of the image to be detected; The white light image and the NBI image are respectively subjected to feature extraction processing to obtain the color depth features of the white light image and the texture depth features of the NBI image. Then, a depth feature fusion processing is performed based on the color depth features and the texture depth features to obtain the fused depth features. This includes: obtaining initial optimal image features; and performing feature fusion processing based on the color depth features, the texture depth features, and the initial optimal image features to obtain the fused depth features. Image quality detection processing is performed based on the fused deep features, and a target image detection result is generated for the image to be detected. This includes: extracting temporal features based on the fused deep features to obtain temporal features, and performing image quality detection processing based on the temporal features to obtain the target image detection result; wherein, the temporal features represent the sequence relationship between tn and t in the current feature map; t is used to represent the current time, and n is a positive integer; When the target image detection result indicates that the image to be detected is the optimal target image, the temporal features are input into the optimal feature extraction network in the target image detection model, and the optimal feature extraction network is used to calculate the current optimal image features based on the initial optimal image features and the temporal features; feature fusion processing is performed based on the color depth features, the texture depth features and the current optimal image features to obtain new fused depth features; new temporal features are obtained based on the new fused depth features, and a new target image detection result is generated based on the new temporal features.

2. The image detection method according to claim 1, characterized in that, The method further includes: Both the white light image and the NBI image are input into the fully trained target image detection model; The white light image is processed by the first feature extraction network in the target image detection model to obtain the color depth feature, and the NBI image is processed by the second feature extraction network in the target image detection model to obtain the texture depth feature.

3. The image detection method according to claim 2, characterized in that, Before inputting both the white light image and the NBI image into the fully trained target image detection model, the method further includes: Acquire training data; wherein the training data carries image classification and annotation information; The training data is input into a preset initial image detection model for training to obtain the training detection results; The loss function result is calculated based on the training detection results and the image classification annotation information, and the gradient of the loss function result is backpropagated to the initial image detection model for iterative training to generate the target image detection model.

4. The image detection method of claim 1, wherein, The step of performing image quality detection processing based on the temporal features and generating target image detection results includes: The temporal features are input into the classification network of the target image detection model to perform image classification, and the image classification result is obtained. Based on the image classification result, the optimal target image in the image to be detected is determined. The target image detection result is generated based on the optimal target image.

5. An image detection apparatus characterized by comprising: The device includes: an acquisition module, a feature module, a generation module, and a loop module; The acquisition module is used to acquire the white light image and NBI image of the image to be detected; The feature module is used to perform feature extraction processing on the white light image and the NBI image respectively to obtain the color depth features of the white light image and the texture depth features of the NBI image, and to perform depth feature fusion processing based on the color depth features and the texture depth features to obtain fused depth features, including: obtaining initial optimal image features; and performing feature fusion processing based on the color depth features, the texture depth features and the initial optimal image features to obtain the fused depth features. The generation module is used to perform image quality detection processing based on the fusion depth features and generate a target image detection result for the image to be detected, including: performing temporal feature extraction processing based on the fusion depth features to obtain temporal features, and performing image quality detection processing based on the temporal features to obtain the target image detection result; The loop module is used to input the temporal features into the optimal feature extraction network in the target image detection model when the target image detection result indicates that the image to be detected is the target optimal image, and to use the optimal feature extraction network to calculate the current optimal image features based on the initial optimal image features and the temporal features. The loop module is further configured to perform feature fusion processing based on the color depth feature, the texture depth feature, and the current optimal image feature to obtain a new fused depth feature; obtain a new temporal feature based on the new fused depth feature; and generate a new target image detection result based on the new temporal feature.

6. An image detection system, characterized by The system includes: an image acquisition device and a main control device; wherein the main control device is connected to the image acquisition device; The image acquisition device is used to acquire the image to be detected and send the image to be detected to the main control device; The main control device is used to execute the image detection method as described in any one of claims 1 to 4. 7.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to run the computer program to perform the image detection method according to any one of claims 1 to 4.