Image detection method and apparatus, electronic device, and storage medium

By fusing the features of the multi-frame images acquired by the lens module into the features of the current frame image, the problem of inaccurate detection of the lens module in the prior art is solved, and higher detection accuracy is achieved.

WO2025124114A1PCT designated stage expired Publication Date: 2025-06-19SHINING 3D TECH CO LTD

Patent Information

Application Number
PCT/CN2024/134024
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-11-23
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The prior art is prone to various interferences when detecting the dirty lens module, resulting in inaccurate detection results and ineffective identification of the dirty lens module.

Method used

By acquiring the current frame image and multi-frame reference images collected by the lens module, feature extraction is performed separately to obtain a shallow local feature map, and the feature map of the reference image is fused into the feature map of the current frame image to enhance the feature value of the target pixel position to improve the accuracy of dirty detection.

Benefits of technology

By integrating the feature map, the significance of dirty features is enhanced, the accuracy of dirty detection of lens modules is significantly improved, and the misidentification rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024134024_19062025_PF_FP_ABST
    Figure CN2024134024_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image detection method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a current image frame collected by a lens module, and a plurality of frames of reference images collected before and / or after collecting the current image frame; performing extraction to obtain a shallow local feature map of the current image frame and a shallow local feature map of each reference image frame; and then fusing the shallow local feature map of each reference image frame into the shallow local feature map of the current image frame, wherein in a selected specific fusion mode, a feature value, having high similarity with a feature value on a corresponding position in the shallow local feature map of the reference image, on a position of the shallow local feature map of the current image frame is enhanced, so as to obtain a fused shallow local feature map. Because a dirty feature becomes more remarkable in the finally obtained fused shallow local feature map, a dirt detection result obtained on the basis of the fused shallow local feature map is also more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Image detection method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 14, 2023, with application number 202311719012.5, and invention name “Image detection method, device, electronic device and storage medium”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of image processing technology, and in particular to an image detection method, device, electronic device, and storage medium. Background Art

[0003] When using a lens module to capture images, it's often contaminated. This can severely impact the quality of the captured images, and consequently, the subsequent use of the images. Therefore, it's necessary to promptly detect lens contamination to prompt the user to address it. However, existing technologies are susceptible to interference from various factors, misidentifying portions of an image that resemble dirt as contamination, resulting in inaccurate detection results. Therefore, a more accurate solution for detecting lens contamination is needed. Summary of the Invention

[0004] The present disclosure provides an image detection method, device, electronic device, and storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image detection method, comprising: acquiring a current frame image and a reference image captured by a lens module, wherein the reference image comprises a plurality of frame images captured before and / or after the current frame image; performing feature extraction on the current frame image and the reference image respectively to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image; fusing the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein, during the fusion process, a feature value at a target pixel position in the shallow local feature map of the current frame image is enhanced, and a similarity between the feature value at the target pixel position and the feature value at a corresponding pixel position at the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity; and determining whether there is dirt in the lens module based on the fused shallow local feature map.

[0006] According to a second aspect of an embodiment of the present disclosure, an image detection method is provided, the method comprising: obtaining an image to be detected currently captured by a lens module on a three-dimensional scanning device; inputting the image to be detected into a pre-trained detection model, and determining the current scanning environment type of the three-dimensional scanning device through the detection model; wherein the detection model is trained based on the following method: obtaining at least two frames of sample images, and the scanning environment types corresponding to the at least two frames of sample images are the same; using a preset initial model to perform feature extraction on the at least two frames of sample images respectively to obtain global features of the at least two frames of sample images; determining a target loss based on the difference between the similarity of the global features of each pair of sample images in the at least two frames of sample images and a preset similarity threshold, and adjusting the model parameters of the initial model based on the target loss to train a detection model.

[0007] According to a third aspect of an embodiment of the present disclosure, an image detection method is provided, the method comprising: obtaining an image to be detected currently captured by a lens module on a three-dimensional scanning device; inputting the image to be detected into a pre-trained detection model, and determining the current scanning environment type of the three-dimensional scanning device through the detection model; wherein the detection model is trained based on the following method: obtaining a sample image triple, each group of sample image triples including a first sample image, a second sample image having the same scanning environment type as the first sample image, and a third sample image having a different scanning environment type from the first sample image; performing feature extraction on the first sample image, the second sample image, and the third sample image through a preset initial model to obtain respective global features; determining a target loss based on the similarity between the global features of the second sample image and the global features of the first sample image, and the similarity between the global features of the third sample image and the global features of the first sample image, and adjusting the model parameters of the initial model based on the target loss to train a detection model.

[0008] According to a fourth aspect of an embodiment of the present disclosure, an image detection method is provided, the method comprising: obtaining an image to be detected currently captured by a lens module on a three-dimensional scanning device; performing feature extraction on the image to be detected to obtain a global feature; and determining whether the current scanning environment type of the three-dimensional scanning device is a target environment based on a degree of proximity between the global feature and a preset feature clustering center; wherein the feature clustering center is a clustering center of the global features of multiple frames of sample images, and the sample image is an image captured by the three-dimensional scanning device in the target environment.

[0009] According to a fifth aspect of an embodiment of the present disclosure, an image detection device is provided, which includes: an acquisition module, configured to acquire a current frame image and a reference image captured by a lens module, wherein the reference image includes multiple frame images captured before and / or after the current frame image; a feature extraction module, configured to perform feature extraction on the current frame image and the reference image, respectively, to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image; a fusion module, configured to fuse the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein during the fusion process, the feature value at the target pixel position in the shallow local feature map of the current frame image is enhanced, and the similarity between the feature value of the target pixel position and the feature value of the corresponding pixel position of the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity; a prediction module, configured to determine whether there is dirt in the lens module based on the fused shallow local feature map.

[0010] According to the sixth aspect of an embodiment of the present disclosure, an electronic device is provided, which includes a processor, a memory, and computer instructions stored in the memory for execution by the processor. When the processor executes the computer instructions, the methods mentioned in the first, second, third and / or fourth aspects above can be implemented.

[0011] According to the seventh aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed, the methods mentioned in the first, second, third and / or fourth aspects above are implemented.

[0012] In the disclosed embodiments, when performing contamination detection on a lens module based on images captured by the lens module, the fixed location of contamination in multiple frames of images captured by the lens module can be incorporated into the detection mechanism to improve the accuracy of the detection results. A current frame image captured by the lens module and multiple reference frames captured by the lens module before and / or after the current frame image can be obtained, and feature extraction is performed on the current frame image and the multiple reference frames to obtain a shallow local feature map of the current frame image and shallow local feature maps of each reference frame image. Among them, the shallow local features are mainly some texture, edge, corner and other information in the image, that is, the shallow local features cover the dirt features in the image. In order to highlight the dirt features in the shallow local feature map, the shallow local feature maps of each frame reference image can be fused into the shallow local feature map of the current frame image, and a specific fusion method can be selected so that in the fusion process, the feature values ​​at the positions with high similarity in the shallow local feature map of the current frame image and the shallow local feature map of the reference image are enhanced, thereby obtaining a fused shallow local feature map, and determining whether there is dirt in the lens module based on the fused shallow local feature map. Since the position of the dirt in each frame image is fixed, the positions with high similarity in the shallow local feature maps of each frame image are most likely the positions where the dirt is located. By enhancing the feature values ​​of these positions, the features of the dirt in the final fused shallow local feature map become more prominent, and the dirt detection result obtained based on the fused shallow local feature map is also more accurate.

[0013] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0015] FIG1 is a schematic diagram of an image detection method according to an embodiment of the present disclosure.

[0016] FIG2 is a flow chart of an image detection method according to an embodiment of the present disclosure.

[0017] FIG3 is a schematic diagram of a method for detecting whether a lens module is dirty according to an embodiment of the present disclosure.

[0018] FIG4 is a schematic diagram of a method for detecting whether a lens module is dirty according to an embodiment of the present disclosure.

[0019] FIG5 is a schematic diagram of a method for detecting whether a lens module is dirty or foggy, and the type of a scanning environment according to an embodiment of the present disclosure.

[0020] FIG6 is an architecture diagram of a detection model according to an embodiment of the present disclosure.

[0021] FIG7 is a flow chart of an image detection device according to an embodiment of the present disclosure.

[0022] FIG8 is a schematic diagram of the logical structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0024] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a" and "the" used in this disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein means any combination of at least two of any one or more of a plurality of.

[0025] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0026] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure are further described in detail below with reference to the accompanying drawings.

[0027] When using a lens module to capture images, it is often possible for the lens module to be dirty. This can seriously affect the quality of the captured image and, in turn, the subsequent use of the image. Therefore, it is necessary to promptly detect lens module dirtiness to prompt the user to address the lens dirtiness.

[0028] For example, if a dental 3D scanner is used to scan a patient's teeth and reconstruct a dental model, if there is dirt in the lens module of the 3D scanner, the dirt will be present in every frame of the image captured by the 3D scanner, seriously affecting the quality of the 3D model reconstructed from the image, forcing the user to rescan, seriously affecting the user experience. Therefore, it is necessary to promptly detect dirt in the lens module and prompt the user to clean it immediately.

[0029] The applicant has found that currently, when performing dirt detection on the lens module, a large number of sample images marked with dirt can be used in advance to train a detection model, and then the detection model can be used to detect the image to be detected to obtain the detection result. In this detection method, the model only predicts the detection result based on the features of the current frame image, and is easily interfered by various situations. For example, some parts of the image that look like dirt are mistakenly identified as dirt. However, the applicant took into account that for scenes where the lens module is dirty, the images captured by the lens module have a significant characteristic: the position of the dirt in the multiple frames of images captured by the lens module is fixed. At present, when using the model to detect whether the lens module is dirty, the features of the single frame image captured by the lens module are extracted, and the detection results are predicted based on the features of the single frame image. The above-mentioned characteristics in this scene are not utilized, resulting in low accuracy of the detection results.

[0030] Based on this, an embodiment of the present application provides an image detection method. When performing dirt detection on a lens module based on images captured by the lens module, the characteristic that the position of dirt in multiple frames of images captured by the lens module is fixed can be incorporated into the detection mechanism to improve the accuracy of the detection results. For example, as shown in Figure 1, a current frame image captured by the lens module and multiple frames of reference images captured by the lens module before and / or after the current frame image are captured can be obtained, and feature extraction is performed on the current frame image and the multiple frames of reference images to obtain a shallow local feature map of the current frame image and shallow local feature maps of each frame of reference image. Among them, the shallow local features are mainly some texture, edge, corner and other information in the image, that is, the shallow local features cover the dirt features in the image. In order to highlight the dirt features in the shallow local feature map, the shallow local feature maps of each frame reference image can be fused into the shallow local feature map of the current frame image, and a specific fusion method can be selected so that in the fusion process, the feature values ​​at the positions with high similarity in the shallow local feature map of the current frame image and the shallow local feature map of the reference image are enhanced, thereby obtaining a fused shallow local feature map, and determining whether there is dirt in the lens module based on the fused shallow local feature map. Since the position of the dirt in each frame image is fixed, the positions with high similarity in the shallow local feature maps of each frame image are most likely the positions where the dirt is located. By enhancing the feature values ​​of these positions, the features of the dirt in the final fused shallow local feature map become more prominent, and the dirt detection result obtained based on the fused shallow local feature map is also more accurate.

[0031] The image detection method provided in the embodiments of the present application can be performed by various electronic devices with detection capabilities, such as mobile phones, computers, cloud servers, etc. The electronic device can be a device that captures images, or other devices that are communicatively connected to the device that captures images, and the embodiments of the present application do not impose any restrictions.

[0032] As shown in FIG2 , a flow chart of the image detection method provided in an embodiment of the present application specifically includes the following steps:

[0033] S202, acquiring a current frame image and a reference image captured by a lens module, wherein the reference image includes multiple frames of images captured before and / or after the current frame image;

[0034] In step S202, a current frame image captured by the lens module and a reference image captured by the lens module may be obtained. The reference image may be a plurality of frames captured before and / or after the current frame image. For example, 2N+1 frames of images captured continuously by the lens module may be obtained, the N+1th frame image may be used as the current frame image, and the N frames before and after the current frame image may be used as reference images. Of course, the 2N+1 frames of images may also be images captured discontinuously, and this embodiment of the present application does not impose any limitation thereto.

[0035] S204, performing feature extraction on the current frame image and the reference image respectively to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image;

[0036] In step S204, feature extraction can be performed on the current frame image and each reference frame image to obtain a shallow local feature map for the current frame image and a shallow local feature map for each reference frame image. Shallow local features are low-level features of the image, such as color, texture, and edges. They contain features of more pixels and, therefore, more details. Features such as dirt in the image are shallow local features and can be represented by the shallow local feature map.

[0037] Among them, feature extraction is performed on the current frame image and the reference image to obtain a shallow local feature map, which can be achieved through some feature extraction networks, such as MobileNetV2 network, ResNet network, VGG network, Transform network, etc.

[0038] S206, fusing the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein during the fusion process, the feature value at the target pixel position in the shallow local feature map of the current frame image is enhanced, and the similarity between the feature value at the target pixel position and the feature value at the corresponding pixel position at the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity;

[0039] Considering that lens dirt usually appears at a fixed position in multiple frames of images captured by the lens module, that is, the lens dirt feature will appear at the same position in the shallow local feature map of the multiple frames, and since the dirt is the same, the feature values ​​of the dirt position in the shallow local feature map of the multiple frames should be close. Based on this, in step S206, the shallow local feature map of each reference image can be fused one by one into the shallow local feature map of the current frame image, and a suitable fusion method can be selected so that during the fusion process, the feature value of the target pixel position in the shallow local feature map of the current frame image is enhanced, wherein the similarity between the feature value of the target pixel position and the feature value of the corresponding pixel position of the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity. Among them, the position corresponding to the dirty feature in the shallow local feature map of the current frame image is most likely a position in the shallow local feature map of the current frame image where the feature values ​​are highly similar to those in the shallow local feature map of the reference image. Therefore, during the fusion process, the feature values ​​at the target pixel position where the two frame feature maps have a high similarity can be enhanced, so that in the final fused shallow local feature map, the dirty feature is enhanced, that is, the dirty feature becomes more prominent.

[0040] Among them, in order to enhance the dirty features in the shallow local feature map of the current frame image during the fusion process, so that the dirty features are more prominent, when the shallow local feature map of the reference image is fused with the shallow local feature map of the current frame image, a fusion method based on the attention mechanism can be adopted, or a pyramid pooling module (PPM) can be used to fuse them, or an atrous spatial pyramid pooling module (ASPP) can be used to fuse them, or a non-local module (Non-Local) can be used to fuse them, etc.

[0041] S208: Determine whether there is dirt in the lens module based on the fused shallow local feature map.

[0042] In step S208, after obtaining the fused shallow local feature map, it can be determined whether the lens module is dirty based on the fused shallow local feature map. Since the dirt features in the fused shallow local feature map are enhanced and become more significant, the dirt prediction result obtained based on the fused feature map will be more accurate. Among them, the dirt detection of the lens module can be performed directly based on the fused shallow local feature map, or the dirt detection of the lens module can be performed in combination with the fused shallow local feature map and other types of features of the current frame image. The embodiments of the present application do not limit this. Among them, if the detection result is that the lens module is dirty, the user can be prompted to clean the dirt. The prompting methods include but are not limited to: voice prompts, pop-up prompts, ringtone prompts, vibration prompts, etc.

[0043] Considering that the shallow local feature map only reflects some low-level features of the image, the detection result of the dirt detection of the lens module based only on the shallow local features may not be accurate enough. Therefore, in some embodiments, as shown in Figure 3, when determining whether there is dirt in the lens module based on the fusion of the shallow local feature map, the current frame image can be further feature extracted to obtain the global features of the current frame image, wherein the global features are higher-level features extracted from the image, which contain some high-level semantic information in the image. The global features and the shallow local feature map can then be fused to obtain fused features, and then the presence of dirt in the lens module can be determined based on the fused features. Among them, the global features can be represented in the form of feature maps or in the form of feature vectors, which is not limited in the embodiments of the present application.

[0044] In some embodiments, if the global feature is represented by a global feature vector, when the global feature is fused with the shallow local feature map to obtain the fused feature, the shallow local feature map can be first pooled to obtain a shallow local feature vector, and then the shallow local feature vector is fused with the global feature vector. There are many ways of fusion. For example, the shallow local feature vector can be directly spliced ​​with the global feature vector to obtain the fused feature. Alternatively, the shallow local feature vector can be superimposed with the vector of the global feature to obtain the fused feature, where superposition refers to adding the values ​​of the corresponding positions of the two to obtain a new feature vector.

[0045] In some embodiments, as shown in FIG4 , before the shallow local feature map of the reference image is fused into the shallow local feature map of the current frame image, the shallow local feature map of the reference image and the shallow local feature map of the current frame image can be subjected to maximum pooling processing respectively, and then the shallow local feature map of the reference image after the maximum pooling processing is fused into the shallow local feature map of the current frame image after the maximum pooling processing. Among them, the maximum pooling processing can extract the significant features of the local area in the shallow local feature map, that is, it can extract the features of the dirty area level (relative to the pixel level), making the dirty features more prominent. In addition, the maximum pooling processing can also reduce the size of the shallow local feature map, thereby accelerating the subsequent fusion operation.

[0046] In some embodiments, when fusing the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, the shallow local feature map of each reference image frame can be fused with the shallow local feature map of the current frame image, wherein, in order to enhance the feature value of the target pixel position in the shallow local feature map of the current frame image after fusion, during fusion, for each pixel position in the shallow local feature map of the current frame image, the higher the similarity between the feature value of the pixel position and the feature value of the corresponding pixel position in the shallow local feature map of the reference image, the greater the fusion weight of the feature value of the pixel position. In this way, the feature value of the pixel position with a high similarity to the shallow local feature map of the reference image in the shallow local feature map of the current frame image can be enhanced. After fusing the shallow local feature map of each reference image frame with the shallow local feature map of the current frame image, the shallow local feature maps of each fused frame can be superimposed to obtain a fused shallow local feature map, wherein the superposition process is to add the feature values ​​of the corresponding pixel positions on the shallow local feature maps of each frame.

[0047] In some embodiments, the lens module can be a lens module on a 3D scanning device. The 3D scanning device can be an oral scanner, a facial scanner, an industrial scanner, or a professional scanner, and can be used for 3D reconstruction of objects such as teeth, faces, bodies, industrial products, industrial equipment, cultural relics, artworks, prostheses, medical devices, and buildings.

[0048] Considering that for 3D scanning equipment, during use, there may be a certain temperature difference between the scanning environment and the external environment, which often causes fog on the lens module. For example, taking the oral 3D scanning equipment as an example, due to a certain temperature difference between the oral environment and the external environment, fog will appear in the lens module, affecting the oral scanning. Therefore, the lens module can be tested for fog. When fog is detected on the lens module, the defog function can be turned on first to remove the fog on the lens module before collecting images, so as to avoid the impact of fog on the collected images, which in turn affects the subsequent 3D reconstruction process.

[0049] In addition, there are usually multiple scanning scenarios for three-dimensional scanning equipment. Under different scanning scenarios, the impurities around the target object to be reconstructed, the brightness of the scanning environment, etc. may vary greatly. Therefore, when the three-dimensional scanning device collects images, the more appropriate operating parameters of the three-dimensional scanning device, or when using the images collected by the three-dimensional scanning device for three-dimensional reconstruction, the more appropriate processing methods for processing the images are different.

[0050] For example, taking oral 3D scanning equipment as an example, it has two usage scenarios: intraoral scanning and extraoral scanning. Due to the large differences between the intraoral and extraoral environments, for example, there are interferences such as gums, tongue, and buccal soft tissue in the intraoral environment, while the extraoral environment does not have the above problems. Or, compared with the extraoral environment, the brightness in the intraoral environment is lower, etc. Therefore, for the two scanning environments, there are also differences in the way the images are processed when acquiring images or when reconstructing teeth using the acquired images.

[0051] Considering the above issues, different operating modes can be set for different usage scenarios, and appropriate processing methods, algorithm parameters, and / or device operating modes can be configured for each operating mode. In one embodiment, the operating modes of a 3D scanning device can include the operating modes of each device in the 3D scanning device and the processing methods of the 3D scanning software that matches the 3D scanning device.

[0052] Before using a 3D scanning device to capture images, the current scanning environment type can also be detected first. After selecting an appropriate working mode based on the current scanning environment type, image capture, 3D reconstruction, and other tasks can be performed.

[0053] It should be noted that for oral 3D scanning equipment, in addition to the two usage scenarios of intraoral scanning and extraoral scanning, it can be further subdivided according to specific application scenarios. Then, different working modes can be set for each subdivided usage scenario, and appropriate processing methods, algorithm parameters, and / or device operation modes can be configured for each working mode.

[0054] In some embodiments, as shown in FIG5 , in addition to detecting dirt on the lens module based on an image, the image detection method can also simultaneously detect lens fog and the type of scanning environment, thereby improving detection efficiency. For example, feature extraction can be performed on the current frame image to obtain the global features of the current frame image, and feature extraction can be performed on the reference image to obtain the global features of each frame of the reference image. The global features of the reference image are fused with the global features of the current frame image to obtain a first fused global feature, and then the current scanning environment type of the three-dimensional scanning device is determined based on the first fused global feature. For example, taking an oral three-dimensional scanning device as an example, it can be determined whether the scanning environment type is an intraoral environment or an extraoral environment.

[0055] Alternatively, in some embodiments, feature extraction can be performed on the current frame image to obtain global features of the current frame image, and feature extraction can be performed on the reference image to obtain global features of each frame reference image. The global features of the reference image are fused into the global features of the current frame image to obtain a second fused global feature, and then it is determined whether there is fog on the lens module based on the second fused global feature.

[0056] Among them, feature extraction of the current frame image and the reference image to obtain global features can be achieved through some feature extraction networks, such as MobileNetV2 network, ResNet network, VGG network, Transform network, etc.

[0057] In some embodiments, the image detection method can be implemented by a pre-trained detection model. For example, the detection model can realize the detection of lens module dirtiness, or the detection model can realize two tasks of detecting lens module dirtiness and detecting lens fog simultaneously, or the detection model can realize two tasks of detecting lens dirtiness and detecting scanning environment types simultaneously, or the detection model can realize three tasks of detecting lens dirtiness, detecting lens fog and detecting scanning environment types simultaneously.

[0058] For example, taking the example of a detection model that simultaneously performs two tasks: detecting lens dirtiness and detecting the type of scanning environment, the detection model can be trained in the following manner: a current frame sample image and a reference sample image can be obtained, where the current frame sample image and the reference frame sample image are acquired by the same 3D scanning device, and the reference sample image includes multiple frames of images acquired before and / or after the current frame sample image. The current frame sample image carries a label that indicates whether the lens module that acquired the current frame sample image is dirty, and the type of scanning environment corresponding to the current frame sample image, i.e., the type of scanning environment in which the 3D scanning device was located when acquiring the current frame sample image.

[0059] Then the current frame sample image and the reference sample image can be input into a preset initial model, and the initial model performs feature extraction on the current frame sample image and the reference sample image respectively to obtain the shallow local feature map and global features of the current frame sample image, as well as the shallow local feature map and global features of the reference sample image. Then the shallow local feature map of the reference sample image is fused into the shallow local feature map of the current frame sample image to obtain a fused shallow local feature map (in order to distinguish it from the fused shallow local feature map in the reasoning stage, hereinafter referred to as the sample fused shallow local feature). The sample fused shallow local feature map and the global feature of the current frame sample image are fused to obtain a fused feature (in order to distinguish it from the fused feature in the reasoning stage, hereinafter referred to as the sample fused feature). Then, based on the sample fused feature, it can be determined whether the lens module includes dirt, and the first loss can be determined based on the difference between the determination result and the above-mentioned label.

[0060] At the same time, the initial model can fuse the global features of the current frame sample image with the global features of the reference sample image to obtain a first sample fusion global feature (hereinafter referred to as the first sample fusion global feature to distinguish it from the first fusion global feature in the inference stage), determine the type of scanning environment corresponding to the current frame sample image based on the first sample fusion global feature, and determine the second loss based on the difference between the determination result and the label. The model parameters of the initial model can then be adjusted based on the first loss and the second loss to train the above-mentioned detection model.

[0061] In some embodiments, as shown in FIG6 , the model of the initial model may include the following subnetworks: a feature extraction subnetwork, a shallow local feature fusion subnetwork, a first global fusion subnetwork, a second global fusion subnetwork, and a loss joint optimization subnetwork.

[0062] The feature extraction subnetwork is used to extract features from the current frame sample image and the reference sample image, respectively, to obtain shallow local feature maps and global features of the current frame sample image, and shallow local feature maps and global features of the reference sample image. The feature extraction subnetwork can be a MobileNetV2 network, a ResNet network, a VGG network, a Transform network, and so on.

[0063] The shallow local feature fusion subnetwork is used to fuse the shallow local feature map of the reference sample image into the shallow local feature map of the current frame sample image to obtain a sample fusion shallow local feature map.

[0064] The first global fusion sub-network is used to fuse the sample fusion shallow local feature map with the global feature of the current frame sample image to obtain the sample fusion feature.

[0065] The second global fusion subnetwork is used to fuse the global features of the current frame sample image with the global features of the reference sample image to obtain the first sample fusion global features.

[0066] A loss joint optimization subnetwork; used to determine whether the lens module includes dirt based on the sample fusion features, and determine a first loss based on the difference between the determination result and the above-mentioned label, and to determine the scanning environment type based on the first sample fusion global features, and determine a second loss based on the difference between the determination result and the above-mentioned label. Then, the model parameters of the initial model can be adjusted based on the first loss and the second loss to train the detection model.

[0067] In some embodiments, as shown in Figure 6, if the detection model can be used to simultaneously implement three tasks: detection of lens dirtiness, detection of lens fog, and detection of scanning environment type, then the label carried by the current frame image is also used to indicate whether the lens module that captures the current frame image includes fog. The initial model also includes: a third global fusion subnetwork: used to fuse the global features of the current frame sample image and the global features of the reference sample image to obtain a second sample fused global feature.

[0068] In addition to the aforementioned functions, this loss joint optimization subnetwork can also be used to determine whether there is fog on the lens module based on the second sample fusion global features, and determine the third loss based on the difference between the determination result and the label. The model parameters of the initial model are adjusted based on the first loss, the second loss, and the third loss to train the detection model. For example, different weights can be set for the above three losses, and the three losses are weighted and summed based on the weights to obtain the target loss. The model parameters of the initial model are then adjusted based on the target loss to train the detection model.

[0069] In the related art, three different detection models need to be pre-trained for the detection of lens dirt, lens fog, and scanning environment type, and are implemented separately by three different detection models. This processing method is less efficient. In the embodiment of the present application, it is taken into account that the three-dimensional scanning equipment usually needs to perform the above three types of detection before scanning the target object to be reconstructed. Only after determining that the lens is not dirty, not foggy, and the scanning environment type is determined, can the corresponding working mode be selected for subsequent scanning operations. Therefore, in the embodiment of the present application, it is thought of to combine the three detection tasks and implement them using one detection model. At the same time, in order to achieve the three detection tasks, a new detection model architecture is provided in the embodiment of the present application. As shown in Figure 6, the detection model includes a feature extraction subnetwork, a shallow local feature fusion subnetwork, a first global fusion subnetwork, a second global fusion subnetwork, a third global fusion subnetwork, and a loss joint optimization subnetwork. Among them, the features extracted by the feature extraction subnetwork are shared by the three detection tasks, thereby avoiding the extraction of features for the three detection tasks separately and improving the detection efficiency. The shallow local feature fusion subnetwork and the first global fusion subnetwork are used to detect dirt, the second global fusion subnetwork is used to detect the scanning environment type, and the third global fusion subnetwork is used to detect lens fog. The loss joint optimization subnetwork is used to determine the total loss based on the difference between the detection results of the three tasks and the true results (i.e., the results indicated by the labels). The model parameters are then adjusted based on the loss and trained. This approach not only achieves the completion of three tasks with a single model, but also improves detection efficiency.

[0070] In some embodiments, after determining the current scanning environment type of the 3D scanning device based on the first fused global feature, the current operating mode of the 3D scanning device can be switched to a target operating mode that matches the scanning environment type, or a user can be prompted to switch the current operating mode of the 3D scanning device to a target operating mode that matches the scanning environment type. When the 3D scanning device is in different operating modes, components on the 3D scanning device operate differently when capturing images using the 3D scanning device, and / or when three-dimensionally reconstructing a target object using images captured by the 3D scanning device, the image processing methods used vary. The processing methods include the processing steps involved in image processing, the processing algorithms employed, the processing parameters used during processing, and so on. The device operating methods can be operating parameters or states of various components in the 3D scanning device, such as the activation state of a fill light in the 3D scanning device, the brightness parameters of the fill light, or the activation state of an anti-fog module. In some embodiments, when a three-dimensional scanning device is scanning a target object to reconstruct a three-dimensional model of the target object, there are two scenarios: one is a scenario in which there are many impurities around the target object to be reconstructed, and the other is a scenario in which there are few or no impurities around the target object to be reconstructed. For example, taking an oral three-dimensional scanning device as an example, in a scenario in which the user's real teeth are scanned in the oral cavity, the captured image includes many impurities because the user's oral cavity includes many interferences around the teeth, such as the tongue, buccal soft tissue, and gums. In a scenario in which a tooth model is scanned outside the oral cavity, for example, a tooth model made of paraffin, metal, resin, etc., there is no interference around the teeth, so the captured image contains fewer impurities.

[0071] Therefore, the operating mode of the 3D scanning device can include a first operating mode corresponding to the first type of scenario described above, and a second operating mode corresponding to the second type of scenario described above. When the operating mode is the first operating mode, the processing steps for the images captured by the 3D scanning device during the 3D reconstruction process include a target step. When the operating mode is the second operating mode, the processing steps do not include a target step. The target step is used to identify data in the image that is irrelevant to the 3D reconstruction of the target object to be reconstructed, so that this irrelevant data can be ignored when generating the 3D model of the target object. For example, in an oral scan, irrelevant data may refer to areas such as the tongue, buccal soft tissue, and gums in the image. The image data corresponding to these areas in the image can be first identified and not used in the 3D reconstruction, thereby eliminating interference from this data in the subsequent 3D reconstruction process. In addition, a certain operating mode can be subdivided into multiple submodes, and the processing method can be adaptively adjusted for each submode. For example, when determining data that is irrelevant to the 3D reconstruction of the target object to be reconstructed (such as teeth), the judgment criteria can be adjusted.

[0072] In some embodiments, if it is determined based on the second fused global feature that there is fog on the lens module, the three-dimensional scanning device can be controlled to turn on the defog function. By automatically detecting whether there is fog on the lens module and automatically turning on the defog function of the three-dimensional scanning device when there is fog to remove the fog on the lens, it is possible to avoid the fog on the lens module affecting the quality of the collected image, thereby affecting the subsequent three-dimensional reconstruction.

[0073] In some embodiments, it is considered that some scanning environment types cover a variety of scenes, that is, they can be further divided into multiple subtypes. For example, taking the extraoral scanning environment type of an oral three-dimensional scanning device as an example, the extraoral environment is usually diverse. For example, in the scene of scanning a tooth model outside the mouth, since the material of the tooth model usually includes many categories, such as metal, resin, paraffin, etc., and the sample images used to train the detection model are difficult to cover all extraoral environment scenes, therefore, the detection model trained using the sample images is difficult to accurately predict some scenes that are not covered by the sample images, resulting in poor accuracy of the detection results of the pre-trained scanning environment detection model. In order to improve the accuracy of the detection model's prediction results of the scanning environment type, in some embodiments, when training the above-mentioned initial model to obtain the detection model, a contrast learning mechanism can be introduced, that is, when determining the loss of the initial model, a contrast loss can be introduced. By introducing the contrast loss, the initial model can be constrained so that when the initial model extracts features from the image, the features extracted are as close as possible for images of the same category (i.e., images of the same scanning environment type), and conversely, the features extracted are as far apart as possible for images of different categories.

[0074] For example, in some embodiments, after inputting a current sample image and a reference sample image into an initial model, the initial model may determine at least one sample image pair from the current sample image and the reference sample image, wherein the two sample images in the sample image pair correspond to the same scanning environment type. The initial model may then determine the similarity of the global features of the two sample images in the at least one sample image pair, and determine a fourth loss based on the difference between the global features and a preset similarity threshold. The model parameters of the initial model may then be adjusted based on the first loss, the second loss, and the fourth loss to train the detection model. Generally, if the two sample images correspond to the same scanning environment type, the features of the two sample images should also be similar. Therefore, a similarity threshold may be pre-set. If the two sample images correspond to the same scanning environment type, the similarity of the features of the two sample images should be very close to the similarity threshold. Therefore, the fourth loss may be determined based on the degree of similarity between the two features. The fourth loss is then used as a constraint to train the initial model, so that when the trained model extracts features from images with the same scanning environment type, the features extracted are more similar.

[0075] Of course, in some embodiments, if the detection model also has a lens fog detection function, the model parameters of the initial model can be adjusted in combination with the first loss, the second loss, the third loss and the fourth loss to train the detection model.

[0076] For example, different weights can be set for the above four losses, and the four losses can be weighted and summed based on the weights to obtain the target loss. Then, the model parameters of the initial model can be adjusted based on the target loss to train the detection model. For example, the target loss can be determined based on the following formula (1): Loss = a1×(La1+a2×La2)+a3×L b+a4*L c

[0077] Among them, La1 is the loss of the scanning environment type, a1 is the weight of the loss of the scanning environment type, La2 is the contrast loss of the scanning environment type, a2 is the weight of the contrast loss of the scanning environment type, Lb is the loss due to lens dirtiness, a3 is the weight of the loss due to lens dirtiness; Lc is the loss due to lens fogging, and a4 is the weight of the loss due to lens fogging.

[0078] In some embodiments, after the current frame sample image and the reference sample image are input into the initial model, the initial model can determine at least one group of sample image triplets from the current frame sample image and the reference sample image, each group of sample image triplets including a first sample image, a second sample image having the same scanning environment type as the first sample image, and a third sample image having a scanning environment type different from the scanning environment type of the first sample image. Feature extraction can then be performed on the first sample image, the second sample image, and the third sample image respectively to obtain their respective global features. Based on the similarity between the global features of the second sample image and the global features of the first sample image, and the similarity between the global features of the third sample image and the global features of the first sample image, the fifth loss can be determined. Then, the model parameters of the initial model are adjusted based on the first loss, the second loss, and the fifth loss to train the detection model. Among them, for two frames of sample images with the same scanning environment type, the similarity of their global features should be higher than the similarity of the global features of two frames of sample images with different scanning environment types. Based on this principle, the fifth loss is determined, and then the fifth loss is used to constrain the initial model to adjust the model parameters of the initial model. This can make the trained detection model more accurate in extracting features from the image, and thus the detection results for the scanning environment type are also more accurate.

[0079] Of course, in some embodiments, if the detection model also has a lens fog detection function, the model parameters of the initial model can be adjusted in combination with the first loss, the second loss, the third loss and the fifth loss to train the detection model.

[0080] In some embodiments, considering that the scenes corresponding to certain scanning environment types are often relatively similar and have little difference, taking the intraoral scanning environment type of an oral 3D scanning device as an example, considering that intraoral environments are often relatively similar and have little difference, the image features of images captured in the intraoral environment are often relatively similar, that is, the feature vectors of the images are distributed within a certain approximate range in the feature space. Therefore, for detecting a target environment (e.g., an intraoral environment) whose corresponding scenes are often relatively similar, a large number of sample images captured by the 3D scanning device in the target environment can be pre-acquired, and then feature extraction can be performed on the sample images to obtain global features of these sample images, and feature cluster centers of the global features of the sample images can be determined. For example, feature vectors representing the global features of the sample images can be obtained, and cluster centers of these global feature vectors can be determined. Then, the current scanning environment type of the 3D scanning device can be determined based on the current frame image captured by the 3D scanning device. For example, feature extraction can be performed on the current frame image to obtain global features of the current frame image, and then, based on the proximity between the features of the current frame image and the feature cluster centers, it can be determined whether the current scanning environment type of the 3D scanning device is the target environment. For example, if the feature of the current frame image is very close to the center of the feature cluster, it means that the scanning environment type corresponding to the current frame image is the target environment, where the degree of closeness can be expressed by the distance between the two, for example, it can be expressed by Euclidean distance, Manhattan distance, etc. If the distance between the two is less than a preset distance threshold, it is considered that the two are close, that is, the scanning environment type corresponding to the current frame image is determined to be the target environment.

[0081] Furthermore, an embodiment of the present application also provides an image detection method, the method comprising:

[0082] Obtain the image to be detected currently captured by the lens module on the 3D scanning device;

[0083] The image to be inspected is input into a pre-trained inspection model, and the inspection model is used to determine the current scanning environment type of the 3D scanning device. The inspection model is trained based on the following method:

[0084] Acquire at least two frames of sample images, where the scanning environments corresponding to the at least two frames of sample images are of the same type;

[0085] Using a preset initial model, feature extraction is performed on at least two frames of sample images to obtain global features of the at least two frames of sample images;

[0086] Based on the difference between the similarity of the global features of each of the sample images in at least two frames and a preset similarity threshold, a target loss is determined, and the model parameters of the initial model are adjusted based on the target loss to train a detection model.

[0087] Of course, the initial model can also be used to predict the scanning environment type corresponding to the sample image, and the total loss is obtained based on the difference between the predicted result and the actual result, as well as the above-mentioned target loss. The model parameters of the initial model are adjusted based on the total loss to train a detection model.

[0088] The specific training method of the detection model can refer to the description in the above embodiments and will not be repeated here.

[0089] Furthermore, an embodiment of the present application also provides an image detection method, the method comprising:

[0090] Obtain the image to be detected currently captured by the lens module on the 3D scanning device;

[0091] The image to be inspected is input into a pre-trained inspection model, and the inspection model is used to determine the current scanning environment type of the 3D scanning device. The inspection model is trained based on the following method:

[0092] Acquire sample image triplets, each sample image triplet comprising a first sample image, a second sample image having the same scanning environment type as the first sample image, and a third sample image having a different scanning environment type from the first sample image;

[0093] Extract features from the first sample image, the second sample image, and the third sample image using a preset initial model to obtain their respective global features;

[0094] Based on the similarity between the global features of the second sample image and the global features of the first sample image, and the similarity between the global features of the third sample image and the global features of the first sample image, a target loss is determined, and model parameters of the initial model are adjusted based on the target loss to train a detection model.

[0095] Of course, the initial model can also be used to predict the scanning environment type corresponding to the sample image, and the total loss is obtained based on the difference between the predicted result and the actual result, as well as the above-mentioned target loss. The model parameters of the initial model are adjusted based on the total loss to train a detection model.

[0096] The specific training method of the detection model can refer to the description in the above embodiments and will not be repeated here.

[0097] Furthermore, an embodiment of the present application also provides an image detection method, the method comprising:

[0098] Obtain the image to be detected currently captured by the lens module on the 3D scanning device;

[0099] Perform feature extraction on the image to be detected to obtain global features;

[0100] Based on the degree of proximity between the global features and the preset feature clustering center, it is determined whether the current scanning environment type of the three-dimensional scanning device is the target environment; wherein the feature clustering center is the clustering center of the global features of multiple frames of sample images, and the sample image is the image collected by the three-dimensional scanning device in the target environment.

[0101] The specific implementation details of the detection method can be referred to the description in the above embodiments and will not be repeated here.

[0102] It is not difficult to understand that the solutions described in the above embodiments can be freely combined to obtain new solutions when there is no conflict. Due to space reasons, they are not listed one by one in the embodiments of this disclosure.

[0103] Accordingly, an embodiment of the present disclosure further provides an image detection device, as shown in FIG7 , comprising:

[0104] An acquisition module 71 is configured to acquire a current frame image and a reference image captured by the lens module, wherein the reference image includes multiple frames of images captured before and / or after the current frame image;

[0105] A feature extraction module 72 is configured to perform feature extraction on the current frame image and the reference image respectively to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image;

[0106] a fusion module 73 configured to fuse the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein during the fusion process, a feature value at a target pixel position in the shallow local feature map of the current frame image is enhanced, and a similarity between the feature value at the target pixel position and a feature value at a corresponding pixel position at the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity;

[0107] The prediction module 74 is configured to determine whether there is dirt in the lens module based on the fused shallow local feature map.

[0108] The specific steps of the image detection method executed by the above-mentioned device can be referred to the description in the above-mentioned method embodiment, which will not be repeated here.

[0109] Furthermore, an embodiment of the present disclosure also provides an electronic device, as shown in Figure 8, the electronic device 80 includes a processor 81, a memory 82, and computer instructions stored in the memory 82 for execution by the processor 81, and when the processor 81 executes the computer instructions, it implements any method in the above embodiments.

[0110] The embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which implements the method of any of the aforementioned embodiments when the program is executed by a processor.

[0111] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0112] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment of the embodiments of the present disclosure or certain parts of the embodiments.

[0113] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0114] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, and the modules described as separate components may or may not be physically separated. When implementing the embodiment of the present disclosure, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0115] The above is only a specific implementation of the embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present disclosure. Industrial Applicability

[0116] In the image detection method provided by the present disclosure, a current frame image captured by a lens module and multiple reference images captured before and / or after the current frame image are captured are obtained, and a shallow local feature map of the current frame image and shallow local feature maps of each frame reference image are extracted. The shallow local feature maps of each frame reference image are then fused into the shallow local feature map of the current frame image. By selecting a specific fusion method, during the fusion process, the feature values ​​at positions with high similarity between the shallow local feature map of the current frame image and the shallow local feature map of the reference image are enhanced, thereby obtaining a fused shallow local feature map. Since the features of dirt become more prominent in the final fused shallow local feature map, the dirt detection result obtained based on the fused shallow local feature map is also more accurate, and has strong industrial applicability.

Claims

1. An image detection method, wherein: The method comprises: Acquire a current frame image and a reference image captured by the lens module, wherein the reference image includes multiple frame images captured before and / or after the current frame image; Performing feature extraction on the current frame image and the reference image respectively to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image; Fusion of the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein during the fusion process, the feature value at the target pixel position in the shallow local feature map of the current frame image is enhanced, and the similarity between the feature value at the target pixel position and the feature value at the corresponding pixel position at the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity; Determine whether there is dirt in the lens module based on the fused shallow local feature map.

2. The method according to claim 1, wherein: The determining whether there is dirt in the lens module based on the fused shallow local feature map comprises: extracting features from the current frame image to obtain global features of the current frame image; fusing the global features with the shallow local feature map to obtain fused features; determining whether there is dirt in the lens module based on the fused features; and / or The global feature is represented by a global feature vector, and the global feature is fused with the shallow local feature map to obtain a fused feature, including: performing pooling processing on the shallow local feature map to obtain a shallow local feature vector; concatenating the shallow local feature vector with the global feature vector to obtain the fused feature; or superimposing the shallow local feature vector with the global feature vector to obtain the fused feature.

3. The method according to claim 1, wherein: Fusion of the shallow local feature map of the reference image into the shallow local feature map of the current frame image, comprising: performing maximum pooling processing on the shallow local feature map of the reference image and the shallow local feature map of the current frame image respectively; fusing the shallow local feature map of the reference image after the maximum pooling processing into the shallow local feature map of the current frame image after the maximum pooling processing; and / or The shallow local feature map of the reference image is fused into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, including: fusing the shallow local feature map of each frame of the reference image with the shallow local feature map of the current frame image, wherein, for each pixel position in the shallow local feature map of the current frame image, the higher the similarity between the feature value of the pixel position and the feature value of the corresponding pixel position in the shallow local feature map of the reference image, the greater the fusion weight of the feature value of the pixel position; and superimposing the fused shallow local feature maps of each frame to obtain the fused shallow local feature map.

4. The method according to claim 1, wherein: The lens module is a lens module on a three-dimensional scanning device, and the method further includes: Extracting features of the current frame image to obtain global features of the current frame image; Extracting features from the reference image to obtain global features of the reference image; The global features of the reference image are fused into the global features of the current frame image to obtain a first fused global feature; based on the first fused global feature, the current scanning environment type of the three-dimensional scanning device is determined; and / or the global features of the reference image are fused into the global features of the current frame image to obtain a second fused global feature; based on the second fused global feature, it is determined whether there is fog on the lens module.

5. The method according to claim 4, wherein: After determining the current scanning environment type of the three-dimensional scanning device based on the first fused global feature, the method further includes: Switching the current working mode of the three-dimensional scanning device to a target working mode that matches the scanning environment type, or prompting the user to switch the current working mode of the three-dimensional scanning device to a target working mode that matches the scanning environment type; Among them, when the three-dimensional scanning device is in different working modes, in the process of three-dimensionally reconstructing the target object using the image captured by the three-dimensional scanning device, the processing method of the image is different, and / or, in the process of capturing the image using the three-dimensional scanning device, the operation mode of each component in the three-dimensional scanning device is different.

6. The method according to claim 4, wherein: If it is determined based on the second fused global feature that fog exists on the lens module, the three-dimensional scanning device is controlled to start a defog function.

7. The method according to claim 4, wherein: The method is performed by a pre-trained detection model, which is trained based on the following method: Acquire a current frame sample image and a reference sample image captured by a three-dimensional scanning device, wherein the reference sample image includes multiple frames of images captured before and / or after the current frame sample image; wherein the current frame sample image carries a label, and the label is used to indicate whether a lens module for capturing the current frame sample image is dirty, and the type of scanning environment corresponding to the current frame sample image; The current frame sample image and the reference sample image are input into a preset initial model, and the initial model performs the following operations: Performing feature extraction on the current frame sample image and the reference sample image respectively to obtain a shallow local feature map and a global feature of the current frame sample image, and a shallow local feature map and a global feature of the reference sample image; Fusing the shallow local feature map of the reference sample image into the shallow local feature map of the current frame sample image to obtain a sample fusion shallow local feature map, fusing the sample fusion shallow local feature map with the global feature of the current frame sample image to obtain a sample fusion feature, determining whether the lens module includes dirt based on the sample fusion feature, and determining a first loss based on a difference between a determination result and the label; Fusing the global features of the current frame sample image with the global features of the reference sample image to obtain a first sample fused global feature, determining the scanning environment type based on the first sample fused global feature, and determining a second loss based on a difference between the determination result and the label; The model parameters of the initial model are adjusted based on the first loss and the second loss to train the detection model.

8. The method according to claim 7, wherein: The initial model includes: A feature extraction subnetwork, used to extract features from the current frame sample image and the reference sample image respectively, to obtain a shallow local feature map and a global feature of the current frame sample image, and a shallow local feature map and a global feature of the reference sample image; Shallow local feature fusion subnetwork: used to fuse the shallow local feature map of the reference sample image into the shallow local feature map of the current frame sample image to obtain a sample fusion shallow local feature map; The first global fusion sub-network is used to fuse the sample fusion shallow local feature map with the global feature of the current frame sample image to obtain a sample fusion feature; The second global fusion subnetwork is used to fuse the global features of the current frame sample image with the global features of the reference sample image to obtain the first sample fusion global features; A loss joint optimization subnetwork; used to determine whether the lens module includes dirt based on the sample fusion features, and determine a first loss based on the difference between the determination result and the label; determine the scanning environment type based on the first sample fusion global features, and determine a second loss based on the difference between the determination result and the label information, and adjust the model parameters of the initial model based on the first loss and the second loss to train the detection model.

9. The method according to claim 8, wherein: The label is also used to indicate whether there is fog on the lens module. The initial model also includes: The third global fusion subnetwork is used to fuse the global features of the current frame sample image with the global features of the reference sample image to obtain the second sample fusion global features; The loss joint optimization subnetwork is also used to determine whether there is fog on the lens module based on the second sample fusion global feature, and determine the third loss based on the difference between the determination result and the label information, and adjust the model parameters of the initial model based on the first loss, the second loss and the third loss to train the detection model.

10. The method according to claim 7, wherein: After the current frame sample image and the reference sample image are input into the initial model, the initial model is further used to perform the following operations: determine at least one set of sample image pairs from the current frame sample image and the reference sample image, wherein the two frames of sample images in the sample image pairs correspond to the same scanning environment type; determine the similarity of global features of the two frames of sample images in the at least one set of sample image pairs, and determine a fourth loss based on the difference between the similarity and a preset similarity threshold; adjust the model parameters of the initial model based on the first loss, the second loss and the fourth loss to train and obtain the detection model; and / or After the current frame sample image and the reference sample image are input into the initial model, the initial model is further used to perform the following operations: determine at least one group of sample image triplets from the current frame sample image and the reference sample image; each group of sample image triplets includes a first sample image, a second sample image having the same scanning environment type as the first sample image, and a third sample image having a different scanning environment type from the first sample image; determine a fifth loss based on the similarity between the global features of the second sample image and the global features of the first sample image, and the similarity between the global features of the third sample image and the global features of the first sample image; adjust the model parameters of the initial model based on the first loss, the second loss and the fifth loss to train the detection model.

11. The method according to claim 1, wherein: The lens module is a lens module on a three-dimensional scanning device, and the method further includes: Extracting features of the current frame image to obtain global features of the current frame image; Based on the degree of proximity between the global feature and the preset feature clustering center, determine whether the scanning environment type when the three-dimensional scanning device collects the current frame image is the target environment; wherein the feature clustering center is the clustering center of the global features of multiple frames of sample images, and the sample image is an image collected by the three-dimensional scanning device in the target environment.

12. An image detection method, wherein: The method comprises: Obtain the image to be detected currently captured by the lens module on the 3D scanning device; The image to be detected is input into a pre-trained detection model, and the current scanning environment type of the three-dimensional scanning device is determined by the detection model; wherein the detection model is trained based on the following method: Acquire at least two frames of sample images, where the scanning environments corresponding to the at least two frames of sample images are of the same type; Using a preset initial model, respectively extracting features from the at least two frames of sample images to obtain global features of the at least two frames of sample images; Based on the difference between the similarity of the global features of each pair of sample images in the at least two frames of sample images and a preset similarity threshold, a target loss is determined, and the model parameters of the initial model are adjusted based on the target loss to train the detection model.

13. An image detection method, wherein: The method comprises: Obtain the image to be detected currently captured by the lens module on the 3D scanning device; The image to be detected is input into a pre-trained detection model, and the current scanning environment type of the three-dimensional scanning device is determined by the detection model; wherein the detection model is trained based on the following method: Acquire sample image triplets, each group of sample image triplets includes a first sample image, a second sample image having the same scanning environment type as that of the first sample image, and a third sample image having a different scanning environment type from that of the first sample image; Extracting features from the first sample image, the second sample image, and the third sample image respectively through a preset initial model to obtain respective global features; Based on the similarity between the global features of the second sample image and the global features of the first sample image, and the similarity between the global features of the third sample image and the global features of the first sample image, a target loss is determined, and the model parameters of the initial model are adjusted based on the target loss to train the detection model.

14. An image detection method, wherein: The method comprises: Obtain the image to be detected currently captured by the lens module on the 3D scanning device; Extracting features from the image to be detected to obtain global features; Based on the degree of proximity between the global feature and the preset feature clustering center, determine whether the current scanning environment type of the three-dimensional scanning device is the target environment; wherein the feature clustering center is the clustering center of the global features of multiple frames of sample images, and the sample images are images collected by the three-dimensional scanning device in the target environment.

15. An image detection device, wherein: The device comprises: An acquisition module is configured to acquire a current frame image and a reference image captured by the lens module, wherein the reference image includes multiple frame images captured before and / or after the current frame image; A feature extraction module is configured to perform feature extraction on the current frame image and the reference image respectively to obtain a shallow local feature map of the current frame image and a shallow local feature map of the reference image; A fusion module is configured to fuse the shallow local feature map of the reference image into the shallow local feature map of the current frame image to obtain a fused shallow local feature map, wherein during the fusion process, the feature value at the target pixel position in the shallow local feature map of the current frame image is enhanced, and the similarity between the feature value at the target pixel position and the feature value at the corresponding pixel position at the target pixel position in the shallow local feature map of the reference image is higher than a preset similarity; The prediction module is configured to determine whether there is dirt in the lens module based on the fused shallow local feature map.

16. An electronic device, wherein: The electronic device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, it implements the method described in any one of claims 1-14.

17. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.

Citation Information

Patent Citations

  • Camera stain detection method, device, device and storage medium

    CN109118498A

  • Image recognition method and device, electronic equipment and storage medium

    CN114283316A

  • Image detection method and device and electronic equipment

    CN117409285A

  • Three-dimensional scanning equipment control method, device, terminal equipment and system

    CN117414110A

  • Occlusion-robust visual object fingerprinting using fusion of multiple sub-region signatures

    US9483839B1

Cited By

  • Industrial defect classification detection method and device, medium and electronic equipment

    CN120374628A

  • Dental pulp state recognition model training method, electronic equipment and storage medium

    CN121147666A

  • Traffic image defogging method based on wavelet convolution and semantic-content guide fusion

    CN121353135A