Dirt detection method, electronic equipment and computer readable storage medium

By introducing an attention mechanism into the dirty detection method for multiple rounds of image feature encoding and cross-fusion, the problem of insufficient accuracy of dirty detection in the prior art is solved, and more efficient dirty detection is achieved.

CN120125855APending Publication Date: 2025-06-10ZHEJIANG HUARAY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510020629.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing dirty detection methods have limitations when detecting dirty areas with small areas and random locations, and are insufficient in accuracy.

Method used

The dirty detection method based on attention mechanism is adopted, through multiple rounds of image feature encoding, the encoding features matching multiple target dimensions are obtained, and cross-fusion is performed to improve the expression ability of decoding features.

Benefits of technology

It improves the accuracy of dirty detection and can more effectively detect dirty with small area and random locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125855A_ABST
    Figure CN120125855A_ABST
Patent Text Reader

Abstract

The invention discloses a smudginess detection method, electronic equipment and a computer readable storage medium. The method comprises the following steps: acquiring a to-be-detected image; carrying out image feature coding on the to-be-detected image based on an attention mechanism to obtain coding features matched with a plurality of target dimensions; wherein the attention mechanism is used for extracting reference features matched with different reference dimensions, and fusing the reference features; performing cross fusion on all the coding features to obtain a target fusion feature, and based on the target fusion feature, obtaining a decoding feature matched with each target dimension; and based on the decoding features matched with all the target dimensions, obtaining a smudginess detection result matched with the to-be-detected image. In this way, the accuracy of dirt detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a dirt detection method, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of machine vision technology, intelligent cameras have been widely used in various fields, including medical, security, transportation, industrial production, etc. However, during the assembly or application process, dirt inside and outside the camera can easily cause various noises in the captured images, seriously affecting the imaging quality. Therefore, the dirt detection of cameras is crucial. Currently, traditional dirt detection methods rely on manual analysis of the captured images to determine whether there is dirt inside and outside the camera; alternatively, existing neural network models are used to detect the captured images, but for dirt with a small area and random position, existing algorithms have limitations in detection.

[0003] In view of this, how to propose a dirt detection method with higher accuracy has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a dirt detection method, an electronic device, and a computer-readable storage medium, which can improve the accuracy of dirt detection.

[0005] To solve the above technical problem, a technical solution adopted by this application is: to provide a dirt detection method, including: obtaining an image to be detected; performing image feature encoding on the image to be detected based on an attention mechanism to obtain encoded features matching multiple target dimensions; wherein, the attention mechanism is used to extract reference features matching different reference dimensions and fuse the reference features; performing cross-fusion on all the encoded features to obtain a target fusion feature, and based on the target fusion feature, obtaining decoded features respectively matching each of the target dimensions; and based on the decoded features matching all the target dimensions, obtaining a dirt detection result matching the image to be detected.

[0006] To solve the above technical problem, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the method mentioned in the above technical solution.

[0007] To solve the above technical problem, another technical solution adopted by this application is: to provide a computer-readable storage medium, storing program instructions that can be run by a processor, and the program instructions are used for the method mentioned in the above technical solution.

[0008] The beneficial effects of the present application are as follows: Different from the prior art, the proposed dirt detection method in the present application obtains encoded features matching the corresponding target dimensions through multiple rounds of encoding using the attention mechanism. In each encoding round, the attention mechanism fuses the reference features matching different reference dimensions to enhance the expressive ability of the obtained encoded features. Furthermore, all the encoded features are cross-fused, so that each decoded feature obtained combines the associated features between different encoded features, further enhancing the expressive ability of the decoded features and contributing to improving the accuracy of the dirt detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0010] Figure 1 is a schematic flowchart of an implementation manner of the dirt detection method of the present application;

[0011] Figure 2 is Figure 1 a schematic flowchart of another implementation manner corresponding to step S102 in

[0012] Figure 3 is a schematic structural diagram of an implementation manner of the dirt detection model of the present application;

[0013] Figure 4 is Figure 2 a schematic flowchart of another implementation manner corresponding to step S202 in

[0014] Figure 5 is a schematic structural diagram of an implementation manner of the attention network of the present application;

[0015] Figure 6 is Figure 4 a schematic flowchart of another implementation manner corresponding to step S303 in

[0016] Figure 7 is Figure 1 a schematic flowchart of another implementation manner corresponding to step S103 in

[0017] Figure 8 is Figure 7 a schematic flowchart of another implementation manner corresponding to step S501 in

[0018] Figure 9 is Figure 1 a schematic flowchart of another implementation manner corresponding to step S101 in

[0019] Figure 10 It is a schematic structural diagram of an embodiment of the electronic device of the present application;

[0020] Figure 11 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Specific embodiments

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments, and different embodiments can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the dirt detection method of the present application. The method includes:

[0023] S101: Obtain an image to be detected.

[0024] In one embodiment, an image to be detected for dirt detection is obtained so that the image to be detected is processed through subsequent steps to identify the dirt in the image to be detected.

[0025] In one implementation scenario, to detect whether there is dirt on the lens of the camera acquisition device, which affects the effect of image acquisition, the image to be detected is an image acquired in real time by the camera acquisition device.

[0026] In another embodiment, to save the consumption of computing resources, for the video stream acquired in real time by the camera acquisition device, a corresponding video frame is extracted from the video stream at a preset time interval as the image to be detected.

[0027] S102: Perform image feature encoding on the image to be detected based on the attention mechanism to obtain encoded features matching multiple target dimensions; wherein, the attention mechanism is used to extract reference features matching different reference dimensions and fuse the reference features.

[0028] In one embodiment, image feature encoding is performed on the image to be detected to extract encoded features matching multiple different target dimensions.

[0029] Specifically, multiple rounds of feature extraction are performed on the image to be detected in sequence, and the input of the current encoding round is the output of the previous encoding round. For the output of the previous encoding round, the current encoding round uses an attention mechanism to obtain reference features matching different reference dimensions, and by fusing the reference features corresponding to the current encoding round, the encoding features corresponding to the current encoding round are obtained. Among them, for the first encoding round, the corresponding encoding features are obtained by performing convolution processing on the image to be detected.

[0030] S103: Cross-fuse all the encoding features to obtain the target fusion features, and based on the target fusion features, obtain the decoding features respectively matching each target dimension.

[0031] In one embodiment, all the obtained encoding features are cross-fused to combine the associated features between different encoding features, so as to obtain the target fusion features. Decompose the target fusion features to obtain the decoding features respectively matching each target dimension. Among them, each decoding feature includes the associated features between the encoding features corresponding to the corresponding target dimension and other encoding features.

[0032] S104: Based on the decoding features matching all the target dimensions, obtain the dirt detection result matching the image to be detected.

[0033] In one embodiment, decode and classify by combining the decoding features matching all the target dimensions to obtain the dirt detection result matching the image to be detected.

[0034] Specifically, the above dirt detection result is the target detection image corresponding to the image to be detected, and the target detection image includes the marked dirty areas.

[0035] The dirt detection method proposed in this application obtains the encoding features matching the corresponding target dimensions by performing multiple rounds of encoding using the attention mechanism. In each encoding round, the attention mechanism fuses the reference features matching different reference dimensions to enhance the expression ability of the obtained encoding features. Furthermore, all the encoding features are cross-fused, so that each obtained decoding feature combines the associated features between different encoding features, further enhancing the expression ability of the decoding features, which helps to improve the accuracy of the dirt detection result.

[0036] Please refer to Figure 2 , Figure 2 is Figure 1 the schematic flowchart of another embodiment corresponding to step S102 in

[0037] S201: Input the image to be detected into the first feature encoding module in the stain detection model, and use the first feature encoding module to perform convolution processing on the image to be detected to obtain the encoded features output by the first encoded feature module.

[0038] In one embodiment, the stain detection model includes a plurality of sequentially connected feature encoding modules. For the first feature encoding module, the image to be detected is used as the input to obtain the corresponding encoded features by using the first feature encoding module to process the image to be detected.

[0039] Specifically, the first feature encoding module in the stain detection model includes at least one convolutional layer. By using the convolutional layer in the first encoded feature module to perform convolution processing on the image to be detected, the corresponding encoded features are obtained.

[0040] In a specific application scenario, the first feature encoding module in the stain detection model includes two 3×3 convolutional layers coupled to each other. The image to be detected is input into the first 3×3 convolutional layer, and the output of this convolutional layer is used as the input to the next 3×3 convolutional layer, and finally the encoded features output by the first feature encoding module model are obtained.

[0041] S202: Input the encoded features output by the previous feature encoding module into the current encoding module, and use the attention network in the current feature encoding module to output the corresponding encoded features until the encoded features output by each feature encoding module are obtained; wherein, the attention network includes a first extraction branch and a second extraction branch for feature extraction in different reference dimensions.

[0042] In one embodiment, the remaining feature encoding modules in the stain detection model except the first feature encoding module include an attention network. For the remaining feature encoding modules in the stain detection model, the encoded features output by the previous feature encoding module are input into the current encoding module to perform further feature extraction by using the first extraction branch and the second extraction branch in the corresponding attention network.

[0043] Among them, in the current feature encoding module, the first extraction branch and the second extraction branch in the attention network are respectively used to perform feature extraction on the encoded features output by the previous feature encoding module in different reference dimensions, so as to fuse the features in different reference dimensions obtained by extraction to obtain encoded features with stronger expression ability.

[0044] In a specific application scenario, please refer to Figure 3 , Figure 3It is a schematic structural diagram of an embodiment of the dirt detection model of the present application. The dirt detection model includes a plurality of feature encoding modules stacked in sequence. Among them, the first feature encoding module includes two 3×3 convolutional layers coupled to each other, and is used to extract features from the image to be detected to obtain encoded features in the corresponding target dimension. The remaining feature encoding modules each include a corresponding attention network, which is used to further extract features from the encoded features output by the previous feature encoding module. In addition, the remaining feature encoding modules each further include a max pooling layer coupled to the corresponding attention network. The max pooling layer is used to perform downsampling on the output result of the corresponding attention network, so as to output encoded features matching the corresponding target dimension. It should be noted that the convolutional layer structure in the first feature encoding module can also be other, Figure 3 only a plurality of feature encoding modules are schematically drawn, and in actual applications, the number of feature encoding modules can be set according to actual needs.

[0045] Please refer to Figure 4 and Figure 5 , Figure 4 is Figure 2 a schematic flow diagram of another embodiment corresponding to step S202 in Figure 5 It is a schematic structural diagram of an embodiment of the attention network of the present application. Specifically, the implementation process of step S202 includes:

[0046] S301: Input the encoded features output by the previous feature encoding module into the first convolutional layer in the attention network to obtain the initial features output by the first convolutional layer.

[0047] In one embodiment, for the current feature encoding module, input the encoded features output by the previous feature encoding module into the first convolutional layer in the attention network, so as to use the first convolutional layer to reduce the dimension of the input encoded features without changing the depth of the input encoded features, thereby obtaining the initial features to reduce the amount of calculation.

[0048] In one implementation scenario, the above first convolutional layer is a 1×1 convolutional layer structure.

[0049] S302: Input the initial features into the first extraction branch and the second extraction branch in the attention network respectively to obtain the first reference features output by the first extraction branch that match the first reference dimension and the second reference features output by the second extraction branch that match the second reference dimension.

[0050] In one embodiment, input the obtained initial features into the first extraction branch and the second extraction branch in the attention network respectively, so as to use the first extraction branch to extract the first reference features that match the first reference dimension and use the second extraction branch to extract the second reference features that match the second reference dimension.

[0051] Specifically, the first extraction branch and the second extraction branch respectively include corresponding convolutional layer structures, and there are differences between the convolutional layer structures of the first extraction branch and the second extraction branch, so that the first reference dimension matching the first reference feature is different from the second reference dimension matching the second reference feature.

[0052] In a specific application scenario, the first extraction branch includes a 3×3 convolutional layer, and the first reference feature is obtained by using the 3×3 convolutional layer to extract the initial feature. The second extraction branch includes two stacked 3×3 convolutional layers in sequence, and the input of the second 3×3 convolutional layer is the output of the previous 3×3 convolutional layer. The second reference feature is obtained by using the two 3×3 convolutional layers to extract the initial feature.

[0053] S303: Based on the first reference feature and the second reference feature, obtain the target extraction feature output by the attention network.

[0054] In one embodiment, according to the first reference feature and the second reference feature, obtain the target extraction feature output by the attention network of the current feature encoding module.

[0055] In one implementation scenario, the first reference feature and the second reference feature are concatenated to obtain the target extraction feature.

[0056] In another embodiment, the obtained first reference feature and second reference feature are feature fused to obtain the target extraction feature. Among them, the target extraction feature includes the correlation feature between the first reference feature and the second reference feature.

[0057] S304: Perform downsampling processing on the target extraction feature to obtain the encoded feature corresponding to the target dimension output by the current feature encoding module.

[0058] In one embodiment, perform downsampling processing on the obtained target extraction feature to obtain the encoded feature corresponding to the target dimension.

[0059] Specifically, the current feature encoding module further includes a max pooling layer coupled to the attention network, and the max pooling layer is used to perform downsampling processing on the target extraction feature output by the corresponding attention network to obtain the encoded feature matching the corresponding target dimension.

[0060] In the above solution, by using the attention network to combine the first reference feature and the second reference feature matching different reference dimensions, the encoded feature matching the current feature encoding module is obtained, which improves the expression ability of the encoded feature, thereby helping to improve the accuracy of subsequent dirt detection.

[0061] Please refer to Figure 6 , Figure 6 is Figure 4The flowchart diagram corresponding to step S303 in another embodiment. Specifically, the implementation process of step S303 includes:

[0062] S401: Respectively perform convolution processing on the first reference feature and the second reference feature to obtain a first processed feature corresponding to the first reference feature and a second processed feature corresponding to the second reference feature.

[0063] In one embodiment, convolution processing is respectively performed on the obtained first reference feature and second reference feature to obtain corresponding first and second processed features.

[0064] Specifically, the attention network further includes at least one stacked convolutional sub-network and a normalization layer. By sequentially inputting the first reference feature into the above-mentioned convolutional sub-network, the convolutional sub-network is used to perform deeper feature extraction and normalization processing on the first reference feature to obtain the corresponding first processed feature. Similarly, the second reference feature is input into the above-mentioned convolutional sub-network and normalization layer to obtain the corresponding second processed feature.

[0065] In a specific application scenario, please continue to refer to Figure 6 , the attention network includes two stacked convolutional sub-networks, and each convolutional sub-network includes a convolutional layer, a batch normalization layer, and a rectified linear units layer (ReLU layer) sequentially coupled. Among them, the convolutional layer in the convolutional sub-network is a 3×3 convolutional layer structure for local feature extraction of the input feature; the batch normalization layer is used to improve the stability of the model structure; the rectified linear units layer (ReLU layer) is used to enhance the expression ability of the model.

[0066] It should be noted that the number of convolutional sub-networks in the attention network and the specific structure of the convolutional layer in the convolutional sub-network can also be other, and can be specifically set according to actual needs; for example, the number of convolutional sub-networks can also be 1 or 3, etc., and the convolutional layer in the convolutional sub-network can also be a 1×1 convolutional layer structure.

[0067] S402: Concatenate the first processed feature and the first reference feature to obtain a first concatenated feature; and, concatenate the second processed feature and the second reference feature to obtain a second concatenated feature.

[0068] In one embodiment, the obtained first processed feature and the first reference feature are concatenated to obtain a first concatenated feature to improve the expression ability of the first concatenated feature. And, the obtained second processed feature and the second reference feature are concatenated to obtain a second concatenated feature to improve the expression ability of the second concatenated feature.

[0069] S403: Based on the first concatenated feature and the second concatenated feature, obtain the target extraction feature.

[0070] In one embodiment, the first splicing feature and the second splicing feature are fused to obtain a target extraction feature.

[0071] Specifically, the first splicing feature and the second splicing feature are spliced, and the obtained feature after splicing is used as the above-mentioned target extraction feature.

[0072] In another embodiment, the implementation process of step S403 may further include: fusing the first splicing feature and the second splicing feature to obtain a first fusion feature. Fusing the first fusion feature and the encoded feature output by the previous feature encoding module to obtain a target extraction feature.

[0073] Specifically, the first splicing feature and the second splicing feature are spliced and then processed by convolution of a convolutional layer to obtain a first fusion feature. According to the obtained first fusion feature and the input of the corresponding current feature encoding module, that is, splicing the first fusion feature and the encoded feature output by the previous feature encoding module, to obtain a target extraction feature.

[0074] In a specific application scenario, the feature obtained after splicing the first splicing feature and the second splicing feature is input into a 1×1 convolutional layer to obtain the first fusion feature output by this convolutional layer.

[0075] In the above solution, after obtaining the first reference feature and the second reference feature, further feature extraction and processing are respectively performed to improve the expression ability of the finally obtained target extraction feature.

[0076] Please refer to Figure 7 , Figure 7 which Figure 1 is the schematic flowchart of another embodiment corresponding to step S103 in

[0077] S501: Input the encoded features matching multiple target dimensions into the connection module in the dirt detection model, and use the connection module to perform cross-fusion on all encoded features to obtain a target fusion feature.

[0078] In one embodiment, the dirt detection model further includes a connection module. Input the encoded features matching multiple target dimensions into the above-mentioned connection module to extract the correlation features between different encoded features by using the connection module, so as to obtain a target fusion feature.

[0079] S502: Decompose the target fusion feature into decoded features matching each feature decoding module in the dirt detection model.

[0080] In one embodiment, the dirt detection model includes a plurality of feature decoding modules, and the plurality of feature decoding modules correspond one-to-one with a plurality of feature encoding modules. The target fusion feature is split to obtain decoding features that match each feature decoding module.

[0081] In a specific application scenario, the dirt detection model further includes a plurality of feature decoding modules. The number of feature decoding modules is the same as the number of feature encoding modules, and the feature decoding modules are used to decode the input decoding features. Please continue to refer to Figure 3 , and the dimension of the image to be detected is H×W×C. Based on this, the dimension of the encoded feature output by the first feature encoding module A is The dimension of the encoded feature output by the feature encoding module B is The dimension of the encoded feature output by the feature encoding module C is The dimension of the encoded feature output by the feature encoding module D is The dimension of the encoded feature output by the feature encoding module E is After splitting the obtained target fusion, decoding features that match each feature decoding module are obtained; among them, the dimension of the decoding feature that matches the feature decoding module e is The feature decoding module e decodes and upsamples the decoding feature to obtain a corresponding output feature, and inputs the output feature to the next feature decoding module d. The dimension of the decoding feature that matches the feature decoding module d is The feature decoding module d uses the corresponding decoding feature and the output feature of the previous feature decoding module e to obtain a corresponding output feature. Repeat the above decoding steps until a dirt detection result consistent with the dimension of the image to be detected and output by the feature decoding module a is obtained.

[0082] In another embodiment, for each feature decoding module, the attention mechanism is used to process the features input to the feature decoding module to obtain corresponding output features. Among them, the features input to the current feature decoding module include the features output by the previous feature decoding module and the decoding features that match the current feature decoding module.

[0083] Specifically, the feature decoding module includes an attention network, and the attention network in the feature decoding module includes a first decoding branch and a second decoding branch. For the current feature decoding module, the input features are concatenated and then input into the convolutional layer in the attention network to obtain the corresponding initial decoded features. The initial decoded features are respectively input into the first decoding branch and the second decoding branch in the attention network to obtain a third reference feature output by the first decoding branch that matches the third reference dimension and a fourth reference feature output by the second decoding branch that matches the fourth reference dimension. Based on the third reference feature and the fourth reference feature, a reference decoded feature is obtained. By performing upsampling processing on the reference decoded feature, the output feature corresponding to the current feature decoding module is obtained. Among them, the step of outputting the corresponding output feature by using the attention network in the current decoding module can refer to the process of obtaining the encoded features in the corresponding above embodiments, and will not be elaborated in detail here.

[0084] Please refer to Figure 8 , Figure 8 is Figure 7 a schematic flowchart of another embodiment corresponding to step S501 in

[0085] S601: Flatten the encoded features that match multiple target dimensions, and concatenate the flattened encoded features to obtain a target concatenated feature.

[0086] In one embodiment, for the encoded features obtained in the corresponding above embodiments, the multi-dimensional encoded features are converted into one-dimensional flattened features. The flattened features corresponding to all the encoded features are concatenated to obtain a target concatenated feature.

[0087] S602: Input the target concatenated feature into the feature analysis network, and use the feature analysis network to extract features from the fused feature to obtain a target fused feature.

[0088] In one embodiment, the obtained target concatenated feature is input into the feature analysis network to perform semantic analysis on the target concatenated feature by using the feature analysis network, so as to extract the corresponding target fused feature.

[0089] In one implementation scenario, the above feature analysis network is a Transformer model. Among them, the Transformer model has better parallel computing capabilities, long-range dependence modeling, global context modeling, scalability, and generalization capabilities, etc.

[0090] In the above solution, the encoded features are flattened so that the resulting target concatenated features can be used as the input to the subsequent feature analysis network. Moreover, by using the feature analysis network deployed in the connection module to process the target concatenated features obtained by concatenating multiple flattened features, the correlation features between the encoded features matching different target dimensions are extracted, thereby improving the expression ability of the target fusion features.

[0091] Please refer to Figure 9 , Figure 9 which Figure 1 is the schematic flowchart of another implementation manner corresponding to step S101. Specifically, the implementation process of step S101 includes:

[0092] S701: Obtain an initial image.

[0093] In one implementation manner, an initial image is obtained, and after subsequent noise reduction processing, the initial image is used as the image to be detected for dirt detection.

[0094] In one implementation scenario, the above initial image is an image collected in real time by a camera acquisition device. Alternatively, the initial image can also be an image extracted from the video stream collected by the camera acquisition device.

[0095] S702: Obtain a trained noise cancellation model, input the initial image into the noise cancellation model, and use the noise cancellation model to obtain the noise information corresponding to the initial image.

[0096] In one implementation manner, a trained noise cancellation model is obtained, and the initial image is input into the noise cancellation model to extract the noise information in the initial image by using the noise cancellation model.

[0097] In one implementation scenario, the above noise cancellation model is a flow-based neural network (Flow-Based Image Denoising Neural Network, FDN). This model is trained by using multiple training samples, and each training sample includes a noisy image containing noise and a target image after noise removal.

[0098] In another implementation manner, the structure of the above noise cancellation model can be other neural network structures.

[0099] S703: Use the initial image with the noise information removed as the image to be detected.

[0100] In one implementation manner, the extracted noise information is removed from the initial image to obtain the image to be detected after denoising processing.

[0101] In the above solution, since the camera acquisition device often easily generates complex-structured noise during image or video acquisition, and some noise changes with temperature, by obtaining the initial image and using the trained noise elimination model to remove the noise in the initial image, the influence of the noise in the initial image on subsequent dirt detection is reduced, and the accuracy of dirt detection is improved.

[0102] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device includes: a memory 10 and a processor 20 that are coupled to each other. Program instructions are stored in the memory 10, and the processor 20 is configured to execute the program instructions to implement the methods mentioned in any of the above embodiments. Specifically, the electronic device includes, but is not limited to: desktop computers, laptop computers, tablet computers, servers, etc., which are not limited herein. In addition, the processor 20 may also be referred to as a CPU (Center Processing Unit, central processing unit). The processor 20 may be an integrated circuit chip with signal processing capabilities. The processor 20 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 20 may be implemented jointly by integrated circuit chips.

[0103] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. A program instruction 40 capable of being run by a processor is stored on the computer-readable storage medium 30, and when the program instruction 40 is executed by the processor, the methods mentioned in any of the above embodiments are implemented.

[0104] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0105] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0106] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0107] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0108] The above is only the embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.

Claims

1. A dirt detection method, characterized in that: include: Acquire the image to be detected; Based on the attention mechanism, image feature encoding is performed on the image to be detected to obtain encoding features matching multiple target dimensions; wherein the attention mechanism is used to extract reference features matching different reference dimensions and fuse the reference features; Cross-fusing all the encoding features to obtain target fusion features, and obtaining decoding features that match each of the target dimensions based on the target fusion features; Based on the decoded features matched by all the target dimensions, a dirt detection result matching the image to be detected is obtained.

2. The method according to claim 1, characterized in that The image feature encoding of the image to be detected based on the attention mechanism to obtain encoding features matching multiple target dimensions includes: Inputting the image to be detected into the first feature encoding module in the dirt detection model, performing convolution processing on the image to be detected using the first feature encoding module, and obtaining the encoding features output by the first feature encoding module; The coding features output by the previous feature coding module are input into the current coding module, and the corresponding coding features are output using the attention network in the current feature coding module until the coding features output by each feature coding module are obtained; wherein the attention network includes a first extraction branch and a second extraction branch, and the first extraction branch and the second extraction branch are used for feature extraction of different reference dimensions.

3. The method according to claim 2, characterized in that The step of inputting the encoding features output by the previous feature encoding module into the current encoding module and using the attention network in the current feature encoding module to output corresponding encoding features includes: Inputting the encoded features output by the previous feature encoding module into the first convolutional layer in the attention network to obtain the initial features output by the first convolutional layer; Inputting the initial features into the first extraction branch and the second extraction branch in the attention network respectively, obtaining a first reference feature output by the first extraction branch that matches the first reference dimension and a second reference feature output by the second extraction branch that matches the second reference dimension; Based on the first reference feature and the second reference feature, obtaining a target extraction feature output by the attention network; The target extraction features are downsampled to obtain the encoding features of the corresponding target dimensions output by the current feature encoding module.

4. The method according to claim 3, characterized in that The step of obtaining the target extraction feature output by the attention network based on the first reference feature and the second reference feature includes: Performing convolution processing on the first reference feature and the second reference feature respectively to obtain a first processing feature corresponding to the first reference feature and a second processing feature corresponding to the second reference feature; Splicing the first processing feature and the first reference feature to obtain a first splicing feature; and splicing the second processing feature and the second reference feature to obtain a second splicing feature; The target extraction feature is obtained based on the first splicing feature and the second splicing feature.

5. The method according to claim 4, characterized in that The obtaining the target extraction feature based on the first splicing feature and the second splicing feature includes: Fusing the first splicing feature and the second splicing feature to obtain a first fused feature; The first fusion feature is fused with the encoding feature output by the previous feature encoding module to obtain the target extraction feature.

6. The method according to claim 2, characterized in that The cross-fusion of all the encoding features to obtain a target fusion feature, and based on the target fusion feature, obtaining a decoding feature that matches each of the target dimensions, includes: Inputting the coding features matched by multiple target dimensions into the connection module in the dirt detection model, and using the connection module to cross-fuse all the coding features to obtain the target fusion features; The target fusion feature is decomposed into the decoded feature that matches each feature decoding module in the dirt detection model.

7. The method according to claim 6, characterized in that The step of inputting the coded features matched with multiple target dimensions into a connection module in the dirt detection model, and using the connection module to cross-fuse all the coded features to obtain the target fusion features includes: Flattening the coding features matched by multiple target dimensions, and splicing the flattened coding features to obtain target splicing features; The target splicing feature is input into a feature analysis network, and the feature analysis network is used to extract the fusion feature to obtain the target fusion feature.

8. The method according to claim 1, characterized in that The step of acquiring the image to be detected comprises: Get the initial image; Obtaining a trained noise elimination model, inputting the initial image into the noise elimination model, and using the noise elimination model to obtain noise information corresponding to the initial image; The initial image after removing the noise information is used as the image to be detected.

9. An electronic device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Camera smudginess detection method and device, electronic equipment and storage medium

    CN121280857A