Image recognition method, electronic equipment and storage medium

By combining multi-level feature extraction and attention modules, the problem of insufficient feature information in small target image recognition is solved, and the accuracy and precision of image recognition are improved.

CN120707918APending Publication Date: 2025-09-26ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510601496.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image processing solutions have limited feature information when processing small target objects, resulting in low image recognition accuracy, especially in low-resolution situations where details are easily lost.

Method used

A multi-level feature extraction method is adopted, and the image recognition model is used to extract features of the image to be identified. At least two different feature extraction methods are used to enhance the feature expression ability of the initial feature map, and the attention module is combined to improve the accuracy of the image recognition results.

Benefits of technology

By combining multi-level feature extraction and attention modules, the accuracy of small target image recognition is improved, the feature expression ability is enhanced, and the precision of image recognition results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707918A_ABST
    Figure CN120707918A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method, electronic equipment and a storage medium. The image recognition method comprises the following steps: acquiring a to-be-recognized image related to an indicator lamp; multi-level feature extraction is carried out on the to-be-recognized image through the image recognition model, an initial feature image set is obtained, the initial feature image set comprises at least two initial feature images, the sizes of the initial feature images are different, and at least two different feature extraction modes are adopted in at least part of level feature extraction. And based on each initial feature map, determining an image recognition result of the to-be-recognized image, the image recognition result including at least one of position information of the indicating lamp and a classification result of the indicating lamp. According to the scheme, the accuracy of the image recognition result obtained by image recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image recognition method, electronic device, and storage medium. Background Art

[0002] Existing image processing solutions primarily rely on the traditional HSV color model or object detection technology to process and recognize targets. However, for smaller objects (e.g., indicator lights), image processing can hinder the effectiveness of image processing for small objects due to the small number of pixels covered. This leads to limited feature information, especially at low resolutions where details are easily lost. This can affect image processing for small objects, resulting in low image processing and recognition accuracy. Therefore, existing solutions for image processing and recognition of small objects remain challenging.

[0003] In view of the existing technical defects, how to provide an effective image recognition solution is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] This application at least provides an image recognition method, an electronic device, and a storage medium.

[0005] The present application provides an image recognition method, comprising: obtaining an image to be recognized associated with an indicator light. Using an image recognition model, performing multi-level feature extraction on the image to be recognized to obtain an initial feature atlas, wherein the initial feature atlas includes at least two initial feature maps, and the scales of the initial feature maps are different, wherein at least two different feature extraction methods are used in at least some of the hierarchical feature extractions. Based on each initial feature map, an image recognition result for the image to be recognized is determined, wherein the image recognition result includes at least one of the position information of the indicator light and the classification result of the indicator light.

[0006] The present application provides an image recognition device, comprising: an acquisition module, a feature extraction module, and a determination module. The acquisition module is used to acquire an image to be recognized related to an indicator light. The feature extraction module is used to perform multi-level feature extraction on the image to be recognized using an image recognition model to obtain an initial feature atlas, wherein the initial feature atlas includes at least two initial feature maps and the scales of the initial feature maps are different, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction. The determination module is used to determine the image recognition result of the image to be recognized based on each initial feature map, wherein the image recognition result includes at least one of the position information of the indicator light and the classification result of the indicator light.

[0007] The present application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the above-mentioned image recognition method.

[0008] The present application provides a computer-readable storage medium having program instructions stored thereon, which implement the above-mentioned image recognition method when the program instructions are executed by a processor.

[0009] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0012] Figure 1 This is a first flow chart of an embodiment of the image recognition method of the present application.

[0013] Figure 2 This is a second flow chart of an embodiment of the image recognition method of the present application.

[0014] Figure 3 This is a third flow chart of an embodiment of the image recognition method of the present application.

[0015] Figure 4 This is a fourth flow chart of an embodiment of the image recognition method of the present application.

[0016] Figure 5 This is a fifth flow chart of an embodiment of the image recognition method of the present application.

[0017] Figure 6 Schematic diagram of an edge perception module in an embodiment of the image recognition method of the present application.

[0018] Figure 7 This is a sixth flow chart of an embodiment of the image recognition method of the present application.

[0019] Figure 8a Schematic diagram of a sample image according to an embodiment of the image recognition method of the present application.

[0020] Figure 8b It is a schematic diagram of a sample binary image of an embodiment of the image recognition method of the present application.

[0021] Figure 9 Schematic diagram of a feature point detection head according to an embodiment of the image recognition method of the present application.

[0022] Figure 10 Schematic diagram of an image recognition model in one embodiment of the image recognition method of the present application.

[0023] Figure 11 It is a structural diagram of an embodiment of the image recognition device of the present application.

[0024] Figure 12 It is a structural diagram of an embodiment of the electronic device of the present application.

[0025] Figure 13 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0026] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0027] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0028] The term "and / or" in this article is simply a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0029] The present application provides some image recognition methods and image recognition devices. The application scenario of the image recognition method includes image recognition of images containing small targets. Specifically, for example, the small target can be an indicator light. The execution subject of the image recognition method can be any device with image recognition function, for example, an image recognition device. For example, the image recognition device can be set in a terminal device or a server or other processing device, wherein the terminal device can be a device for image recognition, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementations, the image recognition method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0030] See also Figure 1 , Figure 1 This is a first flow chart of an embodiment of the image recognition method of the present application. Specifically, the image recognition method may include the following steps:

[0031] Step S11: Acquire an image to be recognized related to the indicator light.

[0032] The indicator light may be a traffic signal light that needs to be collected at a long distance in a traffic scene. In other application scenarios, the indicator light may be a signal light on the road in an autonomous driving scene. In other application scenarios, the indicator light may be a car signal light on the road in an aerial photography scene. The number of indicator lights may be several. Several may refer to one or more. When the number of indicator lights is multiple, it may refer to multiple indicator lights belonging to the same category, or it may refer to multiple indicator lights belonging to different categories. Exemplarily, the indicator light may be an indicator light on a detection device in an industrial control scene. For example, the indicator light may be one or more indicator lights on a detection device. The detection device may be a device capable of realizing industrial control. This application takes the example of the number of indicator lights being several, which will not be elaborated here. The image to be identified may be an image collected of the indicator light. The image to be identified may be an image containing the indicator light or an image collected of the area where the indicator light is located.

[0033] In some application scenarios, step S11 may be performed by capturing the image to be identified in real time using an image acquisition device. In other application scenarios, step S11 may be performed by selecting any image to be identified from a set of images to be identified related to the indicator light. The set of images to be identified may be images captured by the image acquisition device at a historical moment.

[0034] Step S12: Use the image recognition model to perform multi-level feature extraction on the image to be recognized to obtain an initial feature atlas.

[0035] The image recognition model can be a model with image recognition capabilities deployed on an image recognition device. The image recognition model can be pre-trained. The input of the image recognition model can be an image to be recognized, and the output of the image recognition model can be the image recognition result of the image to be recognized. Multi-level feature extraction can be multi-scale feature extraction of the image to be recognized, with each feature extraction level corresponding to one of the multi-scale feature extraction methods. It is understood that the image recognition model includes a backbone network. The backbone network can also be referred to as a backbone network. The backbone network is used to perform multi-level feature extraction on the image to be recognized input into the image recognition model to obtain an initial feature map set. The backbone network includes several feature extraction levels (i.e., the backbone network includes several extraction modules), and the several extraction modules include at least two extraction modules, each of which is used to extract features from the input image / feature map. Each of the several extraction modules is used to perform feature extraction on the input image to be recognized / the initial feature map output by the previous feature extraction module to obtain the initial feature map output by the current feature extraction module. It is understood that the extraction module can refer to a feature extraction module, which will not be further described below.

[0036] At least partial hierarchical feature extraction can refer to a feature extraction method that can be implemented by at least some of the extraction modules in a number of extraction modules. It is understandable that the extraction modules in a number of extraction modules may use the same / different feature extraction methods. At least some of the extraction modules in a number of extraction modules may use the same preset feature extraction method, and at least two different feature extraction methods may be used in the preset feature extraction method. Exemplarily, the extraction module is used to extract features from the input image to be identified / the initial feature map output by the previous extraction module. Each extraction module in the backbone network corresponds to a preset feature extraction method. The preset feature extraction method can be implemented by a preset feature extraction network set on the extraction module. Specifically, the preset feature extraction network can be at least two of feature extraction networks such as a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, and an attention mechanism. In some application scenarios, the preset feature extraction network set on at least some of the extraction modules in a number of extraction modules can be a long short-term memory network (LSTM) in a recurrent neural network and a preset perceptron network. In other application scenarios, the preset feature extraction network may be a feature extraction network corresponding to the attention mechanism or other feature extraction networks.

[0037] The initial feature atlas includes at least two initial feature maps, each with a different scale. At least two different feature extraction methods are used in at least some of the hierarchical feature extractions. In some application scenarios, step S12 may involve directly inputting the image to be identified into an image recognition model to obtain an initial feature atlas corresponding to the image to be identified, as output by the backbone network of the image recognition model. In other application scenarios, step S12 may involve first performing a preset test on the image to be identified to obtain a preset test result. The preset test result indicates whether the image to be identified is qualified. A qualified image indicates a quality assessment result for the image to be identified, and the quality assessment result indicates whether the image to be identified is an image capable of generating a valid image recognition result. For example, if the image to be identified does not contain an indicator light or the indicator light is not located in a preset area, the image to be identified is unqualified, and the image recognition result for the image to be identified is therefore of low reference value. If the preset test result indicates a qualified image, the image to be identified is input into the backbone network to obtain an initial feature atlas corresponding to the image to be identified, as output by the backbone network.

[0038] It is understood that each image to be identified can correspond to an initial feature atlas. In some application scenarios, when there are multiple indicator lights, each image to be identified can correspond to the initial feature atlas of all indicator lights, or to the initial feature atlas corresponding to each indicator light. This application uses the example of each image to be identified corresponding to the initial feature atlas of all indicator lights.

[0039] Step S13: Determine the image recognition result of the image to be recognized based on each initial feature map.

[0040] The image recognition result includes at least one of the position information of the indicator light and the classification result of the indicator light. The position information of the indicator light may refer to the position of the indicator light in the image to be recognized. In the case where there are multiple indicator lights, the position information of the indicator light can distinguish the indicator lights at different positions in the image to be recognized. In some application scenarios, the classification result of the indicator light is used to indicate that the indicator light is in any one of the preset working states. The preset working state may include the indicator light being on and the indicator light being off. In other application scenarios, the classification result of the indicator light is used to indicate that the indicator light is in any one of the preset attribute information. The preset attribute information includes several preset colors of the indicator light. In other application scenarios, the classification result of the indicator light is related to the preset working state and preset attribute information of the indicator light. In other application scenarios, the image recognition device can determine whether the indicator light is in a flashing state and the flashing frequency in combination with continuous frames of the image to be recognized.

[0041] In some application scenarios, the above step S13 may be a step of using a determination module in an image recognition model to implement an image recognition result of the image to be recognized based on each initial feature map. The initial feature map set is used as the input of the determination module in the image recognition model to obtain the image recognition result of the image to be recognized. In other application scenarios, the above step S13 may be a step of using at least part of the initial feature maps in the initial feature map set as the input of the determination module in the image recognition model to obtain the image recognition result of the image to be recognized. Each initial feature map corresponds to a candidate image recognition result. The candidate image recognition result with the highest confidence is selected from the candidate image recognition results corresponding to each initial image feature as the image recognition result of the image to be recognized.

[0042] It is understandable that each image to be recognized corresponds to an image recognition result. In the case where there are multiple indicator lights, the image recognition result output by the image recognition model includes the recognition results of each indicator light.

[0043] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0044] See also Figure 2 , Figure 2 This is a second flow chart of an embodiment of the image recognition method of the present application.

[0045] In some embodiments, the image recognition model includes several extraction modules arranged in cascade, the several extraction modules include at least one intermediate extraction module, each intermediate extraction module includes a first module and a second module arranged in cascade, one of the first module and the second module is a preset feature extraction module, and the other is an attention module. The above-mentioned step S12 may include the following steps: Step S21: Use the first extraction module to perform the first preset feature extraction on the image to be recognized to obtain an initial feature map. Step S22: Input the initial feature map into each intermediate extraction module in turn to obtain the initial feature map output by each intermediate extraction module. Step S23: Input the initial feature map output by the last intermediate extraction module into the last extraction module among the several extraction modules to perform the second preset feature extraction to obtain the initial feature map output by the last extraction module.

[0046] The backbone network in the image recognition model also includes several extraction modules arranged in cascade. The several extraction modules include a first extraction module, at least one intermediate extraction module, and a last extraction module arranged in cascade. It is understandable that the several extraction modules are cascaded in the order in which the features are input in sequence. The types and / or quantities of feature extraction networks set on the first extraction module and the last extraction module may be the same or different. It is understandable that setting different numbers and / or different types of feature extraction networks in the extraction module can realize processing corresponding to different feature extraction methods for the input of the extraction module.

[0047] In each intermediate extraction module, the above-mentioned preset feature extraction method can be used to perform feature extraction on the feature map output by the first extraction module / the feature map output by the previous intermediate extraction module to obtain the initial feature map output by the intermediate extraction module. The above-mentioned preset feature extraction method can be a feature extraction method corresponding to at least two of the feature extraction networks such as convolutional neural network, recurrent neural network, long short-term memory network, gated recurrent unit, BiGRU neural network, attention mechanism, etc. It can be understood that at least some of the feature extraction methods adopted by each intermediate extraction module can be the same or different. In the case that at least some of the feature extraction methods adopted by each intermediate extraction module are the same, the order of at least two feature extraction methods used by each intermediate extraction module can be the same or different. This application takes the case where the feature extraction methods adopted by each intermediate extraction module are the same and each intermediate extraction module uses two feature extraction methods as an example, which will not be repeated here.

[0048] Specifically, each intermediate extraction module includes a first module and a second module arranged in cascade. One of the first module and the second module is a preset feature extraction module, and the other is an attention module. Among them, the setting order of the first module and the second module in each intermediate extraction module is not limited. The setting order of the first module and the second module may be different in different intermediate extraction modules. The types of feature extraction networks set on the first module and the second module are different. The preset feature extraction module can be provided with any one of the feature extraction networks such as the above-mentioned convolutional neural network, recurrent neural network, long short-term memory network, gated recurrent unit, BiGRU neural network, etc. The attention module can be provided with the above-mentioned attention mechanism.

[0049] In some application scenarios, the input feature map of the current intermediate extraction module is used as the input of the first module, the output of the first module is used as the input of the second module, and the output of the second module is used as the output of the current intermediate feature extraction module. In other application scenarios, the input feature map of the current intermediate extraction module is used as the input of the second module, the output of the second module is used as the input of the first module, and the output of the first module is used as the output of the current intermediate feature extraction module.

[0050] In some application scenarios, the first preset feature extraction in step S21 above may be feature extraction that can be performed by the feature extraction network set on the first extraction module. The image to be identified is input into the first extraction module for first preset feature extraction, and the first extraction network outputs a starting feature map. In some application scenarios, the first extraction module may be provided with a convolutional neural network. In other application scenarios, the first extraction module may be a submodule of the backbone network in a preset target detection model. The preset target detection model may be a YOLO series model, such as a YOLOv8 model. In some application scenarios, the second preset feature extraction in step S23 above may be feature extraction that can be performed by the feature extraction network set on the last extraction module. Specifically, the second preset feature extraction may be a feature extraction method corresponding to at least two different feature extraction networks set on the last extraction module. For example, the last extraction module includes a convolutional neural network and a pooling layer. For example, the pooling layer may be a pyramid pooling (SPP) layer for pooling the feature map output by the convolutional neural network in the last extraction module.

[0051] The backbone network includes at least one intermediate extraction module in a cascade arrangement. It is understandable that each intermediate extraction module is cascaded in the order in which the features are input in sequence. In some application scenarios, when there is only one intermediate extraction module, the above step S22 may be inputting the starting feature map into the intermediate extraction module to obtain the initial feature map output by the intermediate extraction module. The above step S23 may be using the initial feature map output by the intermediate extraction module as the input of the last extraction module to perform a second preset feature extraction, to obtain the initial feature map output by the last extraction module. In other application scenarios, when there are multiple intermediate extraction modules, the above step S22 may be inputting the starting feature map into the first intermediate extraction module to obtain the initial feature map output by the first intermediate extraction module. For the last intermediate extraction module, the initial feature map output by the previous intermediate extraction module of the last intermediate extraction module is used as the input of the last intermediate extraction module to obtain the initial feature map output by the last intermediate extraction module. The above step S23 may be to use the initial feature map output by the last intermediate extraction module as the input of the last extraction module to perform the second preset feature extraction to obtain the initial feature map output by the last extraction module.

[0052] In some application scenarios, the first module in each intermediate extraction module is an attention module, and the second module is a preset feature extraction module, but the order of the cascade setting of the first module and the second module in each intermediate extraction module is not limited.

[0053] See also Figure 3 , Figure 3 This is a third flow chart of an embodiment of the image recognition method of the present application.

[0054] In some embodiments, the first module is an attention module, and the second module is a preset feature extraction module. The above step S22 may include the following steps: for each intermediate extraction module, perform the following Figure 3 The following steps are shown: Step S31: Input the candidate feature map into the first module of the current intermediate extraction module to obtain the advanced feature map corresponding to the candidate feature map. Step S32: Input the advanced feature map into the second module of the current intermediate extraction module to perform the third preset feature extraction to obtain the initial feature map corresponding to the current intermediate extraction module.

[0055] It is understood that this application uses the example of the first module in each intermediate extraction module being an attention module, the second module being a preset feature extraction module, and the order of the cascade arrangement of the first and second modules in each intermediate extraction module being the same, and will not be repeated here. The intermediate extraction modules have the same structure, but the parameters of the intermediate extraction modules can be the same or different.

[0056] The candidate feature map includes the starting feature map or the initial feature map output by the previous intermediate extraction module. The advanced feature map can be a feature map obtained by processing the input candidate feature map using the attention mechanism matched by the attention module in the current intermediate extraction module. In some application scenarios, the third preset feature extraction in the above step S32 can be a feature extraction that can be performed by the feature extraction network set on the second module in the current intermediate extraction module.

[0057] In some application scenarios, when the current intermediate extraction module is the first intermediate extraction module, the candidate feature map is the starting feature map output by the first extraction module. The above step S31 can input the starting feature map into the first module in the first intermediate extraction module to obtain the advanced feature map corresponding to the starting feature map output by the first module in the first intermediate extraction module. The above step S32 can be inputting the advanced feature map corresponding to the starting feature map into the second module in the first intermediate extraction module to perform a third preset feature extraction, obtaining the feature map output by the second module in the first intermediate extraction module, and using the feature map output by the second module in the first intermediate extraction module as the initial feature map output by the first intermediate extraction module.

[0058] In other application scenarios, when the current intermediate extraction module is not the first intermediate extraction module, the candidate feature map is the initial feature map output by the previous intermediate extraction module. The above step S31 can input the initial feature map output by the previous intermediate extraction module into the first module in the non-first intermediate extraction module to obtain the advanced feature map output by the first module in the non-first intermediate extraction module. The above step S32 can be inputting the advanced feature map output by the first module into the second module in the non-first intermediate extraction module to perform a third preset feature extraction, obtaining the feature map output by the second module in the non-first intermediate extraction module, and using the feature map output by the second module in the non-first intermediate extraction module as the initial feature map output by the non-first intermediate extraction module.

[0059] In some embodiments, step S31 may include the following steps: first, obtaining the similarity between each pixel in the candidate feature map. Then, based on each similarity, determining the attention weight corresponding to the candidate feature map. Then, combining the candidate feature map and the attention weight to obtain an advanced feature map corresponding to the candidate feature map.

[0060] Similarity can refer to the similarity between the target pixel and other pixels in the candidate feature map.

[0061] The above-mentioned method of obtaining the similarity between each pixel point in the candidate feature map can be to use at least one similarity calculation method to determine the similarity between each pixel point in the candidate feature map. The at least one similarity calculation method includes determining the Euclidean distance, hash similarity, cosine similarity, etc. between each pixel point in the candidate feature map. In the case where there are multiple similarity calculation methods, the similarities determined by the different similarity calculation methods are averaged to obtain a new similarity, and the new similarity is used as the similarity between each pixel point in the candidate feature map.

[0062] Exemplarily, the above-mentioned attention module may be a SimAM module. The attention weight corresponding to the candidate feature map includes the weight of each pixel in the candidate feature map, which is used to represent the local importance of each pixel in the candidate feature map. In some application scenarios, the above-mentioned method of determining the attention weight corresponding to the candidate feature map based on each similarity may be to directly use the similarity between the target pixel and other pixels in the candidate feature map as the attention weight of the target pixel in the candidate feature map, or to normalize the similarity between the target pixel and other pixels and use the normalized result as the attention weight of the target pixel. In other application scenarios, the above-mentioned method of determining the attention weight corresponding to the candidate feature map based on each similarity may be to pre-process the similarity between the target pixel and other pixels and use the result of the preset processing as the attention weight of the target pixel.

[0063] In some application scenarios, the method of combining the candidate feature map with the attention weight to obtain the advanced feature map corresponding to the candidate feature map can be to fuse at least part of the candidate feature map with the attention weight of at least part of the target pixels in the candidate feature map to obtain a new feature map, and use the new feature map as the advanced feature map. In other application scenarios, the method of combining the candidate feature map with the attention weight to obtain the advanced feature map corresponding to the candidate feature map can be to multiply the candidate feature map with the attention weight of each target pixel in the candidate feature map to obtain a new feature map, and use the new feature map as the advanced feature map. It can be understood that the w in the SimAM module i is the attention weight of each pixel, which reflects the similarity between the point and other features, w i As a weight applied to the candidate feature map, a new feature map is generated. The value of each pixel in the new feature map is the original feature value of the target pixel in the candidate feature map and its corresponding attention weight w i This can enhance the expression of important information in the feature map while suppressing unimportant information. SimAM calculates weights based on local self-similarity, without introducing additional parameters and complex calculations.

[0064] Specifically, the process of determining the attention weight corresponding to the candidate feature map based on each similarity can refer to the following formula (1):

[0065]

[0066] Among them, N i It can represent the set of pixels in the candidate feature map. i It can represent the attention weight of the i-th pixel in any candidate feature map. k can represent the normalized weight. s(f i -f j ) can represent the similarity between the i-th pixel and the j-th pixel in any candidate feature map. For example, the similarity here uses the simple and effective Euclidean distance.

[0067] The input of the last extraction module among the several extraction modules is the initial feature map output by the last intermediate extraction module in the at least one intermediate extraction module mentioned above. The last extraction module is used to perform a second preset feature extraction on the initial feature map output by the last intermediate extraction module to obtain the initial feature map output by the last extraction module. In some application scenarios, the second preset feature extraction in the above step S23 may be a feature extraction that can be performed by the feature extraction network set on the last extraction module. Specifically, the second preset feature extraction may be a feature extraction method corresponding to at least two different feature extraction networks set on the last extraction module. For example, the last extraction module includes a convolutional neural network and a pooling layer. For example, the pooling layer may be a pyramid pooling SPP layer for performing pooling processing on the feature map output by the convolutional neural network in the last extraction module.

[0068] It can be understood that when at least part of the feature extraction methods in the first preset feature extraction, the second preset feature extraction, and the third preset feature extraction all include convolutional neural networks, the structures and / or model parameters of the convolutional neural networks in different modules may be the same or different.

[0069] It is understandable that in order to improve the expression of features, the present application introduces an attention module (i.e., SimAM module) in the backbone network (i.e., backbone) in the image recognition model. The introduction of the attention module can enhance the feature expression ability of the image recognition model without increasing the number of image recognition model parameters, which can avoid increasing time-consuming. In addition, the addition of the attention module, especially in the scenario of small target detection, can improve the detection accuracy of the image recognition model for small targets. The attention module generates 3D attention weights by calculating the local self-similarity of the input feature map, and applies these weights to the original feature map to enhance its expression ability. The whole process does not add any additional network parameters, and is simple to implement and easy to integrate into the corresponding convolutional neural network architecture in the backbone network and neck stage.

[0070] See also Figure 4 , Figure 4 This is a fourth flow chart of an embodiment of the image recognition method of the present application.

[0071] In some embodiments, step S13 may include the following steps: Step S41: performing feature fusion on each initial feature map to obtain at least one initial fused feature map. Step S42: determining a target feature map corresponding to each initial feature map based on the last initial feature map in the initial feature map set and each initial fused feature map. Step S43: determining an image recognition result for the image to be recognized based on each target feature map.

[0072] The initial feature maps in the initial feature map set are sorted according to the time in which the extraction modules in the backbone network output the initial feature maps. The first initial feature map is the initial feature map output by the first intermediate extraction module. The last initial feature map can be the initial feature map output by the last extraction module.

[0073] The at least one initial fused feature map may be one or more. The initial fused feature map may represent the fusion result obtained by fusing features of at least some of the initial feature maps. Feature fusion may be performed by direct multiplication, concatenation, or fusion based on the weights of the initial feature maps. In some application scenarios, step S41 may involve fusing features of any two of the initial feature maps to obtain at least one initial fused feature map. In other application scenarios, step S41 may involve fusing features of two adjacent initial feature maps to obtain at least one initial fused feature map. In other application scenarios, the initial feature maps in the initial feature map set are sorted according to the time at which the initial feature maps are output by the extraction modules in the backbone network. Step S41 may involve upsampling the last initial feature map to obtain the last upsampled initial feature map. The candidate initial feature maps in the candidate initial feature map set are respectively used as the non-last initial feature maps in the initial feature map set. The candidate initial feature map set is sorted according to the time at which the initial feature maps are output by the extraction modules in the backbone network. The first candidate initial feature map is the last candidate initial feature map. The first candidate initial feature map is fused with the last initial feature map after upsampling to obtain the first initial fused feature map. The feature map in the candidate initial feature map set that is adjacent to the first candidate initial feature map and whose output time of each extraction module is earlier is used as the second candidate initial feature map. The second candidate initial feature map is fused with the first fused feature map after upsampling to obtain the second initial fused feature map. This process is repeated until all candidate initial feature maps in the candidate initial feature map set are processed, and the first initial fused feature map, the second initial fused feature map, and the Nth initial fused feature map are collected to obtain the initial fused feature map set.

[0074] See also Figure 5 , Figure 5 This is a fifth flow chart of an embodiment of the image recognition method of the present application.

[0075] In some embodiments, the above step S41 may include the following steps: for each non-last initial feature map in the initial feature map set, perform the following steps: Figure 5The following steps are shown: Step S51: Obtain the first feature map to be fused corresponding to the current non-last initial feature map. Step S52: Perform a first fusion process on the current non-last initial feature map and the first feature map to be fused to obtain a candidate fused feature map. Step S53: Input the candidate fused feature map into the preset attention module to obtain the features output by the preset attention module, and use the features output by the preset attention module as the initial fused feature map corresponding to the current non-last initial feature map.

[0076] The first feature map to be fused includes the last initial feature map after upsampling, or a feature map obtained by upsampling the initial fused feature map corresponding to the next non-last initial feature map. The non-last initial feature maps in the initial feature map set are the initial feature maps other than the last initial feature map in the initial feature map set. The initial feature maps in the initial feature map set are sorted according to the time in which the initial feature maps were output by the extraction modules in the backbone network. The next non-last initial feature map is the non-last initial feature map in the initial feature map set that is adjacent to the current non-last initial feature map and whose output by the extraction modules precedes the current non-last initial feature map. Specifically, the candidate initial feature maps in the candidate initial feature map set are respectively used as the non-last initial feature maps in the initial feature map set. The candidate initial feature map set is sorted according to the time in which the initial feature maps were output by the extraction modules in the backbone network. The first fusion process may be performing a residual connection or residual processing on the at least two input feature maps. In some application scenarios, the preset attention module may be a module having the same structure and / or parameters as the attention module in the backbone network. In other application scenarios, the preset attention module can be a module with the same function as the attention module in the above-mentioned backbone network.

[0077] In some application scenarios, the above step S51 may be: when the current non-last initial feature map is the last candidate initial feature map, upsampling the last initial feature map to obtain the last initial feature map after upsampling, and using the last initial feature map after upsampling as the first feature map to be fused corresponding to the current non-last initial feature map. The above step S52 may be: when the current non-last initial feature map is the last candidate initial feature map, performing a first fusion process on the last candidate initial feature map and the last initial feature map after upsampling to obtain a candidate fused feature map corresponding to the last candidate initial feature map. The above step S53 may be: inputting the candidate fused feature map corresponding to the last candidate initial feature map into a preset attention module to obtain the features output by the preset attention module, and using the features output by the preset attention module as the initial fused feature map corresponding to the last candidate initial feature map, and the initial fused feature map corresponding to the last candidate initial feature map is the initial fused feature map corresponding to the current non-last initial feature map.

[0078] In some application scenarios, the above step S51 may be: when the current non-last initial feature map is the non-last candidate initial feature map, the initial fusion feature map corresponding to the next non-last initial feature map is upsampled to obtain the first feature map to be fused corresponding to the current non-last initial feature map. The above step S52 may be: when the current non-last initial feature map is the non-last candidate initial feature map, the non-last candidate initial feature map is first fused with the last initial feature map after upsampling to obtain the candidate fusion feature map corresponding to the non-last candidate initial feature map. The above step S53 may be: inputting the candidate fusion feature map corresponding to the non-last candidate initial feature map into the preset attention module, obtaining the features output by the preset attention module, and using the features output by the preset attention module as the initial fusion feature map corresponding to the non-last candidate initial feature map, and the initial fusion feature map corresponding to the non-last candidate initial feature map is the initial fusion feature map corresponding to the current non-last initial feature map.

[0079] It is understandable that the above steps S41 and S42 are executed by the neck stage (i.e., neck) in the image recognition model, and the above step S43 is executed by the head stage (i.e., head) in the image recognition model. It can be considered that in order to improve the expression of features, the present application introduces an attention module (i.e., SimAM module) into the backbone network (i.e., backbone) and the neck stage (i.e., neck) in the image recognition model. The introduction of the attention module can enhance the feature expression capability of the image recognition model without increasing the number of image recognition model parameters, which can avoid increasing time consumption. In addition, the addition of the attention module, especially in the scenario of small target detection, can improve the detection accuracy of the image recognition model for small targets. The attention module generates 3D attention weights by calculating the local self-similarity of the input feature map, and applies these weights to the original feature map to enhance its expression capability. The whole process does not add any additional network parameters, and is simple to implement and easy to integrate into the corresponding convolutional neural network architecture in the backbone network and the neck stage.

[0080] Step S42: Based on the last initial feature map in the initial feature map set and each initial fusion feature map, determine the target feature map corresponding to each initial feature map.

[0081] The last initial feature map and each initial fusion feature map are used as the feature maps to be processed in the feature map set to be processed corresponding to each initial feature map. The order of the feature maps to be processed in the feature map set to be processed is based on the time sequence in which the image recognition model obtains the feature maps to be processed. The later the image recognition model obtains any feature map to be processed, the higher the order of the feature map to be processed in the feature map set to be processed.

[0082] In some application scenarios, step S42 may be performed by fusing features of any two of the feature maps to be processed to obtain target feature maps corresponding to the initial feature maps. In other application scenarios, step S42 may be performed by fusing features of two adjacent feature maps in each feature map to be processed to obtain target feature maps corresponding to the initial feature maps.

[0083] In some embodiments, step S42 may include the following steps: first, the last initial feature map and each initial fused feature map are used as the feature maps to be processed corresponding to each initial feature map. Then, for the first feature map to be processed, the first feature map to be processed is downsampled to obtain a target feature map corresponding to the first feature map to be processed. Then, for each non-first feature map to be processed, the second feature map to be fused is fused with the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed.

[0084] The first feature map to be processed may be the initial fused feature map corresponding to the first initial feature map in the feature map set to be processed. Each non-first feature map to be processed includes other feature maps to be processed in the feature map set except the first feature map to be processed. Specifically, each non-first feature map to be processed includes at least the last initial feature map. The last feature map to be processed may be the last initial feature map in the feature map set to be processed. The non-last feature map to be processed adjacent to the last feature map to be processed may be the initial fused feature map adjacent to the last initial feature map in the feature map set to be processed, that is, the initial fused feature map corresponding to the last candidate initial feature map when the current non-last initial feature map is the last candidate initial feature map. The second feature map to be fused includes the target feature map corresponding to the first feature map to be processed or the target feature map corresponding to the previous non-first feature map to be processed.

[0085] In some application scenarios, the processing steps for each non-first feature map to be processed may be: directly fusing the second feature map to be fused with the current non-first feature map to be processed to obtain the target feature map corresponding to the current non-first feature map to be processed. Specifically, when the current non-first feature map to be processed is the non-first feature map to be processed adjacent to the first feature map to be processed in the feature map set to be processed, the second feature map to be fused is the target feature map corresponding to the first feature map to be processed. Directly fusing the target feature map corresponding to the first feature map to be processed with the current non-first feature map to be processed to obtain the target feature map corresponding to the current non-first feature map to be processed. The fusion process may be a residual connection / jump connection. Specifically, when the current non-first feature map to be processed is the non-first feature map to be processed that is not adjacent to the first feature map to be processed in the feature map set to be processed, the second feature map to be fused is the target feature map corresponding to the previous non-first feature map to be processed. The previous non-first feature map to be processed may be a feature map that is generated later in the image recognition model than the current feature map to be processed. The target feature map corresponding to the previous non-first feature map to be processed is directly fused with the current non-first feature map to be processed to obtain the target feature map corresponding to the current non-first feature map to be processed. The fusion process can be a residual connection / skip connection.

[0086] In other application scenarios, the processing steps for each non-first feature map to be processed may include: performing a preset processing on the second feature map to be fused to obtain a new feature map to be fused. The preset processing may include sequentially inputting the second feature map to be fused into at least one convolution layer for convolution processing to obtain a new feature map to be fused. The new feature map to be fused is then fused with the current non-first feature map to be processed to obtain a target feature map corresponding to the current non-first feature map to be processed.

[0087] In some embodiments, for each non-first feature map to be processed, the step of fusing the second feature map to be fused with the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed includes: for the last feature map to be processed, fusing the last feature map to be processed, the feature map to be processed adjacent to the last feature map to be processed, and the second feature map to be fused to obtain a target feature map corresponding to the last feature map to be processed. For the non-last feature map to be processed, fusing the non-last feature map to be processed and the second feature map to be fused to obtain a target feature map corresponding to the last feature map to be processed.

[0088] In some embodiments, the above-mentioned step of fusing the second feature map to be fused with the non-first feature map to be processed for each non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed may include the following steps: first, performing a preset convolution process on the second feature map to be fused to obtain a first-order feature map corresponding to the second feature map to be fused. Then, performing information fusion processing based on the first-order feature map and edge information to obtain a second-order feature map corresponding to the second feature map to be fused. Subsequently, performing a preset fusion process on the second-order feature map and the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed.

[0089] For example, the preset convolution processing may be to perform convolution processing on the input feature map using a convolution kernel of a fixed size (such as 1×1, 3×3, 5×5). The second feature map to be fused after the preset convolution processing is used as the first-order feature map corresponding to the second feature map to be fused. The edge information is obtained by extracting the edge information of the first-order feature map. The edge information of the first-order feature image is obtained by extracting the edge information of the first-order feature map using the edge perception module. Exemplarily, the edge perception module may be a Sobel operator. Specifically, the process of extracting the edge information of the first-order feature map to obtain the edge information of the first-order feature image may be: performing convolution processing on the first-order feature map in at least one preset direction to obtain the edge information of the first-order feature image. Specifically, the at least one preset direction may be a horizontal direction and a vertical direction.

[0090] The preset fusion processing may be performing the above-mentioned residual connection / jump connection on the input data. In some application scenarios, the step of performing information fusion processing based on the first-order feature map and edge information to obtain the second-order feature map corresponding to the second feature map to be fused may be: directly splicing the first-order feature map with the edge information corresponding to the first-order feature map to obtain the second-order feature map corresponding to the second feature map to be fused. The step of performing preset fusion processing on the second-order feature map and the non-first feature map to be processed to obtain the target feature map corresponding to the non-first feature map to be processed may be: directly performing residual connection on the second-order feature map and the non-first feature map to be processed to obtain the target feature map corresponding to the non-first feature map to be processed. In other application scenarios, the step of performing information fusion processing based on the first-order feature map and edge information to obtain the second-order feature map corresponding to the second feature map to be fused may be: splicing the first-order feature map with the edge information corresponding to the first-order feature map to obtain the second-order feature map corresponding to the second feature map to be fused. Perform preset convolution processing on the second-order feature map to obtain a new second-order feature map. It is understandable that the preset convolution processing performed on the second-order feature map and the preset convolution processing performed on the second feature map to be fused can be implemented by using the same / different convolution kernels for convolution operations. The step of performing the preset fusion processing on the second-order feature map and the non-first feature map to be processed to obtain the target feature map corresponding to the non-first feature map to be processed can be: performing a residual connection on the new second-order feature map and the non-first feature map to be processed to obtain the target feature map corresponding to the non-first feature map to be processed.

[0091] See also Figure 6 , Figure 6 Schematic diagram of an edge perception module in an embodiment of the image recognition method of the present application.

[0092] The neck stage in the image recognition model includes an edge perception module. The input of the edge perception module is the second feature map to be fused. The output of the edge perception module can be the above-mentioned second-order feature map or the above-mentioned new second-order feature map. This application takes the output of the edge perception module as the above-mentioned new second-order feature map as an example. It can be understood that since the shape of the indicator light can be circular or square, adding an edge perception module to the feature fusion process of the neck stage of the image recognition model can further increase the feature expression of the indicator light. The edge perception module includes a first convolutional layer, an edge detection operator, a preset fusion module and a second convolutional layer.

[0093] The edge detection operator may be the Sobel operator, which is used to extract edge information from the input first-order feature map to obtain edge information of the first-order feature map. Specifically, the edge detection operator is used to calculate the edge information in the first-order feature map. The gradient amplitude of each pixel point is calculated in at least one preset direction to obtain the edge information of the first-order feature map. For example, the edge information of the first-order feature map can represent the edge map corresponding to the first-order feature map. Specifically, the calculation process of the above-mentioned gradient amplitude can refer to the following formula (2):

[0094] Gradient=Sqrt(Gx* Gx +Gy* Gy) formula (2);

[0095] Gradient represents the gradient magnitude. Edge detection operators are used to extract edge information from the input feature map. Edge detection operators typically include two convolution kernels: horizontal (Gx) and vertical (Gy). Gx and Gy represent the horizontal and vertical gradient components of each pixel in the first-order feature map, respectively. These two convolution kernels are convolved with the input feature map to produce horizontal and vertical gradient maps. * represents multiplication.

[0096] Step S43: Determine the image recognition result of the image to be recognized based on each target feature map.

[0097] The number of initial feature maps, the number of feature maps to be processed, and the number of target feature maps may be equal. The number of image recognition results for the image to be identified is one. Specifically, step S43 may be to determine the image recognition result for the image to be identified from the image recognition results corresponding to the target feature maps. The image recognition result corresponding to each target feature map includes at least one of the following: position information corresponding to the target feature map, and a classification result corresponding to the target feature map.

[0098] Specifically, step S43 may include inputting each target feature map into at least one linear layer to obtain an image recognition result corresponding to each target feature map. An image recognition result that satisfies a preset condition is selected from the image recognition results corresponding to each target feature map as the image recognition result for the image to be recognized. The preset condition may be selecting the classification result and / or position information with the highest confidence among the image recognition results corresponding to the target feature map as the image recognition result for the image to be recognized. In some application scenarios, the preset condition may be selecting the classification result with the highest confidence among the image recognition results corresponding to the target feature map as the classification result for the indicator light in the image recognition result for the image to be recognized. The target feature map corresponding to the classification result with the highest confidence is used as the reference feature map, and the position information corresponding to the reference feature map is used as the position information for the indicator light in the image recognition result for the image to be recognized. In some application scenarios, the preset condition may be selecting the position information with the highest confidence among the image recognition results corresponding to the target feature map as the position information for the indicator light in the image recognition result for the image to be recognized. The target feature map corresponding to the position information with the highest confidence is used as the reference feature map, and the classification result corresponding to the reference feature map is used as the classification result for the indicator light in the image recognition result for the image to be recognized. In other application scenarios, the classification result and position information with the highest confidence in the image recognition result corresponding to the target feature map are respectively used as the classification result of the indicator light and the position information of the indicator light in the image recognition result of the image to be recognized.

[0099] In the case where there are multiple indicator lights in the image to be identified, in some application scenarios, the position information in the image recognition result corresponding to each target feature map can represent the position information of multiple indicator lights in the image to be identified, and the classification result in the image recognition result corresponding to each target feature map can represent the classification results of multiple indicator lights (for example, preset working status, preset attribute information). The position information of each indicator light in the image to be identified in the image recognition result corresponding to the target feature map can represent the detection frame of the indicator light in the image to be identified. Exemplarily, the image recognition results corresponding to each target feature map will be obtained by non-maximum suppression (NMS) filtering, where the NMS filtering will filter out the optimal detection frame by setting a threshold, suppress repeated or unnecessary detection frames, and thus obtain the accurate position and classification results of each indicator light.

[0100] See also Figure 7 , Figure 7 This is a sixth flow chart of an embodiment of the image recognition method of the present application.

[0101] In some embodiments, the above-mentioned image recognition method may further include a step of training the image recognition model, the training step including:

[0102] Step S61: Acquire a sample image and a sample binary image of the sample image.

[0103] The sample image is an image of the same type as the image to be identified, or an image of a different type. The sample image is an image captured of a sample object. The sample object can be the indicator light, or a small indicator light of the same or different type as the indicator light. The specific definition and acquisition method of the sample image can be referred to above for the image to be identified and will not be further described here.

[0104] See also Figure 8a and Figure 8b , Figure 8a is a schematic diagram of a sample image of an embodiment of the image recognition method of the present application, Figure 8b It is a schematic diagram of a sample binary image of an embodiment of the image recognition method of the present application.

[0105] The number of sample objects in the sample image is multiple. For example, the sample objects in the sample image are as follows: Figure 8a a, b, c, d, e, f, g, h and k shown. The sample object can be an indicator light. Different sample objects can represent different indicator lights. The preset attribute information and preset working status of different sample objects can represent the status of different functions corresponding to the detection device. Figure 8b The sample binary image of the sample image shown can be an image with a black background that matches the sample image and a white area where the sample object is located. The area where the sample object is located can represent information such as the location and size of the sample object in the sample image. The sample binary image of the sample image can be obtained by directly obtaining a preset binary image that matches the sample image, or by binarizing each sample image to obtain the sample binary image.

[0106] Step S62: Input each sample target feature map of the sample image into at least one feature point detection head to obtain a predicted binary map output by each feature point detection head.

[0107] The head stage of the image recognition model may include at least one feature point detection head. Step S62 may include inputting each target feature map of the sample image into each feature point detection head, obtaining a predicted binary map output by each feature point detection head. For each target feature map, at least one convolution operation is performed on the target feature map to obtain a corresponding predicted binary map.

[0108] See also Figure 9 , Figure 9 Schematic diagram of a feature point detection head according to an embodiment of the image recognition method of the present application.

[0109] The head stage of the image recognition model includes a detection head corresponding to each target feature map. Each detection head corresponding to each target feature map includes a feature point detection head. Each feature point detection head includes a first preset convolutional layer and a second preset convolutional layer in a cascaded configuration. The convolution kernel sizes in the first and second preset convolutional layers can be different. For example, the convolution kernel size in the first preset convolutional layer is 1×1, while the convolution kernel size in the second preset convolutional layer is 5×5.

[0110] The method for determining each sample target feature map of the sample image can be the same as the method for determining each target feature map of the above-mentioned image to be identified. Please refer to the above for the specific definition method, which will not be repeated here.

[0111] Step S63: Obtaining a target loss based on the difference between each predicted binary image and the sample binary image of the sample image.

[0112] In some application scenarios, the above step S63 may directly use the difference between the predicted binary image corresponding to each target feature map and the sample binary image of the sample image as the target loss. The above step S63 may include the following steps: including: determining the first loss based on the difference between each predicted binary image and the sample binary image of the sample image. Obtaining the sample recognition result corresponding to the sample image, the sample recognition result is obtained through pre-labeling, and the sample recognition result includes the sample position information of the sample object. Inputting each target feature map of the sample image into at least one target frame detection head to obtain the predicted position information of the sample object output by each target frame detection head. Determining the second loss based on the difference between each predicted position information and the sample position information. In some application scenarios, the sample recognition result also includes the sample classification result of the sample object. Inputting each target feature map of the sample image into at least one classification head to obtain the predicted classification result of the sample object output by each classification head. Determining the third loss based on the difference between each predicted classification result and the sample classification result.

[0113] In some application scenarios, the first loss is directly used as the target loss. In some application scenarios, the first loss and the second loss are weighted and fused to obtain the target loss. In some application scenarios, the first loss and the third loss are weighted and fused to obtain the target loss. In some application scenarios, the first loss, the second loss, and the third loss are weighted and fused to obtain the target loss. In some application scenarios, the second loss and the third loss are weighted and fused to obtain the target loss.

[0114] Step S64: Using the target loss, adjust the parameters in the image recognition model to obtain a trained image recognition model.

[0115] It can be considered that when the sample image is sufficient data, the parameters in the image recognition model are trained using the determined target loss until the image recognition model converges to obtain a trained image recognition model. The above steps S11 to S13 are performed using the trained image recognition model. In order to better capture the subtle features of small targets (such as indicator lights) in the head stage of the image recognition module, the model structure of the present application adds a feature point detection head. The specific structure of the feature point detection head is at least two preset convolutional layers. The size of the input target feature map and the output predicted binary map are consistent. The feature point detection head added by the present application can assist in the detection of small targets.

[0116] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0117] See also Figure 10 , Figure 10 Schematic diagram of an image recognition model in one embodiment of the image recognition method of the present application.

[0118] The image recognition model includes a backbone network (backbone), a neck stage (neck) and a head stage (head). Specifically, the backbone network includes several extraction modules. Exemplarily, the several extraction modules include a first extraction module, an intermediate extraction module 1, an intermediate extraction module 2 and a last extraction module. Among them, P1, P2, P3, P4 and P5 respectively included in the first extraction module, the intermediate extraction module 1, the intermediate extraction module 2 and the last extraction module can be set with a convolutional neural network with the same structure, wherein the parameters of the convolutional neural network can be the same or different, which is not limited here. For the attention module in the intermediate extraction module 1 and the intermediate extraction module 2, please refer to the relevant definitions of the attention module in the above-mentioned intermediate extraction modules, which will not be repeated here. Exemplarily, the above step S21 is performed using the first extraction module. The above step S22 is performed in sequence using the intermediate extraction module 1 and the intermediate extraction module 2. The above step S23 is performed using the last extraction module.

[0119] For example, in the neck stage, the image recognition model includes: two upsampling modules, two first target modules (not shown), one downsampling module and two second target modules (not shown). Each first target module includes a cascaded attention module and a fusion module. The attention module in the first target module can be the preset attention module in step S53 above. This application takes the structure of the preset attention module and the attention module in the backbone network as an example, and whether the parameters of the two are the same is not limited. The fusion module in the first target module can be a module for residual connection. Exemplarily, step S41 is performed using each first target module. Specifically, step S51 is performed using the upsampling module. Step S52 is performed using the attention module in each first target module. Step S53 is performed using the fusion module in each first target module. Exemplarily, the downsampling module is used to perform the above-mentioned step of downsampling the first feature map to be processed to obtain the target feature map corresponding to the first feature map to be processed. Each second target module includes a cascaded edge perception module and a fusion module. The fusion module in the second target module can be a module for residual connection. Exemplarily, each second target module is used to perform the above-mentioned step of performing a fusion process on the second feature map to be fused with the non-first feature map to be processed for each non-first feature map to be processed, to obtain a target feature map corresponding to the non-first feature map to be processed. Specifically, the edge perception module in each second target module is used to perform the above-mentioned step of performing a preset convolution process on the second feature map to be fused to obtain a first-order feature map corresponding to the second feature map to be fused, and the above-mentioned step of performing information fusion process based on the first-order feature map and edge information to obtain a second-order feature map corresponding to the second feature map to be fused. The fusion module in each second target module is used to perform the above-mentioned step of performing a preset fusion process on the second-order feature map and the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed.

[0120] In the head stage, the image recognition model includes several detection modules, specifically including: a first detection module, a second detection module and a third detection module. Each detection module includes a target frame detection head, a classification head and a feature point detection head. It can be understood that only the target frame detection head and the classification head in each detection module are used in the application stage of the image recognition model. The target frame detection head, the classification head and the feature point detection head in each detection module need to be used in the training stage of the image recognition model. Among them, the target frame detection head in each detection module is used to determine the position information corresponding to each target feature map. The target frame detection head in each detection module is used to determine the classification result corresponding to each target feature map. The feature point detection head in each detection module is used to determine the predicted binary map corresponding to each sample target feature map during the training stage of the image recognition model.

[0121] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0122] See also Figure 11 , Figure 11 It is a structural diagram of an embodiment of the image recognition device of the present application. The image recognition device 110 includes an acquisition module 111, a feature extraction module 112 and a determination module 113. The acquisition module 111 is used to acquire an image to be recognized related to the indicator light. The feature extraction module 112 is used to perform multi-level feature extraction on the image to be recognized using the image recognition model to obtain an initial feature atlas, wherein the initial feature atlas includes at least two initial feature maps and the scales between the initial feature maps are different, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction. The determination module 113 is used to determine the image recognition result of the image to be recognized based on each initial feature map, and the image recognition result includes at least one of the position information of the indicator light and the classification result of the indicator light.

[0123] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0124] Please refer to the image recognition method for the functions performed by each module, which will not be repeated here.

[0125] See also Figure 12 , Figure 12 1 is a schematic diagram of the structure of an embodiment of an electronic device of the present application. Electronic device 120 includes memory 121 and processor 122. Processor 122 is configured to execute program instructions stored in memory 121 to implement the steps of the above-described image recognition method embodiment. In a specific implementation scenario, electronic device 120 may include, but is not limited to, a multi-camera device, a microcomputer, and a server. Furthermore, electronic device 120 may also include mobile devices such as laptops and tablet computers, which are not limited herein.

[0126] Specifically, the processor 122 is used to control itself and the memory 121 to implement the steps in the above-mentioned image recognition method embodiment. The processor 122 can also be called a CPU (Central Processing Unit). The processor 122 may be an integrated circuit chip with signal processing capabilities. The processor 122 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 122 can be implemented by an integrated circuit chip.

[0127] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0128] See also Figure 13 , Figure 13 The computer-readable storage medium 130 stores program instructions 1301, which, when executed by a processor, implement the steps of any of the above-mentioned image recognition method embodiments.

[0129] The above-mentioned scheme is different from the existing technology which may lead to poor image processing effect due to limited / lost feature information. After obtaining the image to be identified, the present application uses the image recognition model to perform multi-level feature extraction on the image to be identified to obtain an initial feature map set, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction to enhance the feature expression ability of each initial feature map for the indicator light, thereby determining the image recognition result of the image to be identified based on each initial feature map to improve the accuracy of the image recognition result.

[0130] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0131] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0132] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0133] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. An image recognition method, characterized in that: The method comprises: Acquire an image to be identified related to the indicator light; Performing multi-level feature extraction on the image to be recognized using an image recognition model to obtain an initial feature atlas, wherein the initial feature atlas includes at least two initial feature maps with different scales, wherein at least two different feature extraction methods are used in at least part of the hierarchical feature extraction; Based on each of the initial feature maps, an image recognition result of the image to be recognized is determined, where the image recognition result includes at least one of the position information of the indicator light and the classification result of the indicator light.

2. The method according to claim 1, characterized in that The image recognition model includes a plurality of extraction modules arranged in cascade, wherein the plurality of extraction modules include at least one intermediate extraction module, and each intermediate extraction module includes a first module and a second module arranged in cascade, wherein one of the first module and the second module is a preset feature extraction module, and the other is an attention module. The image recognition model is used to perform multi-level feature extraction on the image to be recognized to obtain an initial feature atlas, including: Using a first extraction module to extract a first preset feature from the image to be identified to obtain a starting feature map; Inputting the initial feature map into each of the intermediate extraction modules in sequence to obtain an initial feature map output by each of the intermediate extraction modules; The initial feature map output by the last intermediate extraction module is input into the last extraction module among the plurality of extraction modules to perform second preset feature extraction, so as to obtain the initial feature map output by the last extraction module.

3. The method according to claim 2, characterized in that The first module is the attention module, the second module is the preset feature extraction module, and the initial feature map is sequentially input into each of the intermediate extraction modules to obtain the initial feature map output by each of the intermediate extraction modules, including: For each intermediate extraction module, perform the following steps: Inputting a candidate feature map into the first module of the current intermediate extraction module to obtain an advanced feature map corresponding to the candidate feature map, wherein the candidate feature map includes the starting feature map or the initial feature map output by the previous intermediate extraction module; The advanced feature map is input into the second module in the current intermediate extraction module to perform a third preset feature extraction to obtain an initial feature map corresponding to the current intermediate extraction module.

4. The method according to claim 3, characterized in that The step of inputting the candidate feature map into the first module of the current intermediate extraction module to obtain an advanced feature map corresponding to the candidate feature map includes: Obtaining the similarity between each pixel in the candidate feature map; Based on each of the similarities, determining an attention weight corresponding to the candidate feature map; Combining the candidate feature map with the attention weight, an advanced feature map corresponding to the candidate feature map is obtained.

5. The method according to any one of claims 1 to 4, characterized in that The determining, based on each of the initial feature maps, an image recognition result of the image to be recognized includes: Performing feature fusion on each of the initial feature maps to obtain at least one initial fused feature map; Determining a target feature map corresponding to each of the initial feature maps based on the last initial feature map in the initial feature map set and each of the initial fusion feature maps; Based on each of the target feature maps, an image recognition result of the image to be recognized is determined.

6. The method according to claim 5, characterized in that The performing feature fusion on the initial feature maps to obtain at least one initial fused feature map includes: For each non-last initial feature map in the initial feature map set, perform the following steps: Obtaining a first feature map to be fused corresponding to the current non-last initial feature map, where the first feature map to be fused includes the last initial feature map after upsampling, or a feature map obtained by upsampling the initial fused feature map corresponding to the next non-last initial feature map; Performing a first fusion process on the current non-last initial feature map and the first feature map to be fused to obtain a candidate fused feature map; The candidate fused feature map is input into a preset attention module to obtain the features output by the preset attention module, and the features output by the preset attention module are used as the initial fused feature map corresponding to the current non-last initial feature map.

7. The method according to claim 5, characterized in that The determining, based on the last initial feature map in the initial feature map set and each of the initial fusion feature maps, a target feature map corresponding to each of the initial feature maps, includes: The last initial feature map and each of the initial fused feature maps are respectively used as the feature maps to be processed corresponding to each of the initial feature maps; For a first feature map to be processed, downsampling the first feature map to be processed is performed to obtain a target feature map corresponding to the first feature map to be processed; For each non-first feature map to be processed, the second feature map to be fused is fused with the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed, where the second feature map to be fused includes the target feature map corresponding to the first feature map to be processed or the target feature map corresponding to the previous non-first feature map to be processed.

8. The method according to claim 7, characterized in that For each non-first feature map to be processed, fusing the second feature map to be fused with the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed, including: Performing a preset convolution process on the second feature map to be fused to obtain a first-order feature map corresponding to the second feature map to be fused; Performing information fusion processing based on the first-order feature map and edge information to obtain a second-order feature map corresponding to the second feature map to be fused, wherein the edge information is obtained by extracting edge information from the first-order feature map; The second-order feature map is subjected to a preset fusion process with the non-first feature map to be processed to obtain a target feature map corresponding to the non-first feature map to be processed.

9. The method according to any one of claims 1 to 4, characterized in that The method further comprises a step of training the image recognition model, wherein the training step comprises: Acquire a sample image and a sample binary image of the sample image; Inputting each sample target feature map of the sample image into at least one feature point detection head to obtain a predicted binary image output by each feature point detection head; Obtaining a target loss based on a difference between each of the predicted binary images and the sample binary image of the sample image; The target loss is used to adjust the parameters in the image recognition model to obtain a trained image recognition model.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor calls the program instructions from the memory to execute the method according to any one of claims 1 to 9.

11. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 9.