Image recognition method and device, computer equipment and storage medium
By combining feature extraction and fusion of amplitude and depth maps, and performing attention enhancement processing, the accuracy problem of small target detection in no light or low light environments is solved, and efficient target recognition under low light conditions is achieved.
Patent Information
- Application Number
- CN202511935366.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-12-22
AI Technical Summary
In the absence of light or in low light conditions, existing technologies struggle to accurately detect small targets, especially with reduced signal-to-noise ratios in RGB images and insufficient detection accuracy when aided by infrared imaging.
Feature extraction is performed by combining amplitude and depth maps, feature fusion is performed by channel saliency, and attention enhancement processing is applied to generate enhanced fused features, which are then input into the object detection network for recognition.
It improves the accuracy and reliability of target detection under no light or low light conditions, ensures stable feature representation in key areas, and enhances the ability to identify small targets.
Smart Images

Figure CN121366397A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image recognition, and in particular to an image recognition method and device, a computer device and a storage medium. BACKGROUND
[0002] In industrial production lines, warehouse scenes and night work workshops, it is often necessary to continuously monitor the work area to avoid personnel misentry, foreign object intrusion or equipment operation abnormalities. The commonly used detection method is to collect color images through an RGB camera and perform target recognition. Such a method can obtain relatively stable recognition results in normal lighting environments, but in dark or unlit environments, the signal-to-noise ratio of the RGB image is significantly reduced, the detailed information is difficult to retain, the target outline is blurred, and the detection accuracy and stability of the existing recognition model will be significantly reduced.
[0003] In order to compensate for the shortcomings of visible light imaging in low-illumination environments, existing technical solutions attempt to combine infrared imaging or thermal imaging images with RGB images for joint recognition to enhance the perception of targets. However, such solutions still rely on RGB images as the main source of information, and in the absence of light or extremely weak light, the RGB image itself still cannot provide structured detailed information that can be used for effective recognition, resulting in limited overall recognition results, and for small-sized targets or complex background areas, relying solely on infrared imaging or thermal imaging assistance still cannot guarantee detection accuracy.
[0004] Therefore, how to provide an image recognition method for accurately detecting small-sized targets in the absence of light or in dark light environments has become a technical problem that needs to be solved in the field. SUMMARY
[0005] Therefore, it is necessary to provide an image recognition method, device, computer device and storage medium that can accurately detect small-sized targets in the absence of light or in dark light environments.
[0006] An image recognition method, the method comprising: obtaining an amplitude map to be recognized and a corresponding depth map; performing feature extraction processing based on the amplitude map and the depth map to obtain amplitude features of the amplitude map and depth features of the depth map; performing feature fusion processing on the amplitude features and the depth features based on channel saliency of the amplitude features and the depth features to obtain fusion features; performing attention enhancement processing based on the fusion features to obtain enhanced fusion features; providing the enhanced fusion features to a pre-set target detection network to obtain a corresponding image recognition result.
[0007] Optionally, the obtaining the to-be-identified amplitude map and the corresponding depth map comprises: obtaining an initial amplitude map and an initial depth map collected by a target depth camera; mapping pixel values of each pixel in the initial amplitude map and the initial depth map to a preset numerical interval to obtain a to-be-aligned amplitude map and a to-be-aligned depth map; performing alignment processing based on the to-be-aligned amplitude map and the to-be-aligned depth map to obtain the to-be-identified amplitude map and the corresponding depth map.
[0008] Optionally, the performing feature fusion processing on the amplitude feature and the depth feature based on the channel saliency of the amplitude feature and the depth feature to obtain a fusion feature comprises: generating a first fusion weight of the amplitude feature and a second fusion weight of the depth feature according to the channel saliency; performing weighted fusion on the amplitude feature and the depth feature according to the first fusion weight and the second fusion weight to obtain the fusion feature.
[0009] Optionally, the generating the first fusion weight of the amplitude feature and the second fusion weight of the depth feature according to the channel saliency comprises: calculating a first weight score for representing an importance degree of the amplitude feature and a second weight score for representing an importance degree of the depth feature based on the channel saliency; calculating the first fusion weight based on the first weight score and the second weight score; calculating the second fusion weight based on the first fusion weight.
[0010] Optionally, the performing attention enhancement processing based on the fusion feature to obtain an enhanced fusion feature comprises: performing channel attention processing based on the fusion feature to obtain a first enhanced feature; performing spatial attention processing based on the fusion feature to obtain a second enhanced feature; performing element-by-element multiplication on the first enhanced feature and the second enhanced feature to obtain the enhanced fusion feature.
[0011] Optionally, the image recognition result comprises coordinates of a recognition box and a recognition confidence, and after the providing the enhanced fusion feature to a preset target detection network to obtain a corresponding image recognition result, the method further comprises: calculating an intersection over union of the recognition box and a preset alert area based on the coordinates; When the intersection-over-union ratio is not less than a preset intersection-over-union ratio threshold and the recognition confidence is not less than a preset confidence threshold, it is determined that the alert area has an intrusion.
[0012] Optionally, the method further comprises: obtaining a plurality of image recognition results in a preset time window, the plurality of image recognition results comprising coordinates of recognition boxes of each frame and corresponding recognition confidences; calculating a stability index for representing stability of the recognition results based on changes in the coordinates of the recognition boxes and changes in the recognition confidences in the plurality of image recognition results; dynamically adjusting the intersection-over-union ratio threshold and the recognition confidence threshold based on the stability index, wherein when the stability index satisfies a first preset condition, the intersection-over-union ratio threshold and / or the recognition confidence threshold is reduced, and when the stability index satisfies a second preset condition, the intersection-over-union ratio threshold and / or the recognition confidence threshold is increased.
[0013] An image recognition device, the device comprising: a first obtaining module configured to obtain an amplitude map to be recognized and a corresponding depth map; a feature extraction module configured to perform feature extraction processing based on the amplitude map and the depth map to obtain amplitude features of the amplitude map and depth features of the depth map; a feature fusion module configured to perform feature fusion processing on the amplitude features and the depth features based on channel saliency of the amplitude features and the depth features to obtain fusion features; a feature enhancement module configured to perform attention enhancement processing based on the fusion features to obtain enhanced fusion features; a target detection module configured to provide the enhanced fusion features to a preset target detection network to obtain a corresponding image recognition result.
[0014] A computer device comprising a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, the processor implementing the above image recognition method when executing the computer readable instructions.
[0015] A readable storage medium having computer readable instructions stored thereon, the computer readable instructions being executable by a processor to implement the above image recognition method.
[0016] The image recognition method, device, computer device and storage medium described above obtain an amplitude graph to be recognized and a corresponding depth graph; perform feature extraction processing based on the amplitude graph and the depth graph to obtain amplitude features of the amplitude graph and depth features of the depth graph; perform feature fusion processing on the amplitude features and the depth features based on channel saliency of the amplitude features and the depth features to obtain fusion features; perform attention enhancement processing based on the fusion features to obtain enhanced fusion features; and provide the enhanced fusion features to a preset target detection network to obtain a corresponding image recognition result. By performing feature extraction on the amplitude graph and the depth graph respectively, then performing weighted fusion based on channel saliency, and combining attention enhancement processing, the reflection intensity information from the amplitude graph and the distance structure information from the depth graph can be effectively complementary at the feature level, thereby avoiding the problem of missing details in a dark light condition, making the feature expression provided to the target detection network more stable and highlighting the key region, and enabling higher recognition accuracy and recognition reliability to be obtained in a lightless or weak light scene. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0018] Figure 1 is a flowchart of an image recognition method in an embodiment of the present application; Figure 2 is a structural schematic diagram of an image recognition device in an embodiment of the present application; Figure 3 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0020] In an embodiment, as shown in Figure 1 An image recognition method is provided, including the following steps: 101. An amplitude graph to be recognized and a corresponding depth graph are obtained.
[0021] In the embodiments of the present application, the image recognition method can be applied to a depth camera. The depth camera can acquire the amplitude map and the depth map. When acquiring the amplitude map and the corresponding depth map, the depth camera can output the amplitude map reflecting the reflection intensity and the depth map reflecting the distance between the object and the camera by the infrared emission and measurement mechanism inside the depth camera. In actual application, the amplitude map is usually generated by the depth camera after emitting infrared light according to the intensity of the returned light. The pixel value of the amplitude map can reflect the reflection characteristics of different surfaces in the scene. For example, the pixel value of the object with smooth surface or high reflectivity in the amplitude map is usually large, and the pixel value of the object with strong light absorption or dark material is low. The depth map calculates the distance information of each pixel point based on the time of flight (ToF) principle or the structured light principle, and represents the distance relationship in the form of gray value or 16-bit pixel value. For example, the pixel value of the area close to the camera is small, and the pixel value of the area far away from the camera is large.
[0022] By simultaneously acquiring the two kinds of images, relatively stable imaging information can be obtained in a dark or lightless environment, thereby providing reliable data basis for subsequent feature extraction and recognition.
[0023] 102. Perform feature extraction processing based on the amplitude map and the depth map to obtain the amplitude feature of the amplitude map and the depth feature of the depth map.
[0024] In the embodiments of the present application, the two types of images can be input into different feature extraction channels. By convolution operation, nonlinear activation and downsampling operation, higher level representation information can be gradually extracted from the original pixels. For the amplitude map, the original pixels mainly reflect the reflection intensity distribution of each position in the scene. The feature extraction process can convert the intensity change into edge, texture, local contrast and other more abstract feature representations. For the depth map, the pixel value corresponds to the distance between each object in the scene and the camera. The feature extraction process can gradually convert the simple distance distribution into depth features related to the object contour, shape and spatial structure. For example, at the position close to the edge or step of the device, there will be obvious distance mutation in the depth map. The mutation information can be aggregated into a feature response reflecting the shape boundary by convolution operation.
[0025] In a possible embodiment, a separate convolutional feature extraction network can be set up for the amplitude map and the depth map respectively, for example, the amplitude map is input into a first convolutional branch, and the depth map is input into a second convolutional branch. Each branch can include multiple convolutional layers, batch normalization layers, and nonlinear activation layers, and gradually compress the spatial resolution and increase the number of channels through a step or pooling operation, to extract feature representations of different scales and different semantic levels. After the above processing, the intermediate features output by the first convolutional branch are taken as amplitude features, and the intermediate features output by the second convolutional branch are taken as depth features, which are used for subsequent channel saliency evaluation and feature fusion.
[0026] 103. Based on the channel saliency of the amplitude features and the depth features, the amplitude features and the depth features are subjected to feature fusion processing to obtain fused features.
[0027] In the embodiment of the present application, the responses of the two types of features in each channel dimension can be analyzed first to determine the importance of different channels in the current scene. Since the amplitude features and the depth features come from different sources, they may contain different types of effective information at the same position, for example, the amplitude features respond more strongly in areas with obvious reflection differences, while the depth features are more sensitive to shape profiles and distance mutations. By statistically and metrically analyzing the response values of each channel (i.e., convolutional channel), a set of channel saliency indicators can be obtained to represent the relative importance of the channel.
[0028] After obtaining the channel saliency, the corresponding fusion weight can be generated according to its size, so that the channels with higher importance obtain larger weights in fusion, and the channels with lower importance obtain relatively smaller weights. The fusion method can adopt a weighted summation method, that is, the amplitude features and the depth features are respectively scaled according to the corresponding weights, and then the scaled features are added element by element to obtain the fused feature representation. In this way, the effective information from the two types of features can be selectively retained, while the overall structural information is maintained and the feature components that are more discriminative in the current scene are enhanced.
[0029] In a possible embodiment, the channel saliency can be obtained by calculating the mean, variance, maximum value or other statistical quantities of each channel feature, or a lightweight weight generation network can be used to input the two types of features and output the corresponding channel weights. In this case, the fusion weights in different scenes can be dynamically adjusted according to the changes of the input features, so that the fused features are more adaptive to the current image content, and provide more discriminative expressions for subsequent attention enhancement and target detection.
[0030] 104. Based on the fused features, attention enhancement processing is performed to obtain enhanced fused features.
[0031] In the embodiments of the present application, the effective area related to the target can be further emphasized on the basis of the fused features, and the background noise or unimportant feature components can be suppressed. By applying the attention mechanism to the fused features, the response of the features can be analyzed in the channel dimension and the spatial dimension respectively, so that the model can pay more attention to the part that has greater influence on the identification result. For example, in the channel dimension, by counting the overall response of each channel, the channel that contributes more to the current scene can be highlighted; in the spatial dimension, by weighting the responses of different positions of the feature map, the feature strength near the target area can be enhanced, and the interference information of the background area can be weakened.
[0032] In a possible embodiment, the channel attention processing can be performed on the fused features first. The average value or the maximum value of the fused features in the spatial dimension is aggregated to obtain a description vector reflecting the overall importance of each channel, and then a series of nonlinear transformations are performed to generate a channel attention weight, which is applied to the corresponding channel of the fused features, so as to obtain the first enhanced feature after channel enhancement.
[0033] Subsequently, the spatial attention processing can be performed on the fused features. The spatial feature description map is constructed by aggregation in the channel dimension, and then the convolution operation is performed to generate a spatial attention map representing the importance of each spatial position. The attention map is multiplied with the fused features position by position to obtain the second enhanced feature after spatial enhancement. Finally, the two enhanced features are multiplied element by element, so that the effects of channel enhancement and spatial enhancement are superimposed on each other, thereby obtaining the final enhanced fused features, which have stronger expression ability and discrimination ability in the subsequent target detection stage.
[0034] In another possible embodiment, when the attention enhancement processing is performed based on the fused features, the channel attention and the spatial attention can not be executed separately, but a joint attention mechanism can be used to model the fused features simultaneously. The joint attention mechanism can uniformly depict the multiple dimension relationships of the features by constructing a joint weight matrix of the fused features in the channel dimension and the spatial dimension. For example, the fused features can be input into a lightweight attention encoding network, and a weight map consistent with the size of the fused features can be generated through multiple convolution or point-by-point convolution. The weight map can describe the importance of the features in the channel and the spatial simultaneously. Subsequently, the weight map is multiplied with the fused features element by element, so as to obtain the enhanced fused features.
[0035] In another possible embodiment, the attention enhancement processing can be implemented in a self-attention manner. Specifically, the fused features can be divided into several feature blocks, the similarity between the feature blocks is calculated to obtain a weight reflecting the correlation of each region, and then the feature values are redistributed based on the weight, so that the feature map can pay more attention to the regions with structural correlation. Through the above manner, even in the case of small target or sparse feature distribution, the long-range dependency information related to the target can be enhanced, and the stability of subsequent detection can be improved.
[0036] 105. providing the enhanced fused features to a preset target detection network to obtain a corresponding image recognition result.
[0037] In the embodiments of the present application, the target detection network can adopt a single-stage detection structure, for example, a detection architecture similar to the YOLO style, that is, feature extraction and candidate box prediction are completed in the same network at the same time, without the need for an additional candidate region generation process, to reduce the inference delay and reduce the overall computational complexity. In this structure, a backbone network constructed by depth separable convolution can be used as the feature extraction part, by decomposing the standard convolution into pointwise convolution and pointwise convolution, the parameter amount and computational amount of convolution operation are reduced, and by channel reduction strategy, the number of intermediate feature channels is appropriately reduced in network design, so that the network is more lightweight and suitable for running on resource-limited edge devices.
[0038] On the basis of the above-mentioned backbone network, a feature pyramid structure can be further introduced to fuse intermediate feature maps of different scales to realize multi-scale target detection. Specifically, the feature maps from different levels of the backbone network can be up-sampled, down-sampled and horizontally connected to construct a feature pyramid from top to bottom or from bottom to top, so that high-level semantic information and low-level spatial detail information are combined, so that the detection needs of large targets and small targets are simultaneously considered. For small-sized objects, such as slender tools, small parts or locally protruding structures, the joint participation of multi-scale features can improve the distinguishability of the target in the feature space, thereby improving the detection rate of small targets. The output of the target detection network can include a group of candidate boxes (i.e. recognition boxes) predicted on different scale feature maps, each candidate box being associated with a corresponding target class prediction result and a confidence prediction result.
[0039] To obtain the preset target detection network, the network can be supervised trained in an offline stage. In the training process, the Adam optimizer can be used to update the network parameters, the initial learning rate is set to 0.01, and the cosine annealing strategy is used to dynamically adjust the learning rate, so that the learning rate gradually decays in the training process, thereby having a faster convergence speed in the early training stage and a more stable detail fitting capability in the later training stage. The batch size can be set to 8, and under this configuration, a large number of collected labeled samples are iteratively trained, and the training round number can be set to 100 epochs, so that the network fully learns the appearance features and spatial distribution rules of the target under different working conditions. The nonlinear activation function in the network can use ReLU to introduce nonlinear expression capability to the features in the forward propagation, while keeping the calculation simple and easy to implement on an embedded platform.
[0040] The loss function used in the training process can be defined as the weighted sum of the classification loss and the regression loss, that is, L = L cls + λL reg, where L cls can be a classification loss based on cross-entropy, used to constrain the target class prediction result, L reg can be a regression loss based on CioU (Complete IoU), used to constrain the position deviation, scale difference and overlap between the predicted bounding box and the real labeled box, and λ is a weight coefficient for balancing the influence of classification loss and regression loss. Through the cooperation of the above optimizer, learning rate strategy, batch size, training round number and loss function, the target detection network can obtain better fitting effect on the amplitude map and depth map fusion features under the premise of ensuring the stability of the training, thereby forming a preset target detection network with fixed parameters that can be directly called in the inference stage.
[0041] After completing the network training and obtaining the initial model, the model can be optimized for edge deployment. In one possible embodiment, the trained target detection network can be first pruned to delete channels or convolution kernels that contribute less to the final recognition result, to reduce the model size and computation amount; then, the pruned network can be quantized to convert the network weights and part of the intermediate features from floating-point representation to fixed-point representation with lower bit width, thereby further reducing the storage occupation and computation resource requirement. The pruned and quantized model can be deployed on a Jetson Nano or other embedded computing platform, and still be able to achieve an inference frame rate of not less than 30 frames per second under the condition that the input resolution is 640x480, to meet the real-time detection requirements of industrial sites.
[0042] In the inference stage, after the enhanced fusion features are processed by the preset target detection network, the output image recognition result can include the prediction information of several detection targets. Each detection target can at least include the coordinate information of one recognition box, the corresponding recognition confidence and the recognition category; the coordinates of the recognition box can be represented in the form of left upper corner and right lower corner coordinates, or in the form of center point coordinates and width and height; the recognition confidence is used to represent the confidence degree of the network that the target belongs to a certain category; the recognition category is used to indicate the belonging category of the target in the preset category set. In a possible implementation, the number information and internal sorting information of the detection targets can also be attached to the recognition result of each frame of image, so that the subsequent modules can be screened or further processed according to the confidence degree or the target position. The above image recognition result can be used as the input data of the subsequent alarm area judgment and alarm logic, to realize the automatic recognition and response of the intrusion or abnormal target in the industrial scene.
[0043] In the embodiment of the present application, the amplitude graph to be recognized and the corresponding depth graph are obtained; feature extraction processing is performed based on the amplitude graph and the depth graph to obtain the amplitude features of the amplitude graph and the depth features of the depth graph; feature fusion processing is performed on the amplitude features and the depth features based on the channel saliency of the amplitude features and the depth features to obtain fusion features; attention enhancement processing is performed based on the fusion features to obtain enhanced fusion features; and the enhanced fusion features are provided to a preset target detection network to obtain corresponding image recognition results. By performing feature extraction on the amplitude graph and the depth graph respectively, then performing weighted fusion based on the channel saliency, and combining the attention enhancement processing, the reflection intensity information from the amplitude graph and the distance structure information from the depth graph can be effectively complementary at the feature level, thereby avoiding the problem of missing details in dark light conditions, making the feature expression provided to the target detection network more stable and highlighting the key areas, and enabling higher recognition accuracy and recognition reliability in lightless or low light scenes.
[0044] It can be understood that in the specific embodiments of the present application, data related to amplitude graphs, depth graphs, image recognition results, etc. are involved. When the embodiments in the present application are applied to specific products or technologies, the user's permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0045] Optionally, in the step of obtaining the amplitude graph to be recognized and the corresponding depth graph, the initial amplitude graph and the initial depth graph collected by the target depth camera can also be obtained; the pixel values of each pixel in the initial amplitude graph and the initial depth graph are mapped to a preset numerical interval to obtain the amplitude graph to be aligned and the depth graph to be aligned; alignment processing is performed based on the amplitude graph to be aligned and the depth graph to be aligned to obtain the amplitude graph to be recognized and the corresponding depth graph.
[0046] In the embodiments of the present application, the initial amplitude map and the initial depth map synchronously collected by the depth camera can be obtained first. The depth camera outputs an amplitude map reflecting the infrared return intensity and a depth map reflecting the distance information of the scene during actual collection. The pixel range and value distribution of the two types of images are usually inconsistent, and direct use for subsequent feature extraction can cause a large difference in feature scale between different modalities. Therefore, the pixel value normalization processing can be performed on the two types of initial images first, and each pixel value is mapped to a preset value interval, for example, the interval [0, 1]. Through the normalization processing, the value scale between the amplitude map and the depth map can be kept consistent, avoiding the feature of one image dominating in the subsequent fusion due to the too large scale, thereby improving the stability of feature extraction and fusion.
[0047] After normalization, spatial registration processing can be further performed based on the amplitude map to be aligned and the depth map to be aligned, for compensating the possible angle deviation or geometric distortion of the two types of images during collection. Specifically, an affine transformation matrix H can be used to perform pixel-level alignment of the amplitude map and the depth map, and through rotation, translation, scaling and other operations, the pixel positions of the amplitude map and the depth map on the two-dimensional plane are one-to-one corresponding. The affine matrix can be obtained through a calibration process or a robust estimation method (such as RANSAC), and after applying the transformation, the physical positions corresponding to the amplitude map and the depth map under the same scene can be aligned to the greatest extent. After the above normalization and spatial alignment processing, a double-channel input tensor containing the two types of images can be formed , as the input of the subsequent feature extraction stage, so that the data collected by the depth camera can enter the neural network in a more consistent and standardized form, providing a more reliable input data basis for subsequent feature extraction and recognition.
[0048] Optionally, in the step of performing feature fusion processing on the amplitude feature and the depth feature based on the channel saliency of the amplitude feature and the depth feature to obtain the fusion feature, the first fusion weight of the amplitude feature and the second fusion weight of the depth feature can also be generated according to the channel saliency; and the amplitude feature and the depth feature are weighted and fused according to the first fusion weight and the second fusion weight to obtain the fusion feature.
[0049] In the embodiments of the present application, the response of the amplitude feature and the depth feature in each channel dimension can be taken as the judgment basis, and the channel saliency distribution for measuring the relative contribution degree of each channel in the amplitude feature and the depth feature can be obtained by calculating the statistics (such as mean, variance, entropy or other indicators reflecting the importance of the channel) of each channel. When the channel saliency of a certain channel is higher, it indicates that the channel has stronger discrimination ability in the current scene, and therefore a higher weight can be given to the channel in the fusion process; otherwise, a smaller weight is given.
[0050] In generating the fusion weight, the first fusion weight a for adjusting the amplitude feature and the second fusion weight β for adjusting the depth feature can be respectively calculated according to the channel saliency described above, so that the sum of the two can satisfy a certain constraint relationship, for example, a + β = 1. In a typical implementation, the weight generation module can be adopted to input the amplitude feature and the depth feature , dynamically obtain the values of a and β by converging and nonlinearly transforming the feature channels. The dynamic weight generation mechanism can adaptively adjust the weights according to the actual feature distribution each time a different image is input, so that the fusion process is not dependent on a fixed ratio, but can be automatically adapted with scene changes. At this time, the fusion feature can be calculated in the following manner:
[0051] wherein, is the amplitude feature, is the depth feature, and a and β are the fusion weights dynamically generated above. Through the weighted fusion manner, the reflection intensity information from the amplitude map and the spatial structure information from the depth map can each play their respective advantages in the fusion, so that the finally generated fusion feature is more comprehensive and stable in the representation ability, and especially in the scene where the light is unstable or the image noise is large, the performance of subsequent attention enhancement and target detection can be effectively improved.
[0052] In a possible embodiment, the weight generation module can adopt a lightweight neural network structure, including one or more fully connected layers, channel-by-channel convolutional layers or gating units, to learn the dynamic generation law of the fusion weight from the input features. In addition, Softmax, Sigmoid or other normalization activation methods can also be adopted to make the generated weights satisfy the non-negative constraint and the proportion constraint, so that the fusion process is more stable and reliable. The above-mentioned manners can all be used as alternative implementations of the feature fusion module for realizing the dynamic weighted fusion of the amplitude feature and the depth feature.
[0053] Finally, it should be noted that the above-mentioned channel can be a feature channel for representing the convolution feature map in different semantic dimensions, that is, the channel dimension in the multi-dimensional feature tensor formed after the convolution network processing. Each channel usually corresponds to a certain pattern or response type extracted by the convolution kernel, such as edge texture, local structure change, reflection intensity distribution difference, depth gradient change, etc., so the importance of different channels in different scenes can not be consistent. In the embodiments of the present application, the calculation of the channel saliency is based on the response of the convolution feature map channel, and by distinguishing which channel has a more discriminative response, the weights that the amplitude feature and the depth feature should bear in the fusion are determined.
[0054] Optionally, in the step of generating the first fusion weight of the amplitude feature and the second fusion weight of the depth feature according to the channel saliency, a first weight score for representing the importance of the amplitude feature and a second weight score for representing the importance of the depth feature can be calculated based on the channel saliency; the first fusion weight is calculated based on the first weight score and the second weight score; and the second fusion weight is calculated based on the first fusion weight.
[0055] In the embodiments of the present application, when the first fusion weight of the amplitude feature and the second fusion weight of the depth feature are generated according to the channel saliency, the normalization calculation through the intermediate weight score can be further performed to make the weight distribution between the two features more stable and controllable.
[0056] Specifically, the amplitude feature and the depth feature can be respectively weighted and summed or linearly transformed based on the channel saliency to obtain the first weight score for representing the importance of the amplitude feature and the second weight score for representing the importance of the depth feature.
[0057] For example, the weight parameter for the amplitude feature and the weight parameter for the depth feature can be pre-set, the amplitude feature is multiplied by the weight parameter to obtain the first weight score, and the depth feature is multiplied by the weight parameter to obtain the second weight score, so as to introduce the learnable parameter based on the channel saliency and finely depict the contribution of different channels or different modalities.
[0058] After obtaining the first weight score and the second weight score, the first fusion weight can be calculated in an exponential normalization manner to make the weight distribution between the two modalities have good numerical stability and interpretability. In an implementation manner, the first fusion weight α can be calculated in the following form:
[0059] wherein, and respectively represent the weighted results of the amplitude feature and the depth feature after the corresponding weight parameters are applied, represents the exponential operation. Through the above normalization processing, it can be ensured that the first fusion weight α is always between 0 and 1, and increases accordingly with the increase of the amplitude feature weight score.
[0060] Subsequently, the second fusion weight β can be determined based on the first fusion weight, for example, β = 1 α, so that the sum of the two is a constant, thereby forming a set of complementary fusion weights.
[0061] Through the above manner, the proportion of the amplitude feature and the depth feature in the fusion process can be adaptively adjusted under the joint action of the channel saliency and the learnable parameter, so that the fusion feature can maintain a suitable modal balance in different scenes, and the stability and the discriminability of the overall feature expression are improved.
[0062] Optionally, in the step of performing attention enhancement processing based on the fusion feature to obtain the enhanced fusion feature, channel attention processing can also be performed on the fusion feature to obtain first enhanced features; spatial attention processing can be performed on the fusion feature to obtain second enhanced features; and element-wise multiplication is performed on the first enhanced features and the second enhanced features to obtain the enhanced fusion feature.
[0063] In the embodiment of the application, the channel attention processing can be first performed on the fusion feature, the fusion feature is aggregated in the spatial dimension to obtain a channel description vector reflecting the overall response strength of each channel, and a channel attention weight is generated based on the channel description vector; then, the channel attention weight is applied to each channel corresponding to the fusion feature, the channel with a stronger response and more related to the target is enhanced, and the channel with a weaker response mainly from the background is inhibited, so as to obtain the first enhanced feature which is enhanced in the channel dimension.
[0064] Meanwhile, the spatial attention processing can also be performed on the fusion feature, the feature is aggregated in the channel dimension to obtain a spatial description map for characterizing the importance of each spatial position, and a spatial attention weight map is generated based on the spatial description map; the spatial attention weight map and the fusion feature are multiplied position by position, so that the feature located near the target region is enhanced and the feature located in the background region is relatively weakened, thereby obtaining the second enhanced feature which is enhanced in the spatial dimension.
[0065] Subsequently, the element-wise multiplication operation can be performed on the first enhanced feature and the second enhanced feature to superimpose the effects of channel enhancement and spatial enhancement into the same feature map, thereby obtaining the final enhanced fusion feature.
[0066] When expressed by symbols, the enhanced fusion feature can be expressed as:
[0067] wherein, is the fusion feature, is the feature after the channel attention is applied to the fusion feature, is the feature after the spatial attention is applied to the fusion feature, and the symbol represents the element-wise multiplication of two feature maps. Through the above channel-spatial joint attention mechanism, the response strength of the fusion feature to the target region can be effectively improved without increasing too much computing overhead, and the accuracy and robustness of subsequent target detection are further improved.
[0068] Optionally, the image recognition result includes coordinates of the recognition box and recognition confidence, and after the step of providing the enhanced fusion feature to the preset target detection network to obtain the corresponding image recognition result, the intersection over union between the recognition box and the preset alert area can be calculated based on the coordinates; when the intersection over union is not less than a preset intersection over union threshold and the recognition confidence is not less than a preset confidence threshold, it is determined that the alert area has an intrusion.
[0069] In the embodiments of the present application, after the enhanced fusion feature is input into the preset target detection network and the image recognition result is obtained, whether the preset alert area has an intrusion can be further judged based on the recognition result.
[0070] Specifically, the image recognition result output by the target detection network for each frame of image can include multiple candidate targets, each candidate target at least including coordinate information of a recognition box and recognition confidence, and the recognition result can further include a class identifier for representing a target class. To avoid the same target being repeatedly labeled by multiple similar recognition boxes, a non-maximum suppression process can be first performed on the multiple candidate recognition boxes output by the network, for example, using the NMS algorithm, based on the overlap degree between the candidate boxes and the corresponding recognition confidence, to filter out recognition boxes with high overlap degree and low confidence, and to retain a set of final recognition boxes with relatively independent positions and high confidence.
[0071] One or more virtual alert areas can be configured in advance for the monitoring picture, for representing areas that need to be protected. The virtual alert area can be defined in the form of a polygon mask, that is, a polygon area is determined on the image plane by a plurality of vertex coordinates, and in implementation, whether a position falls into the alert area is judged according to the polygon mask.
[0072] After the non-maximum suppression is completed and the final recognition box is obtained, the intersection over union between the coordinates of the recognition box and the geometric shape of the alert area can be calculated, wherein the intersection over union can be defined as the ratio of the area of the intersection region of the recognition box and the alert area to the area of the union region of the two.
[0073] After the intersection over union is calculated, the intersection over union is compared with a preset intersection over union threshold, and the recognition confidence is compared with a preset confidence threshold, when the intersection over union is not less than the preset intersection over union threshold and the recognition confidence is not less than the preset confidence threshold, it can be determined that the alert area has an intrusion target, and accordingly a sound and light alarm device is triggered, for example, a buzzer is controlled to sound, a warning light is controlled to flash, etc., and the alarm occurrence time, the corresponding recognition result and related image information are recorded into a log for facilitating subsequent tracking and analysis.
[0074] In a possible embodiment, the intrusion judgment can also be made based on the relationship between the center point of the recognition box and the virtual warning area. For example, the position of the center point of the recognition box can be calculated according to the coordinates of the recognition box, and when the center point of the recognition box falls within the warning area defined by the polygon mask, and at the same time the conditions that the intersection-over-union is greater than a preset threshold and the recognition confidence is greater than a preset threshold are met, the sound-light alarm is triggered. In the above manner, the recognition result output by the target detection network can be combined with the geometric relationship of the preset virtual warning area on the basis of suppressing redundant detection boxes, so that the automatic recognition and real-time alarm of the intrusion behavior in the key area of the industrial site are realized.
[0075] Optionally, the method can further acquire a plurality of image recognition results in a preset time window, the plurality of image recognition results including coordinates of recognition boxes of each frame and corresponding recognition confidences; based on changes in the coordinates of the recognition boxes and changes in the recognition confidences in the plurality of image recognition results, a stability index for representing stability of the recognition result is calculated; and based on the stability index, the intersection-over-union threshold and the recognition confidence threshold are dynamically adjusted, wherein when the stability index meets a first preset condition, the intersection-over-union threshold and / or the recognition confidence threshold are reduced, and when the stability index meets a second preset condition, the intersection-over-union threshold and / or the recognition confidence threshold are increased.
[0076] In the embodiments of the present application, the intersection-over-union threshold and the recognition confidence threshold for alarm judgment can be dynamically adjusted in combination with the plurality of image recognition results in the preset time window, so as to adapt to the changes in the environment and the fluctuations in the detection results in the industrial site.
[0077] Specifically, during the system operation, the plurality of image recognition results in the preset time window, such as the recognition box coordinates and the corresponding recognition confidences of the recent several frames, can be continuously acquired, and the plurality of recognition results are analyzed as time series data. In each frame, the recognition box associated with the warning area can be selected, the coordinate information (such as the position, width and height of the center point of the recognition box) and the recognition confidence corresponding to the recognition box are recorded, and the above data are cumulatively stored in the time dimension.
[0078] After obtaining the multiple frames of recognition results within the preset time window, the change of the recognition box coordinates and the change of the recognition confidence can be respectively counted to construct a stability index for representing the stability of the recognition result. For example, for the recognition box coordinates, the offset of the center point of the recognition box between consecutive frames, the amplitude of the change of the width and height of the recognition box, or the variance of the coordinate change can be calculated to measure the smoothness of the target position over time; for the recognition confidence, the mean, variance or fluctuation range of the recognition confidence of each frame can be counted to reflect the stability of the network to the target recognition result. If the recognition box position changes little and the recognition confidence fluctuates little within the preset time window, it can be considered that the recognition result is relatively stable; on the contrary, if the recognition box position jumps obviously between multiple frames or the recognition confidence fluctuates greatly, it can be considered that the recognition result is unstable. Based on the above position change and confidence change, a comprehensive stability index can be constructed, for example, the coordinate change and the confidence fluctuation are combined according to certain weights to quantitatively evaluate the stability of the recognition result.
[0079] After calculating the stability index, the intersection-over-union threshold and the recognition confidence threshold for intrusion judgment can be dynamically adjusted according to the value of the stability index. When the stability index meets the first preset condition, for example, the center point of the recognition box in consecutive multiple frames is less than the preset offset threshold, and the fluctuation range of the recognition confidence is lower than the preset fluctuation threshold, it can be determined that the current recognition result is in a stable state. At this time, in order to improve the sensitivity of the alarm response, the intersection-over-union threshold and / or the recognition confidence threshold can be appropriately reduced, so that the system can trigger the alarm earlier when the target continuously approaches or stays near the warning area, thereby improving the response speed to potential risks. When the stability index meets the second preset condition, for example, the recognition box position fluctuates greatly between multiple frames or the recognition confidence changes dramatically over time, it can be determined that the recognition result has great uncertainty. At this time, in order to reduce the false alarm rate, the intersection-over-union threshold and / or the recognition confidence threshold can be appropriately increased, so that the system triggers the alarm only when the overlap is high enough and the recognition confidence is stable enough, thereby suppressing false alarms caused by noise, reflection interference or transient false detection in complex environments.
[0080] In one possible embodiment, the stability index can be designed in a hierarchical or interval format. For example, the stability index can be divided into three intervals: "high stability," "medium stability," and "low stability." When the stability index falls into the "high stability" interval, a relatively low set of cross-union ratio (CUI) and recognition confidence thresholds is selected; when the stability index falls into the "low stability" interval, a relatively high set of thresholds is selected; and when the stability index is in the intermediate interval, the default thresholds remain unchanged. Through an adaptive threshold adjustment mechanism based on multi-frame recognition results, this invention improves the overall robustness of the system while balancing the timeliness and accuracy of alarms under different operating conditions, reducing false alarms or missed alarms caused by single-frame anomalies, making it more suitable for long-term continuous monitoring scenarios deployed in industrial sites.
[0081] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0082] In one embodiment, an image recognition device is provided, which corresponds one-to-one with the image recognition methods described in the above embodiments. For example... Figure 2 As shown, the image recognition device includes a first acquisition module 201, a feature extraction module 202, a feature fusion module 203, a feature enhancement module 204, and a target detection module 205. Detailed descriptions of each functional module are as follows: The first acquisition module 201 is used to acquire the amplitude map to be identified and the corresponding depth map; The feature extraction module 202 is used to perform feature extraction processing based on the amplitude map and the depth map to obtain the amplitude features of the amplitude map and the depth features of the depth map; The feature fusion module 203 is used to perform feature fusion processing on the amplitude feature and the depth feature based on the channel saliency of the amplitude feature and the depth feature to obtain fused features; Feature enhancement module 204 is used to perform attention enhancement processing based on the fused features to obtain enhanced fused features; The target detection module 205 is used to provide the enhanced fusion features to a preset target detection network to obtain the corresponding image recognition result.
[0083] Optionally, the first acquisition module 201 is further configured to: Obtain the initial amplitude map and initial depth map acquired by the target depth camera; The pixel values of each pixel in the initial amplitude map and the initial depth map are mapped to a preset numerical range to obtain the amplitude map and the depth map to be aligned. align the amplitude map to be aligned with the depth map to be aligned to obtain the amplitude map to be recognized and the corresponding depth map.
[0084] Optionally, the feature fusion module 203 is further configured to: generate a first fusion weight of the amplitude feature and a second fusion weight of the depth feature according to the channel saliency; perform weighted fusion on the amplitude feature and the depth feature according to the first fusion weight and the second fusion weight to obtain the fusion feature.
[0085] Optionally, the feature fusion module 203 is further configured to: calculate a first weight score for representing the importance of the amplitude feature and a second weight score for representing the importance of the depth feature based on the channel saliency; calculate the first fusion weight based on the first weight score and the second weight score; calculate the second fusion weight based on the first fusion weight.
[0086] Optionally, the feature enhancement module 204 is further configured to: perform channel attention processing based on the fusion feature to obtain a first enhanced feature; perform spatial attention processing based on the fusion feature to obtain a second enhanced feature; perform element-wise multiplication on the first enhanced feature and the second enhanced feature to obtain the enhanced fusion feature.
[0087] Optionally, the image recognition result includes coordinates of a recognition box and a recognition confidence, and the apparatus further includes: a first calculation module configured to calculate an intersection over union of the recognition box and a preset alert area based on the coordinates; a first determination module configured to determine that the alert area is invaded when the intersection over union is not less than a preset intersection over union threshold and the recognition confidence is not less than a preset confidence threshold.
[0088] Optionally, the apparatus further includes: a second acquisition module configured to acquire a plurality of image recognition results in a preset time window, the plurality of image recognition results including coordinates of recognition boxes of each frame and corresponding recognition confidences; a second calculation module configured to calculate a stability index for representing stability of the recognition result based on changes in the coordinates of the recognition boxes and changes in the recognition confidences in the plurality of image recognition results. An adjustment module is used to dynamically adjust the cross-union ratio threshold and the identification confidence threshold based on the stability index. When the stability index meets a first preset condition, the cross-union ratio threshold and / or the identification confidence threshold are decreased. When the stability index meets a second preset condition, the cross-union ratio threshold and / or the identification confidence threshold are increased.
[0089] Each module in the aforementioned image recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0090] In one embodiment, a computer device is provided, which may be a terminal device, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium storing computer-readable instructions. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer-readable instructions implement an image recognition method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.
[0091] In this application embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the image recognition method described above.
[0092] In one embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, they implement the steps of the image recognition method described above.
[0093] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through computer readable instructions, and the computer readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer readable instructions are executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0095] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An image recognition method characterized by, The method comprises: obtaining an amplitude map to be identified and a corresponding depth map; based on the amplitude map and the depth map, feature extraction processing is performed to obtain amplitude features of the amplitude map and depth features of the depth map; based on the channel saliency of the amplitude features and the depth features, the amplitude features and the depth features are subjected to feature fusion processing to obtain fusion features; based on the fusion features, attention enhancement processing is performed to obtain enhanced fusion features; the enhanced fusion features are provided to a preset target detection network to obtain a corresponding image recognition result.
2. The image recognition method of claim 1, wherein, The method comprises: obtaining an amplitude map to be identified and a corresponding depth map; obtaining an initial amplitude map and an initial depth map collected by a target depth camera; mapping the pixel values of each pixel in the initial amplitude map and the initial depth map to a preset numerical interval to obtain an amplitude map to be aligned and a depth map to be aligned; 3. The image recognition method of claim 1, wherein, based on the amplitude map to be aligned and the depth map to be aligned, alignment processing is performed to obtain the amplitude map to be identified and the corresponding depth map. The method comprises: based on the channel saliency, first fusion weights of the amplitude features and second fusion weights of the depth features are generated; 4. The image recognition method of claim 3, wherein, the amplitude features and the depth features are weighted and fused according to the first fusion weights and the second fusion weights to obtain the fusion features. The method comprises: based on the channel saliency, first weight scores representing the importance of the amplitude features and second weight scores representing the importance of the depth features are calculated; based on the first weight scores and the second weight scores, the first fusion weights are calculated; 5. The image recognition method of claim 1, wherein, based on the first fusion weights, the second fusion weights are calculated. The method comprises: based on the fusion features, channel attention processing is performed to obtain first enhanced features; based on the fusion features, spatial attention processing is performed to obtain second enhanced features; 6. The image recognition method according to any one of claims 1 to 5, wherein the first enhanced features and the second enhanced features are subjected to element-by-element multiplication to obtain the enhanced fusion features. The image recognition result comprises the coordinates of the recognition box and the recognition confidence. After the enhanced fusion features are provided to the preset target detection network to obtain the corresponding image recognition result, the method further comprises: based on the coordinates, the intersection over union of the recognition box and a preset warning area is calculated; 7. The image recognition method of claim 6, wherein, when the intersection over union is not less than a preset intersection over union threshold and the recognition confidence is not less than a preset confidence threshold, it is determined that the warning area is invaded. The method further comprises: obtaining a plurality of image recognition results within a preset time window, the plurality of image recognition results comprising the coordinates of the recognition box of each frame and the corresponding recognition confidence; Based on the change of the bounding box coordinates and the change of the recognition confidence in the multi-frame image recognition result, a stability index for representing the stability of the recognition result is calculated; Based on the stability index, the IoU threshold and the recognition confidence threshold are dynamically adjusted, wherein when the stability index meets a first preset condition, the IoU threshold and / or the recognition confidence threshold are reduced, and when the stability index meets a second preset condition, the IoU threshold and / or the recognition confidence threshold are increased.
8. An image recognition apparatus characterized by comprising: The device comprises: A first acquisition module is configured to acquire an amplitude map and a corresponding depth map to be recognized. A feature extraction module is configured to perform feature extraction processing based on the amplitude map and the depth map to obtain amplitude features of the amplitude map and depth features of the depth map. A feature fusion module is configured to perform feature fusion processing on the amplitude features and the depth features based on channel saliency of the amplitude features and the depth features to obtain fused features. A feature enhancement module is configured to perform attention enhancement processing based on the fused features to obtain enhanced fused features. A target detection module is configured to provide the enhanced fused features to a preset target detection network to obtain a corresponding image recognition result. 9.A computer device, comprising a memory, a processor, and computer readable instructions stored on the memory and running on the processor, wherein, The processor executes the computer-readable instructions to implement the image recognition method of any one of claims 1-7.
10. A readable storage medium, having stored thereon computer readable instructions, characterized in that, The computer-readable instructions are executed by the processor to implement the image recognition method of any one of claims 1-7.
Citation Information
Patent Citations
Power transmission work vehicle detection method and system based on frame sequence
CN119478607A
Dual-mode target detection algorithm and system based on steel surface defects
CN120219315A
Salient target detection method and system based on deep interactive fusion of three-modal features
CN120374946A
Warning area intrusion alarm method, device and equipment based on computer vision
CN120451730A