Image Processing Method Based on Double-Coupled Deep Neural Network Learning

By combining RetinaNet and ShuffleNet networks, attention mechanism and depth separation convolution are introduced, and the number of people in the construction elevator is optimized, which solves the accuracy of people statistics in a small space and achieves efficient image processing effect.

CN117935134BActive Publication Date: 2025-07-18TAIZHOU JIAOJIANG XUETIAN CRANE MASCH PLANT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310792352.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-07-18
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

It is difficult to accurately count the number of people in the construction lift cage, especially in a small space, people are easily blocked, and existing image tracking methods are prone to lose targets, resulting in insufficient detection accuracy.

Method used

The image processing method based on dual-coupled deep neural network learning is adopted, and the channel and spatial attention mechanism are introduced through the combination of RetinaNet and ShuffleNet networks, and the deep separation convolution and feature map weighting processing are used to optimize feature extraction and classification regression results.

Benefits of technology

It improves the accuracy and speed of the number of people in the construction elevator, and can accurately identify the number of people and other objects in a narrow space, reducing the impact of shading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117935134B_ABST
    Figure CN117935134B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method based on the learning of a dual-coupled deep neural network. First, the internal monitoring image of the elevator collected by the camera is passed through the backbone network of RetinaNet to obtain the first internal monitoring feature map. Then, the first internal monitoring feature map is passed through ShuffleNet to obtain the second internal monitoring feature map. Finally, the second internal monitoring feature map is passed through the sub-network of RetinaNet to obtain the classification result and the regression result. In this way, the detection accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent image processing, and more specifically, to an image processing method based on dual-coupled deep neural network learning. Background Art

[0002] With the advancement of urbanization construction, there are more and more high-rise buildings, and construction hoists are used more and more frequently. In China, the number of people carried by construction hoists is restricted, and it is required that the safety monitoring system realizes the statistics of the number of people in the hoist cage. According to the working characteristics of construction hoists, since the car needs to transport goods or staff on different floors of the building, a simple safety calculation device outside the door is not feasible. Therefore, a feasible solution is to use the method of image recognition.

[0003] The size of the hoist cage of a construction hoist is generally 3.1m * 1.5m * 2.4m, and the space is narrow. In this area, the movement distance of people is short, and people are easily blocked from each other, and the method of image tracking is prone to losing the target.

[0004] Therefore, an optimized image processing solution is expected. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. An embodiment of this application provides an image processing method based on dual-coupled deep neural network learning. First, the internal monitoring image of the hoist collected by the camera is passed through the backbone network of RetinaNet to obtain a first internal monitoring feature map. Then, the first internal monitoring feature map is passed through ShuffleNet to obtain a second internal monitoring feature map. Finally, the second internal monitoring feature map is passed through the subnet of RetinaNet to obtain a classification result and a regression result. In this way, the detection accuracy can be improved.

[0006] According to one aspect of this application, an image processing method based on dual-coupled deep neural network learning is provided, which includes:

[0007] Obtain the internal monitoring image of the hoist collected by the camera;

[0008] Pass the internal monitoring image through the backbone network of RetinaNet to obtain a first internal monitoring feature map;

[0009] Pass the first internal monitoring feature map through ShuffleNet to obtain a second internal monitoring feature map; and

[0010] Pass the second internal monitoring feature map through the subnet of RetinaNet to obtain a classification result and a regression result.

[0011] In the above image processing method based on dual-coupled deep neural network learning, passing the second internal monitoring feature map through the sub-network of the RetinaNet to obtain a classification result and a regression result includes:

[0012] Performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map;

[0013] Passing the third internal monitoring feature map through a channel attention mechanism module to obtain a fourth internal monitoring feature map;

[0014] Fusing the first internal monitoring feature map and the fourth internal monitoring feature map to obtain a fifth internal monitoring feature map; and

[0015] Inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0016] In the above image processing method based on dual-coupled deep neural network learning, passing the second internal monitoring feature map through the sub-network of the RetinaNet to obtain a classification result and a regression result includes:

[0017] Performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map;

[0018] Passing the third internal monitoring feature map through a spatial attention mechanism module to obtain a sixth internal monitoring feature map;

[0019] Fusing the first internal monitoring feature map and the sixth internal monitoring feature map to obtain a seventh internal monitoring feature map; and

[0020] Inputting the seventh internal monitoring feature map into the regression sub-network of the sub-network to obtain the regression result.

[0021] In the above image processing method based on dual-coupled deep neural network learning, passing the first internal monitoring feature map through ShuffleNet to obtain a second internal monitoring feature map includes:

[0022] Performing feature map slicing of the first internal monitoring feature map along the channel dimension to obtain a plurality of first internal monitoring sub-feature maps;

[0023] Performing grouped convolution on the plurality of first internal monitoring sub-feature maps to obtain a plurality of first internal monitoring depth sub-feature maps; and

[0024] Performing feature map aggregation and channel random mixing on the plurality of first internal monitoring depth sub-feature maps to obtain the second internal monitoring feature map.

[0025] In the above image processing method based on dual-coupled deep neural network learning, obtaining a fourth internal monitoring feature map from the third internal monitoring feature map through a channel attention mechanism module includes:

[0026] Performing explicit spatial encoding on the third internal monitoring feature map using the channel attention mechanism module to obtain a third internal monitoring correlation feature map;

[0027] Calculating the global mean of each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain a channel feature vector;

[0028] Inputting the channel feature vector into a Sigmoid activation function to obtain a channel attention weighted feature vector;

[0029] Based on the autocovariance matrix of the channel attention weighted feature vector, correcting the eigenvalues at each position in the channel attention weighted feature vector to obtain an optimized channel attention weighted feature vector; and

[0030] Using the eigenvalues at each position in the optimized channel attention weighted feature vector as weights to weight each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain the fourth internal monitoring feature map.

[0031] In the above image processing method based on dual-coupled deep neural network learning, inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result includes:

[0032] Calculating the position information schema scene attention unbiased estimation factor of the eigenvalue at each position of the fifth internal monitoring feature map;

[0033] Using the position information schema scene attention unbiased estimation factor of the eigenvalue at each position of the fifth internal monitoring feature map as a weight to weight the eigenvalue at each position of the fifth internal monitoring feature map to obtain an optimized fifth internal monitoring feature map; and

[0034] Inputting the optimized fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0035] In the above image processing method based on dual-coupled deep neural network learning, calculating the position information schema scene attention unbiased estimation factor of the eigenvalue at each position of the fifth internal monitoring feature map includes:

[0036] Calculating the position information schema scene attention unbiased estimation factor of the eigenvalue at each position of the fifth internal monitoring feature map using the following factor calculation formula;

[0037] Among them, the factor calculation formula is as follows:

[0038]

[0039] Among them, and respectively represent functions that map two-dimensional real numbers and three-dimensional real numbers to one-dimensional real numbers. W, H, and C are the width, height, and number of channels of the fifth internal monitoring feature map respectively. f i is the eigenvalue at the i-th position of the fifth internal monitoring feature map. (x i , y i , z i ) are the coordinates of the eigenvalue at the i-th position of the fifth internal monitoring feature map. And is the global mean of all eigenvalues of the fifth internal monitoring feature map. log represents the logarithmic function with base 2. w i is the position information schema scene attention unbiased estimation factor of the eigenvalue at the i-th position of the fifth internal monitoring feature map.

[0040] In the above image processing method based on dual-coupled deep neural network learning, passing the third internal monitoring feature map through the spatial attention mechanism module to obtain the sixth internal monitoring feature map includes:

[0041] Using the convolutional encoding part of the spatial attention mechanism module to perform deep convolutional encoding on the third internal monitoring feature map to obtain the third internal monitoring convolutional feature map;

[0042] Inputting the third internal monitoring convolutional feature map into the spatial attention part of the spatial attention mechanism module to obtain the third internal monitoring spatial attention map;

[0043] Passing the third internal monitoring spatial attention map through the Softmax activation function to obtain the third internal monitoring spatial attention feature map; and

[0044] Calculating the element-wise multiplication of the third internal monitoring spatial attention feature map and the third internal monitoring feature map to obtain the sixth internal monitoring feature map.

[0045] Compared with the prior art, the image processing method based on dual-coupled deep neural network learning provided by this application first passes the internal monitoring image of the elevator collected by the camera through the backbone network of RetinaNet to obtain the first internal monitoring feature map. Then, passing the first internal monitoring feature map through ShuffleNet to obtain the second internal monitoring feature map. Finally, passing the second internal monitoring feature map through the subnet of RetinaNet to obtain the classification result and the regression result. In this way, the detection accuracy can be improved. Description of the Drawings

[0046] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. The following accompanying drawings are not deliberately drawn to scale in actual size, and the focus is on showing the gist of the present application.

[0047] Figure 1 It is an application scenario diagram of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0048] Figure 2 It is a flowchart of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0049] Figure 3 It is a schematic diagram of the architecture of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0050] Figure 4 It is a flowchart of sub-step S130 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0051] Figure 5 It is a flowchart of some steps of sub-step S140 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0052] Figure 6 It is a flowchart of sub-step S412 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0053] Figure 7 It is a flowchart of sub-step S414 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0054] Figure 8 It is a flowchart of another part of the steps of sub-step S140 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0055] Figure 9 It is a flowchart of sub-step S422 of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0056] Figure 10 It is a block diagram of an image processing system based on dual-coupled deep neural network learning according to an embodiment of the present application.

[0057] Figure 11 Schematic diagram of a new classification sub-network according to an embodiment of the present application.

[0058] Figure 12 Schematic diagram of a new regression sub-network according to an embodiment of the present application.

[0059] Figure 13 Schematic diagram of the detection principle of an ultrasonic sensor according to an embodiment of the present application. Detailed implementation manners

[0060] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall also fall within the scope of protection of the present application.

[0061] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0062] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0063] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations in the front or below do not necessarily need to be executed precisely in sequence. On the contrary, as needed, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.

[0064] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0065] According to the working characteristics of construction elevators, the number statistics only need to be carried out at the moment when each door is closed. This patented technology adopts static image recognition. Specifically, taking advantage of the fact that all workers on the construction site must wear safety helmets, a target color and the minimum bounding rectangle are used based on static images to construct an image processing solution.

[0066] Accordingly, the number of people in the elevator, the number of people not wearing safety helmets, as well as electric tricycles and handcarts are automatically recognized through artificial intelligence image recognition technology. This study selects RetinaNet as the basic model for research and makes improvements in three aspects: data preprocessing, network structure, and bounding box selection. For the two sub-network structures, a channel and spatial attention mechanism is introduced to model and generate attention weights and perform weighted processing with the feature map to strengthen the features containing important information; combined with the ShuffleNet network, depthwise separable convolutions are used to extract features, and small-sized convolutions are stacked instead of large-sized convolutions to reduce the number of parameters; a new branch is created in the feature pyramid, and k-means is used to obtain data features, optimize the generation strategy of candidate boxes, adjust the pyramid structure, and improve the detection speed while obtaining rich image information, thereby improving the detection accuracy and average detection rate of some targets of the model on the elevator image dataset.

[0067] In many cases, the computing resources of the human brain are limited and cannot process the images at each position of the overloaded visual information simultaneously. Instead, it is processed through the visual attention mechanism, which can be divided into the channel attention mechanism and the spatial attention mechanism. In the study of the visual attention mechanism, the roles of channel and spatial attention enable the network to better distinguish irrelevant features, focus on the regions where objects exist, and assign more weights to their pixel points during training. Therefore, in order to make the algorithm consider pixel point differences during detection, focus on the regions where objects exist, suppress irrelevant information, and extract more accurate image features, this study introduces an attention mechanism based on the network structure characteristics of RetinaNet to propose a new classification sub-network and regression sub-network. The main role of the classification sub-network in RetinaNet is to classify the detected object categories. After introducing the channel attention mechanism into this branch, a new classification sub-network is created. The channel attention mechanism models the dependencies between the various channels of the feature map to improve the representation ability of important features and increase the ability of the RetinaNet network to correctly classify. In the new classification sub-network, the global information of each channel is obtained through pooling operations on the feature maps of each layer, highlighting the information in the effective regions and compressing the feature map in the spatial dimension. Global average pooling provides feedback to each pixel point on the feature map, while global max pooling only provides feedback at the location of the maximum response gradient in the feature map, which can also help with feature screening. Therefore, when compressing the input feature map in the spatial dimension, both global average pooling and global max pooling are introduced, and two one-dimensional vectors are obtained through the pooling operations. After the compression is completed, a multi-layer perceptron is used for activation, and two fully connected layers and ReLU and Sigmoid activation functions are used to adaptively model the correlation degree between the channels. In order to make the network support input images of any size, two 1×1 convolutional layers are used to replace the two fully connected layers, turning the network into a fully convolutional neural network. In addition, this paper introduces a hyperparameter λ here as the dimensionality reduction ratio of the 1×1 convolutional layer. The dimensionality reduction operation is implemented in the first 1×1 convolutional layer, converting the feature map with c channels into a feature map with c / λ channels, and the dimensionality increase operation is implemented in the second 1×1 convolutional layer, converting the feature map with c / λ channels into c channels. Through experimental verification, when λ is 16, a balance can be achieved in terms of accuracy and speed. Finally, the information of the original feature channels is weighted with the weights after adaptive learning modeling to achieve the effects of feature response and feature recalibration. The channel attention mechanism screens and weights the features in the channel dimension, improving its detection performance.

[0068] ShuffleNet is a lightweight deep network. This network structure introduces group convolution to reduce the computational complexity brought by 1×1 convolution. However, the disadvantage of group convolution is that when several group convolutions overlap, it will hinder the information flow between channels within the group, resulting in the lack of generalization ability of the extracted features. To solve this problem, ShuffleNet uses channel random mixing operation to help the information flow between channels.

[0069] In the technical solution of this application, the two network structures of RetinaNet and ShuffleNet are combined. Depthwise separable convolution is used to extract features in the sub-network. The specific operation is to convert the 3×3 convolution layer into a 3×3 group convolution layer and a 1×1 convolution layer. The number of output channels and the number of groups of the group convolution layer are both set to be the same as the number of input channels. Let DK be the size of the convolution kernel, N be the number of convolution kernels, C be the number of channels, and DF be the input size. Then, when the input and output sizes are the same, depthwise separable convolution performs the convolution operation according to different channels, reducing the parameter calculation and making the model size smaller. And this operation maps the depth information and spatial information separately, decouples the two kinds of information, and can make the network training get better results. After depthwise separable convolution, batch normalization is performed and channel random mixing operation is used to ensure the information flow between channels and improve the model generalization ability. In addition, a residual block is added in the sub-network so that information can be transmitted to the deep network. The final obtained classification and regression sub-network structures are as Figure 11 and Figure 12 shown.

[0070] In addition, in the technical solution of this application, a monitoring camera occlusion recognition function can also be set, including: software algorithm occlusion recognition, which diagnoses the occlusion of the video software based on the detected brightness anomaly; specifically, the definition of brightness anomaly, that is, there is a large area of occlusion on the monitored area screen of the video picture due to human or non-human factors; the algorithm idea is that when the image is occluded: (1) the corresponding area will necessarily be blurred and the edge information will decrease; (2) the pixel variance of this area also decreases; based on this, edge detection is performed on the image, and the variance is calculated for the edge map and the original image respectively; in order to improve the diagnostic accuracy, the image can be divided into N*N blocks, and each small block is diagnosed separately. If the conditions are met, it is considered that the frame image is occluded.

[0071] It also includes hardware physical occlusion recognition: using an ultrasonic sensor to detect the distance, detecting the obstacle in front, and judging by the data returned by the ultrasonic wave and the position of the obstacle in front of the camera determined by the infrared sensor. The principle of the ultrasonic sensor to detect the distance is to measure the time difference from sending out the ultrasonic wave to detecting the sent ultrasonic wave again, and calculate the distance of the object according to the speed of sound at the same time; the detection principle is as Figure 13As shown; in addition, since the speed of ultrasonic waves in air is related to temperature and humidity, in more precise measurements, the changes in temperature and humidity and other factors can also be taken into account.

[0072] Specifically, according to the image processing method based on dual-coupled deep neural network learning of the present application, it includes: acquiring an internal monitoring image of the elevator collected by a camera; passing the internal monitoring image through the backbone network of RetinaNet to obtain a first internal monitoring feature map; passing the first internal monitoring feature map through ShuffleNet to obtain a second internal monitoring feature map; and passing the second internal monitoring feature map through the sub-network of RetinaNet to obtain a classification result and a regression result.

[0073] Particularly, in the technical solution of the present application, passing the second internal monitoring feature map through the sub-network of RetinaNet to obtain a classification result and a regression result includes: performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; passing the third internal monitoring feature map through a channel attention mechanism module to obtain a fourth internal monitoring feature map; fusing the first internal monitoring feature map and the fourth internal monitoring feature map to obtain a fifth internal monitoring feature map; and inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0074] Moreover, in the technical solution of the present application, passing the second internal monitoring feature map through the sub-network of RetinaNet to obtain a classification result and a regression result includes: performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; passing the third internal monitoring feature map through a spatial attention mechanism module to obtain a sixth internal monitoring feature map; fusing the first internal monitoring feature map and the sixth internal monitoring feature map to obtain a seventh internal monitoring feature map; and inputting the seventh internal monitoring feature map into the regression sub-network of the sub-network to obtain the regression result.

[0075] Specifically, in the technical solution of the present application, when fusing the first internal monitoring feature map and the fourth internal monitoring feature map to obtain the fifth internal monitoring feature map, the first internal monitoring feature map is transmitted to the high-dimensional feature space of the fourth internal monitoring feature map by using an information transmission mechanism, which is usually implemented by using a residual mechanism to achieve information transmission. However, since the fourth internal monitoring feature map is obtained by performing convolutional encoding on the third internal monitoring feature map based on a channel attention mechanism, that is, the fourth internal monitoring feature map has feature position attributes in different channel dimensions, and the feature values at each position in the first internal monitoring feature map itself also have position attributes. Therefore, when using the residual mechanism to fuse the first internal monitoring feature map and the fourth internal monitoring feature map to obtain the fifth internal monitoring feature map, the feature values at each position of the fifth internal monitoring feature map also have corresponding position attributes.

[0076] However, when classifying the fifth internal monitoring feature map through the classification sub-network of the sub-network, it is necessary to expand the fifth internal monitoring feature map into a feature vector, that is, it involves the aggregation of the feature values of the fifth internal monitoring feature map by position. Therefore, it is desired to improve the expression effect of each feature value of the fifth internal monitoring feature map on the original feature manifold of the fifth internal monitoring feature map when aggregating by position.

[0077] Based on this, the applicant of the present application calculates the unbiased estimation factor of the position information schema scene attention of the feature value at each position of the fifth internal monitoring feature map, which is expressed as:

[0078]

[0079] Where and respectively represent functions that map two-dimensional real numbers and three-dimensional real numbers to one-dimensional real numbers. For example, they are implemented as non-linear activation functions to activate the representation of the weighted sum plus bias. W, H, and C are the width, height, and number of channels of the fifth internal monitoring feature map respectively. (x i , y i , z i ) are the coordinates of each feature value f i of the fifth internal monitoring feature map. For example, any vertex or center of the feature matrix can be used as the coordinate origin, and is the global mean of all feature values of the fifth internal monitoring feature map.

[0080] Here, the position information schema scene attention unbiased estimation factor uses the schema information representation of the relative geometric direction and relative geometric distance of the fused eigenvalue with respect to the high-dimensional spatial position of the overall feature distribution and the information representation of the high-dimensional feature itself for a higher-order feature expression, so as to further aggregate the shape information of the feature manifold when the eigenvalues are aggregated by position with respect to the overall feature distribution, to achieve an unbiased estimation of the scene geometry of the shape distribution of each sub-manifold set based on the feature manifold in the high-dimensional space, and to accurately express the geometric properties of the manifold shape of the feature map. In this way, by weighting the eigenvalues at each position of the fifth internal monitoring feature map with the position information schema scene attention unbiased estimation factor, the expression effect of each eigenvalue of the fifth internal monitoring feature map on the original feature manifold of the fifth internal monitoring feature map when aggregated by position can be improved, thereby improving the accuracy of the classification result obtained by the fifth internal monitoring feature map through the classification sub-network.

[0081] Figure 1 FIG. is an application scenario diagram of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application. As Figure 1 shown, in this application scenario, first, an internal monitoring image of a lift collected by a camera is obtained (for example, Figure 1 D shown in), and then, the internal monitoring image is input into a server deployed with an image processing algorithm based on dual-coupled deep neural network learning (for example, Figure 1 S shown in), wherein the server can use the image processing algorithm based on dual-coupled deep neural network learning to process the internal monitoring image to obtain a classification result and a regression result.

[0082] After introducing the basic principle of the present application, various non-limiting embodiments of the present application will be specifically introduced below with reference to the accompanying drawings.

[0083] Figure 2 FIG. is a flowchart of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application. As Figure 2 shown, the image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application includes the steps of: S110, obtaining an internal monitoring image of a lift collected by a camera; S120, passing the internal monitoring image through the backbone network of RetinaNet to obtain a first internal monitoring feature map; S130, passing the first internal monitoring feature map through ShuffleNet to obtain a second internal monitoring feature map; and S140, passing the second internal monitoring feature map through the sub-network of RetinaNet to obtain a classification result and a regression result.

[0084] Figure 3Schematic diagram of the architecture of an image processing method based on dual-coupled deep neural network learning according to an embodiment of the present application. As Figure 3 shown, in this network architecture, first, an internal monitoring image of the elevator collected by a camera is obtained; then, the internal monitoring image is passed through the backbone network of RetinaNet to obtain a first internal monitoring feature map; then, the first internal monitoring feature map is passed through ShuffleNet to obtain a second internal monitoring feature map; finally, the second internal monitoring feature map is passed through the sub-network of RetinaNet to obtain a classification result and a regression result.

[0085] More specifically, in step S110, an internal monitoring image of the elevator collected by a camera is obtained. By obtaining the internal monitoring image of the elevator collected by the camera, the number of people inside the elevator, the number of people not wearing safety helmets, and electric tricycles and handcarts are automatically identified using artificial intelligence image recognition technology.

[0086] More specifically, in step S120, the internal monitoring image is passed through the backbone network of RetinaNet to obtain a first internal monitoring feature map.

[0087] More specifically, in step S130, the first internal monitoring feature map is passed through ShuffleNet to obtain a second internal monitoring feature map. ShuffleNet is a lightweight deep network that introduces group convolution to reduce the computational complexity brought by 1×1 convolution. However, the disadvantage of group convolution is that when several group convolutions overlap, it will hinder the information flow between channels within the group, resulting in the lack of generalization ability of the extracted features. To solve this problem, ShuffleNet uses a channel random mixing operation to help the information flow between channels.

[0088] Correspondingly, in a specific example, as Figure 4 shown, passing the first internal monitoring feature map through ShuffleNet to obtain a second internal monitoring feature map includes: S311, performing feature map slicing on the first internal monitoring feature map along the channel dimension to obtain a plurality of first internal monitoring sub-feature maps; S312, performing grouped convolution on the plurality of first internal monitoring sub-feature maps to obtain a plurality of first internal monitoring deep sub-feature maps; and S313, performing feature map aggregation and channel random mixing on the plurality of first internal monitoring deep sub-feature maps to obtain the second internal monitoring feature map.

[0089] More specifically, in step S140, the second internal monitoring feature map is passed through the sub-network of RetinaNet to obtain a classification result and a regression result.

[0090] Correspondingly, in a specific example, asFigure 5 As shown, passing the second internal monitoring feature map through the sub-network of the RetinaNet to obtain a classification result and a regression result includes: S411, performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; S412, passing the third internal monitoring feature map through a channel attention mechanism module to obtain a fourth internal monitoring feature map; S413, fusing the first internal monitoring feature map and the fourth internal monitoring feature map to obtain a fifth internal monitoring feature map; and S414, inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result. The main role of the classification sub-network in RetinaNet is to classify the detected object categories. After introducing the channel attention mechanism in this branch, a new classification sub-network is created. The channel attention mechanism models the dependencies between the channels of the feature map to improve the representation ability of important features and increase the ability of the RetinaNet network to correctly classify.

[0091] Correspondingly, in a specific example, as Figure 6 shown, passing the third internal monitoring feature map through a channel attention mechanism module to obtain a fourth internal monitoring feature map includes: S4121, using the channel attention mechanism module to perform explicit spatial encoding on the third internal monitoring feature map to obtain a third internal monitoring correlation feature map; S4122, calculating the global mean of each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain a channel feature vector; S4123, inputting the channel feature vector into the Sigmoid activation function to obtain a channel attention weighted feature vector; S4124, correcting the feature values at each position in the channel attention weighted feature vector based on the autocovariance matrix of the channel attention weighted feature vector to obtain an optimized channel attention weighted feature vector; and S4125, using the feature values at each position in the optimized channel attention weighted feature vector as weights to weight each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain the fourth internal monitoring feature map.

[0092] Correspondingly, in a specific example, as Figure 7As shown, inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result includes: S4141, calculating the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map; S4142, weighting the feature values at each position of the fifth internal monitoring feature map with the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map to obtain an optimized fifth internal monitoring feature map; and S4143, inputting the optimized fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0093] When classifying the fifth internal monitoring feature map through the classification sub-network of the sub-network, it is necessary to expand the fifth internal monitoring feature map into a feature vector, that is, it involves the position aggregation of the feature values of the fifth internal monitoring feature map. Therefore, it is expected to improve the expression effect of each feature value of the fifth internal monitoring feature map on the original feature manifold of the fifth internal monitoring feature map during position aggregation. Based on this, the applicant of the present application calculates the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map.

[0094] Correspondingly, in a specific example, calculating the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map includes: calculating the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map with the following factor calculation formula; where the factor calculation formula is:

[0095]

[0096] Where and respectively represent functions that map two-dimensional real numbers and three-dimensional real numbers to one-dimensional real numbers, W, H, and C are the width, height, and number of channels of the fifth internal monitoring feature map respectively, f i is the feature value at the i-th position of the fifth internal monitoring feature map, (x i , y i , z i ) are the coordinates of the feature value at the i-th position of the fifth internal monitoring feature map, and is the global mean of all feature values of the fifth internal monitoring feature map, log represents the logarithmic function with base 2, and w i is the position information schema scene attention unbiased estimation factor of the feature value at the i-th position of the fifth internal monitoring feature map.

[0097] Here, the unbiased estimation factor of the position information schema scene attention uses the schema information representation of the relative geometric direction and relative geometric distance of the fused eigenvalue with respect to the high-dimensional space position of the overall feature distribution and the higher-order feature expression of the information representation of the high-dimensional feature itself to further aggregate the shape information of the feature manifold when the eigenvalues are aggregated by position with respect to the overall feature distribution, so as to achieve an unbiased estimation of the scene geometry of the shape distribution of each sub-manifold set based on the feature manifold in the high-dimensional space, and accurately express the geometric properties of the manifold shape of the feature map. In this way, by weighting the eigenvalues at each position of the fifth internal monitoring feature map with the unbiased estimation factor of the position information schema scene attention, the expression effect of each eigenvalue of the fifth internal monitoring feature map on the original feature manifold of the fifth internal monitoring feature map during position aggregation can be improved, thereby improving the accuracy of the classification result obtained by the fifth internal monitoring feature map through the classification sub-network.

[0098] Correspondingly, in a specific example, as Figure 8 shown, obtaining the classification result and the regression result by passing the second internal monitoring feature map through the sub-network of the RetinaNet further includes: S421, performing depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; S422, passing the third internal monitoring feature map through the spatial attention mechanism module to obtain a sixth internal monitoring feature map; S423, fusing the first internal monitoring feature map and the sixth internal monitoring feature map to obtain a seventh internal monitoring feature map; and S424, inputting the seventh internal monitoring feature map into the regression sub-network of the sub-network to obtain the regression result.

[0099] Correspondingly, in a specific example, as Figure 9 shown, passing the third internal monitoring feature map through the spatial attention mechanism module to obtain a sixth internal monitoring feature map includes: S4221, performing depth convolution encoding on the third internal monitoring feature map using the convolution encoding part of the spatial attention mechanism module to obtain a third internal monitoring convolution feature map; S4222, inputting the third internal monitoring convolution feature map into the spatial attention part of the spatial attention mechanism module to obtain a third internal monitoring spatial attention map; S4223, passing the third internal monitoring spatial attention map through the Softmax activation function to obtain a third internal monitoring spatial attention feature map; and S4224, calculating the element-wise product of the third internal monitoring spatial attention feature map and the third internal monitoring feature map by position to obtain the sixth internal monitoring feature map.

[0100] In summary, for the image processing method based on dual-coupled deep neural network learning according to the embodiments of the present application, first, the internal monitoring image of the elevator collected by the camera is passed through the backbone network of RetinaNet to obtain the first internal monitoring feature map. Then, the first internal monitoring feature map is passed through ShuffleNet to obtain the second internal monitoring feature map. Finally, the second internal monitoring feature map is passed through the sub-network of RetinaNet to obtain the classification result and the regression result. In this way, the detection accuracy can be improved.

[0101] Figure 10 FIG. is a block diagram of an image processing system 100 based on dual-coupled deep neural network learning according to an embodiment of the present application. As Figure 10 shown, the image processing system 100 based on dual-coupled deep neural network learning according to an embodiment of the present application includes: a monitoring image acquisition module 110 for acquiring the internal monitoring image of the elevator collected by the camera; a first encoding module 120 for passing the internal monitoring image through the backbone network of RetinaNet to obtain the first internal monitoring feature map; a second encoding module 130 for passing the first internal monitoring feature map through ShuffleNet to obtain the second internal monitoring feature map; and a classification and regression module 140 for passing the second internal monitoring feature map through the sub-network of RetinaNet to obtain the classification result and the regression result.

[0102] In one example, in the above image processing system 100 based on dual-coupled deep neural network learning, the classification and regression module 140 is configured to: perform depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; pass the third internal monitoring feature map through a channel attention mechanism module to obtain a fourth internal monitoring feature map; fuse the first internal monitoring feature map and the fourth internal monitoring feature map to obtain a fifth internal monitoring feature map; and input the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0103] In one example, in the above image processing system 100 based on dual-coupled deep neural network learning, the classification and regression module 140 is further configured to: perform depthwise separable convolution processing on the second internal monitoring feature map to obtain a third internal monitoring feature map; pass the third internal monitoring feature map through a spatial attention mechanism module to obtain a sixth internal monitoring feature map; fuse the first internal monitoring feature map and the sixth internal monitoring feature map to obtain a seventh internal monitoring feature map; and input the seventh internal monitoring feature map into the regression sub-network of the sub-network to obtain the regression result.

[0104] In one example, in the above-mentioned image processing system 100 based on dual-coupled deep neural network learning, the second encoding module 130 is configured to: perform feature map slicing on the first internal monitoring feature map along the channel dimension to obtain a plurality of first internal monitoring sub-feature maps; perform grouped convolution on the plurality of first internal monitoring sub-feature maps to obtain a plurality of first internal monitoring deep sub-feature maps; and perform feature map aggregation and channel random mixing on the plurality of first internal monitoring deep sub-feature maps to obtain the second internal monitoring feature map.

[0105] In one example, in the above-mentioned image processing system 100 based on dual-coupled deep neural network learning, obtaining the fourth internal monitoring feature map by passing the third internal monitoring feature map through the channel attention mechanism module includes: using the channel attention mechanism module to perform explicit spatial encoding on the third internal monitoring feature map to obtain a third internal monitoring correlation feature map; calculating the global mean of each feature matrix of the third internal monitoring correlation feature map along the channel dimension to obtain a channel feature vector; inputting the channel feature vector into the Sigmoid activation function to obtain a channel attention weighted feature vector; correcting the feature values at each position in the channel attention weighted feature vector based on the autocovariance matrix of the channel attention weighted feature vector to obtain an optimized channel attention weighted feature vector; and weighting each feature matrix of the third internal monitoring correlation feature map along the channel dimension with the feature values at each position in the optimized channel attention weighted feature vector as weights to obtain the fourth internal monitoring feature map.

[0106] In one example, in the above-mentioned image processing system 100 based on dual-coupled deep neural network learning, inputting the fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result includes: calculating the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map; weighting the feature value at each position of the fifth internal monitoring feature map with the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map as weights to obtain an optimized fifth internal monitoring feature map; and inputting the optimized fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

[0107] In one example, in the above-mentioned image processing system 100 based on dual-coupled deep neural network learning, calculating the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map includes: calculating the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map using the following factor calculation formula; wherein, the factor calculation formula is:

[0108]

[0109] Among them, and respectively represent functions that map two-dimensional real numbers and three-dimensional real numbers to one-dimensional real numbers. W, H, and C are respectively the width, height, and number of channels of the fifth internal monitoring feature map, and f i is the eigenvalue at the i-th position of the fifth internal monitoring feature map, (x i , y i , z i ) are the coordinates of the eigenvalue at the i-th position of the fifth internal monitoring feature map, and is the global mean of all eigenvalues of the fifth internal monitoring feature map. log represents the logarithmic function with base 2, and w i is the position information schema scene attention unbiased estimation factor of the eigenvalue at the i-th position of the fifth internal monitoring feature map.

[0110] In one example, in the above image processing system 100 based on dual-coupled deep neural network learning, passing the third internal monitoring feature map through the spatial attention mechanism module to obtain a sixth internal monitoring feature map includes: using the convolutional encoding part of the spatial attention mechanism module to perform depth convolutional encoding on the third internal monitoring feature map to obtain a third internal monitoring convolutional feature map; inputting the third internal monitoring convolutional feature map into the spatial attention part of the spatial attention mechanism module to obtain a third internal monitoring spatial attention map; passing the third internal monitoring spatial attention map through the Softmax activation function to obtain a third internal monitoring spatial attention feature map; and calculating the element-wise multiplication of the third internal monitoring spatial attention feature map and the third internal monitoring feature map to obtain the sixth internal monitoring feature map.

[0111] Here, those skilled in the art can understand that the specific functions and operations of each module in the above image processing system 100 based on dual-coupled deep neural network learning have been introduced in detail in the description of the image processing method based on dual-coupled deep neural network learning above with reference to Figures 1 to 9 , and therefore, its repeated description will be omitted.

[0112] As described above, the image processing system 100 based on dual-coupled deep neural network learning according to an embodiment of the present application can be implemented in various wireless terminals, such as a server having an image processing algorithm based on dual-coupled deep neural network learning. In one example, the image processing system 100 based on dual-coupled deep neural network learning according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the image processing system 100 based on dual-coupled deep neural network learning can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the image processing system 100 based on dual-coupled deep neural network learning can also be one of many hardware modules of the wireless terminal.

[0113] Alternatively, in another example, the image processing system 100 based on dual-coupled deep neural network learning and the wireless terminal can also be separate devices, and the image processing system 100 based on dual-coupled deep neural network learning can be connected to the wireless terminal through a wired and / or wireless network, and transmit and interact information according to a predefined data format.

[0114] According to another aspect of the present application, a non-volatile computer-readable storage medium is further provided, on which computer-readable instructions are stored, and when the instructions are executed by a computer, the method described above can be executed.

[0115] The program part in the technology can be considered as a "product" or "article" in the form of executable code and / or related data, which is participated in or implemented by a computer-readable medium. Tangible, permanent storage media can include any memory or storage used by a computer, a processor, or similar devices or related modules. For example, various semiconductor memories, tape drives, disk drives, or any similar devices capable of providing storage functions for software.

[0116] All software, or portions thereof, may sometimes communicate over a network, such as the Internet or other communication networks. Such communication can load software from one computer device or processor to another. For example, from a server or host computer of a video object detection device to a hardware platform of a computer environment, or other computer environments implementing the system, or systems with similar functions related to providing information required for object detection. Therefore, another medium capable of transmitting software elements can also be used as a physical connection between local devices, such as light waves, radio waves, electromagnetic waves, etc., which are propagated through cables, optical fibers, or air. Physical media used to carry the wave, such as cables, wireless connections, or optical fibers and similar devices, can also be considered as media carrying software. As used herein, unless restricted to tangible "storage" media, other terms representing "computer or machine-readable media" refer to media involved in the process of a processor executing any instructions.

[0117] This application uses specific terms to describe embodiments of this application. Such as "the first / second embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.

[0118] In addition, those skilled in the art can understand that various aspects of this application can be illustrated and described by several patentable types or situations, including any new and useful processes, machines, products, or combinations of substances, or any new and useful improvements to them. Accordingly, various aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components", or "systems". In addition, various aspects of this application may be embodied as a computer product located in one or more computer-readable media, which includes computer-readable program code.

[0119] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless explicitly defined as such herein.

[0120] The foregoing is a description of the present invention and should not be construed as limiting thereof. Although several exemplary embodiments of the present invention have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the present invention. Accordingly, all such modifications are intended to be included within the scope of the present invention as defined by the claims. It should be understood that the foregoing is a description of the present invention and should not be considered limited to the specific embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The present invention is defined by the claims and their equivalents.

Claims

1. An image processing method based on learning with a double-coupled deep neural network, characterized in that, Including: Obtain the internal monitoring image of the elevator collected by the camera; Pass the internal monitoring image through the backbone network of RetinaNet to obtain the first internal monitoring feature map; Pass the first internal monitoring feature map through ShuffleNet to obtain the second internal monitoring feature map; Pass the second internal monitoring feature map through the sub-network of RetinaNet to obtain the classification result and the regression result, including: Perform depthwise separable convolution processing on the second internal monitoring feature map to obtain the third internal monitoring feature map; Pass the third internal monitoring feature map through the channel attention mechanism module to obtain the fourth internal monitoring feature map, specifically: Use the channel attention mechanism module to perform explicit spatial encoding on the third internal monitoring feature map to obtain the third internal monitoring correlation feature map; Calculate the global mean of each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain the channel feature vector; Input the channel feature vector into the Sigmoid activation function to obtain the channel attention weighted feature vector; Based on the autocovariance matrix of the channel attention weighted feature vector, correct the feature values at each position in the channel attention weighted feature vector to obtain the optimized channel attention weighted feature vector; Use the feature values at each position in the optimized channel attention weighted feature vector as weights to weight each feature matrix along the channel dimension of the third internal monitoring correlation feature map to obtain the fourth internal monitoring feature map; Fuse the first internal monitoring feature map and the fourth internal monitoring feature map to obtain the fifth internal monitoring feature map; Calculate the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map; Use the position information schema scene attention unbiased estimation factor of the feature value at each position of the fifth internal monitoring feature map as a weight to weight the feature value at each position of the fifth internal monitoring feature map to obtain the optimized fifth internal monitoring feature map; Input the optimized fifth internal monitoring feature map into the classification sub-network of the sub-network to obtain the classification result.

2. The image processing method based on dual-coupled deep neural network learning according to claim 1, wherein Pass the second internal monitoring feature map through the sub-network of RetinaNet to obtain the classification result and the regression result, including: Perform depthwise separable convolution processing on the second internal monitoring feature map to obtain the third internal monitoring feature map; Pass the third internal monitoring feature map through the spatial attention mechanism module to obtain the sixth internal monitoring feature map; Fuse the first internal monitoring feature map and the sixth internal monitoring feature map to obtain the seventh internal monitoring feature map; Input the seventh internal monitoring feature map into the regression sub-network of the sub-network to obtain the regression result.

3. The image processing method based on dual-coupled deep neural network learning according to claim 2, wherein Pass the first internal monitoring feature map through ShuffleNet to obtain the second internal monitoring feature map, including: Perform feature map splitting along the channel dimension on the first internal monitoring feature map to obtain multiple first internal monitoring sub-feature maps; Perform grouped convolution on the multiple first internal monitoring sub-feature maps to obtain multiple first internal monitoring depth sub-feature maps; Perform feature map aggregation and channel random mixing on the multiple first internal monitoring depth sub-feature maps to obtain the second internal monitoring feature map.

4. The image processing method based on dual-coupled deep neural network learning according to claim 3, characterized in that Calculate the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map, including: Calculate the position information schema scene attention unbiased estimation factor of the feature values at each position of the fifth internal monitoring feature map using the following factor calculation formula; wherein, the factor calculation formula is: Among them, and respectively represent functions that map two-dimensional real numbers and three-dimensional real numbers to one-dimensional real numbers. , and are the width, height, and number of channels of the fifth internal monitoring feature map respectively. is the feature value at the th position of the fifth internal monitoring feature map. is the coordinate of the feature value at the th position of the fifth internal monitoring feature map, and is the global mean of all feature values of the fifth internal monitoring feature map. represents the logarithmic function with base 2. is the position information schema scene attention unbiased estimation factor of the feature value at the th position of the fifth internal monitoring feature map.

5. The image processing method based on dual-coupled deep neural network learning according to claim 4, characterized in that Pass the third internal monitoring feature map through the spatial attention mechanism module to obtain the sixth internal monitoring feature map, including: Use the convolutional encoding part of the spatial attention mechanism module to perform depth convolutional encoding on the third internal monitoring feature map to obtain the third internal monitoring convolutional feature map; Input the third internal monitoring convolutional feature map into the spatial attention part of the spatial attention mechanism module to obtain the third internal monitoring spatial attention map; Pass the third internal monitoring spatial attention map through the Softmax activation function to obtain the third internal monitoring spatial attention feature map; Calculate the element-wise multiplication of the third internal monitoring spatial attention feature map and the third internal monitoring feature map to obtain the sixth internal monitoring feature map.

Citation Information

Patent Citations

  • Anti-occlusion pedestrian detection method based on attention mechanism

    CN110929578A

  • Realization method of lightweight convolutional neural network target detection

    CN115311467A