Multi-Stage Deep Network Indoor Scene Recognition Method Based on Multi-Head Attention Mechanism
Through the combination of multi-head attention mechanism and multi-stage deep network, local and deep features of indoor scene images are effectively extracted, solving the problem of easy loss of feature information in traditional methods and improving the recognition accuracy.
Patent Information
- Application Number
- CN202211017228.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-23
AI Technical Summary
In traditional indoor scene recognition methods, local features and depth feature information of indoor scene images are easily lost, resulting in low recognition accuracy.
A multi-stage deep network method based on the multi-head attention mechanism is adopted, and deep feature extraction is performed through Trivial augmentation data augmentation, multi-layer convolutional layer and multi-head self-attention mechanism, combining the pooling layer and the fully connected layer to form the final classifier.
The recognition accuracy of indoor scene images is improved, compared with the traditional method, it is improved by 4.3% on the IndoorCVPR_09 dataset and 3.2% on the Scene15 dataset.
Smart Images

Figure CN115424123B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of recognition and detection, which is the basis for a robot to perceive the environment, and particularly relates to a multi-stage deep network indoor scene recognition method based on a multi-head attention mechanism. Background Technique
[0002] Scene recognition is a core research field of artificial intelligence, mainly studying the classification of scenes in images by using the feature information of scene images. Scene recognition is widely applied in human-computer interaction, so that a machine can recognize the scene information in an image. Therefore, scene recognition plays an important role in fields such as image retrieval, intelligent robots, and intelligent security.
[0003] Compared with outdoor scenes, the image content of indoor scenes containing multiple objects is more complex, and there is occlusion between objects, so it is difficult to extract scene features. Early indoor scene recognition mainly used middle-level features and high-level semantic features, and the recognition effect depended on the selected features, and could not effectively eliminate the interference of indoor objects and accurately extract indoor scene features. Deep learning algorithms have achieved great achievements in many aspects, so more and more scholars have begun to study deep learning algorithms to solve the indoor scene classification problem.
[0004] Traditional deep learning-based scene recognition methods are mainly divided into the following three categories: scene recognition combining deep learning with a bag of visual words, scene recognition based on salient parts, and scene recognition based on multi-layer feature fusion.
[0005] In the scene recognition combining deep learning and a bag of visual words, the bag of words model is based on the idea of text processing, regards an image as a set of unordered visual words, extracts and clusters the features of image patches obtained from the image, and constructs a visual codebook to represent the image. It is simple and easy to use and has achieved good results in some studies, but it is necessary to construct a codebook for a specific task and does not fully utilize the depth feature information of indoor scene images.
[0006] The scene recognition method based on salient parts can be attributed to the fact that the human eye usually can only judge the category of a scene according to the most representative part in the image. Some studies have found that a CNN used for scene recognition can locate the targets in the image that can provide useful information, but a complex scene may contain more than one salient target, and the salient targets in different scenes may have a certain overlap, which has a certain impact on the scene recognition accuracy.
[0007] For scene recognition based on multi-layer feature fusion, each layer structure of the CNN model can learn different features. The deeper the layer, the more abstract and distinctive the learned features are. Use a pre-trained CNN model to extract scene image features, and connect the outputs of the last two fully connected layers as the image representation. This method focuses on using the abstract fully connected layer features to represent the image, while ignoring the rich local information in the convolutional layer, resulting in the underutilization of local feature information in the image information and reducing the recognition accuracy of indoor scenes.
[0008] Therefore, in traditional deep learning-based scene recognition methods, local features and depth feature information of indoor scene images are easily lost and not fully utilized, resulting in relatively low recognition accuracy. Summary of the Invention
[0009] The main technical problem to be solved by the present invention lies in: the problem that local features and depth feature information of indoor scene images in traditional indoor scene recognition methods are easily lost and not fully utilized, resulting in relatively low recognition accuracy. To solve this problem, the present invention proposes a multi-stage deep network indoor scene recognition method based on a multi-head attention mechanism. The input is the original image in the database and a data augmentation function. After passing through the Trivial augmentation data augmentation module, the enhanced indoor scene image is obtained. Local feature extraction is performed through three convolutional layers with different strides, while retaining more feature information. A multi-stage training method is adopted, and deep feature extraction is carried out by sequentially stacking deep convolution and multi-head attention mechanisms. After passing through the final pooling layer and fully connected layer, the finally trained classifier is obtained for the final recognition and detection of indoor scenes.
[0010] According to the first aspect of the present invention, a multi-stage deep network indoor scene recognition method based on a multi-head attention mechanism includes the following steps:
[0011] S1: Obtain a dataset of indoor scenes and divide the dataset into a training set and a test set according to a ratio;
[0012] S2: Preprocess and perform data augmentation on the indoor scene images in the training set to obtain enhanced images;
[0013] S3: Input the enhanced images into a 3-layer convolutional layer network with different strides for downsampling to reduce the size of the images, while retaining more feature information and local information;
[0014] S4: Input the feature information extracted in S3 into the backbone network, and use deep convolution and multi-head self-attention mechanisms to perform deep feature extraction in a multi-stage training manner to obtain deep feature information;
[0015] S5: Input the depth feature information into a pooling layer, a fully connected layer, and a classifier in sequence to obtain the final weights and the trained classifier.
[0016] S6: Use the trained classifier and the final weights to test the indoor scene images in the test set, thereby determining the indoor scene category.
[0017] Preferably, in step S1, the step of obtaining the dataset of indoor scenes includes:
[0018] Use a sentiment robot to collect scene image data of common indoor interaction environments and merge it with the IndoorCVPR_09 related dataset to make a dataset of indoor scenes.
[0019] Preferably, in step S2, use the Trivial augmentation method to perform data augmentation on the preprocessed indoor scene images, specifically including:
[0020] Add a set A of data augmentation functions as input. The data augmentation functions in set A include rotation, translation, flipping, equalization, pixel value flipping, and brightness. Each data augmentation function has its corresponding augmentation range {0, 1, 2…, N}.
[0021] Randomly sample a data augmentation function from A and uniformly sample a value from the augmentation range {0, 1, 2…, N} as the intensity m, where N represents any positive integer. Perform data augmentation on the input image according to the intensity m and return the augmented image.
[0022] Preferably, in step S3, the step of inputting the augmented image into a 3-layer convolutional layer network with different strides for downsampling includes:
[0023] Take the augmented image as the input image and input it into a 3-layer convolutional layer network with different strides.
[0024] The first convolutional layer uses a 3x3 convolution with a stride of 2 and an output channel of 32 to perform a downsampling operation on the input image, reducing the size of the input image and retaining more feature information.
[0025] Take the output of the previous convolutional layer as the input and use two 3x3 convolutions with a stride of 1 to obtain better local information.
[0026] Preferably, in step S4, the backbone network is divided into four stages to generate feature maps of different scales. To produce a hierarchical representation, add a 2x2 convolution layer with a stride of 2 before each stage to reduce the size of the intermediate features and project them to a larger dimension.
[0027] In each stage, there is also a depthwise convolution layer DW before the multi-head attention mechanism, which is used for local feature extraction and reduces the computational complexity at the same time.
[0028] Preferably, in step S5, it ends with a global average pooling layer, a fully connected layer, and a 1000-way classification layer with softmax to obtain the final weights and the trained classifier.
[0029] According to the second aspect of the present invention, a multi-stage deep network indoor scene recognition device based on a multi-head attention mechanism includes the following modules:
[0030] A dataset acquisition module for acquiring a dataset of indoor scenes and dividing the dataset into a training set and a test set according to a ratio;
[0031] A data augmentation module for preprocessing and data augmentation of indoor scene images in the training set to obtain augmented images;
[0032] A downsampling module for inputting the augmented images into a convolutional layer network with 3 different strides for downsampling to reduce the size of the images while retaining more feature information and local information;
[0033] A deep feature extraction module for inputting the feature information extracted by the downsampling module into the backbone network, and using depthwise convolution and multi-head self-attention mechanism to perform deep feature extraction in a multi-stage training manner to obtain deep feature information;
[0034] A classifier acquisition module for sequentially inputting the deep feature information into a pooling layer, a fully connected layer, and a classifier to obtain the final weights and the trained classifier;
[0035] A scene recognition module for using the trained classifier and the final weights to test the indoor scene images in the test set to determine the indoor scene category.
[0036] According to the third aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the multi-stage deep network indoor scene recognition method as described above.
[0037] According to the fourth aspect of the present invention, a storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-stage deep network indoor scene recognition method as described above.
[0038] The technical solution provided by the present invention has the following beneficial effects:
[0039] The present invention proposes a multi-stage deep network indoor scene recognition method based on a multi-head attention mechanism. The input is the original image of the database and the data enhancement function. After the trivial augmentation data enhancement, the enhanced indoor scene image is obtained. After three layers of convolutional layers with different step lengths, local features are extracted while retaining more feature information. A multi-stage training method is adopted, and deep convolution and multi-head attention mechanisms are sequentially superimposed to extract deep features. After the final pooling layer and fully connected layer, the final trained classifier is obtained to perform the final indoor scene recognition and detection. The deep neural network is used to carry out the experiment of scene recognition of the interactive environment of the emotional robot, and the experimental results are analyzed and verified. From the experimental results, the combination of deep convolutional neural network and multi-head attention mechanism effectively improves the recognition accuracy of indoor scene images. The recognition accuracy of the method proposed in the present invention is improved by 4.3% on the IndoorCVPR_09 dataset compared with the VisionTransformer method, and the recognition accuracy is improved by 3.2% on the Scene15 dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The specific effects of the present invention will be further described below in conjunction with the accompanying drawings and embodiments, in which:
[0041] Figure 1 is a flow chart of a multi-stage deep network indoor scene recognition method based on a multi-head attention mechanism in an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of Trivial augmentation data enhancement in an embodiment of the present invention;
[0043] Figure 3 It is a multi-stage deep network structure diagram based on a multi-head attention mechanism in an embodiment of the present invention;
[0044] Figure 4 It is a structural diagram of a multi-stage deep network indoor scene recognition device based on a multi-head attention mechanism in an embodiment of the present invention. DETAILED DESCRIPTION
[0045] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0046] The multi-stage deep network indoor scene recognition method based on multi-head attention mechanism combines deep convolution and multi-head attention mechanism, and adopts multi-stage deep feature extraction to mine the local features and deep feature information of indoor scene images, so as to obtain better detection results.
[0047] Embodiment 1:
[0048] Reference Figure 1 Specifically, a multi-stage deep network indoor scene recognition method based on the multi-head attention mechanism provided in this embodiment specifically includes the following steps:
[0049] S1: Obtain a dataset of indoor scenes, and divide the dataset into a training set and a test set according to a ratio;
[0050] Specifically, in step S1: Use a sentiment robot to collect scene image data of common indoor interaction environments, and merge it with the IndoorCVPR_09 related dataset to make a dataset of indoor scenes. At the same time, divide the dataset into a training set and a test set according to a ratio.
[0051] (1) Manually classify the indoor scene images collected by the chest camera of the sentiment robot. The indoor scene images of the same category are stored in the same folder, and the folder is named the category name of the indoor scene images. And convert all the collected images into RGB images.
[0052] (2) Generate a new indoor scene dataset from the collected indoor scene images of the sentiment robot and the IndoorCVPR_09 dataset, and divide the dataset into a training set and a test set at a ratio of 8:2; The dataset scene images include 20 types of indoor scenes such as meeting rooms, classrooms, kitchens, living rooms, and bedrooms.
[0053] It should be noted that the selection of the IndoorCVPR_09 related dataset, the division ratio of the dataset, and the types of dataset scene images can all be adjusted or replaced according to the actual situation and requirements.
[0054] S2: Preprocess and perform data augmentation on the indoor scene images in the training set to obtain enhanced images;
[0055] Specifically, the preprocessing process is: Resize the image size in the database to 224*224, and change the number of image channels to 3;
[0056] Specifically, this embodiment preferably uses Trivial augmentation data augmentation for data augmentation processing, which can be referred to Figure 2 , Figure 2 is the schematic diagram of Trivial augmentation data augmentation in the embodiment of the present invention. Add a set A of data augmentation functions as the input. The data augmentation functions in set A include rotation, translation, flipping, equalization, pixel value flipping, brightness, etc. Each augmentation function has its corresponding augmentation range. Randomly sample a data augmentation function from A, and then uniformly sample a value from the augmentation range {0, 1, 2…, N} as the intensity m, and then perform data augmentation on the input image and return the enhanced image.
[0057] It should be noted that in this embodiment, it is preferably taken that N = 30. In other embodiments, the value of N can be adjusted according to the actual situation and requirements.
[0058] S3: Input the enhanced image into a convolutional layer network with 3 different strides for downsampling to reduce the size of the image while retaining more feature information and local information;
[0059] Specifically, in step S3: Input the image enhanced through data augmentation into a convolutional layer network with 3 different strides for downsampling to reduce the size of the image while retaining more feature information and local information. The first convolutional layer uses a 3x3 convolution with a stride of 2 and 32 output channels to perform an operation of downsampling the input image, reducing the size of the input image and retaining more feature information for the next layer. Then, take the output of the previous convolutional layer as the input and use two 3x3 convolutions with a stride of 1 to obtain better local information. The settings of the above three convolutional layers enable the model to converge better, and at the same time, the training of the model is not too sensitive to the setting of hyperparameters.
[0060] S4: Input the feature information extracted in S3 into the backbone network, and use depth convolution and multi-head self-attention mechanism to perform deep feature extraction in a multi-stage training manner to obtain deep feature information;
[0061] It should be noted that the backbone network in this embodiment is divided into four stages to generate feature maps of different scales. In order to generate hierarchical representations, a 2x2 convolution layer with a stride of 2 is applied before each stage to reduce the size of the intermediate features (the resolution is reduced by 2 times) and project them to a larger dimension (the dimension is expanded by 2 times). In each stage, several multi-head attention mechanisms are stacked sequentially for feature transformation while maintaining the same resolution of the input.
[0062] In each stage, there is also a depth convolution layer DW before the multi-head attention mechanism for local feature extraction while reducing the computational amount. The formula description of the depth convolution layer DW is:
[0063] DW(X) = DWConv(X) + X
[0064] where X represents the current input image, X ∈ H×W×d, H×W is the resolution of the current input, d represents the dimension of the feature, and DWConv represents depth convolution;
[0065] In the multi-head attention mechanism, a depth convolution with k×k and a stride of k is used, so that the inputs of the multi-head attention mechanism are respectively H i ×W i ×C i 、 and Then the Q corresponding to the multi-head attention mechanism i , K i , V i are respectively:
[0066] Q i =(H i ×W i )×C i
[0067]
[0068]
[0069] Among them, H i , W i and C i respectively represent the height, width and number of channels of the input feature map i;
[0070] According to the principle that the sequences of Q i and K i are consistent and the number of tokens of K i and V i is consistent, the calculation formula of the multi-head attention mechanism (MultiHead) is:
[0071]
[0072] MultiHead(Q i , K i , V i ) = Concat(head1,…,head h )W O
[0073]
[0074] Among them, represents the dimension of K i , T represents transpose, Attention() represents an independent attention mechanism, and the multi-head attention mechanism MultiHead() is to perform a Concat operation on multiple independent attention mechanisms, that is, perform feature layer splicing in the dimension and perform transformation through the transformation matrix W O ; head i represents the independent attention mechanism operation performed by the i-th attention mechanism, h represents the number of attention mechanisms, and one attention mechanism corresponds to one feature map, respectively represent the transformation matrices corresponding to Q i , K i , V i ;
[0075] The multi-head attention mechanism can dynamically adjust the weight values to obtain more local feature information and global feature information. At the same time, the present invention also adopts a multi-stage training mode. In each stage, the depth convolution and the multi-head attention mechanism are sequentially stacked for feature transformation while maintaining the same resolution of the input, so as to be able to extract the depth features of the indoor scene and improve the recognition accuracy of the indoor scene.
[0076] S5: Input the depth feature information into the pooling layer, the fully connected layer, and the classifier in sequence to obtain the final weight value and the trained classifier;
[0077] Specifically, step S5 ends with a global average pooling layer, a fully connected layer, and a 1000-way classification layer with softmax to obtain the final weight value and the trained classifier.
[0078] S6: Use the trained classifier and the final weight value to test the indoor scene images in the test set, so as to determine the indoor scene category;
[0079] Step S6 specifically includes:
[0080] Take the indoor scene images in the test set as the images to be detected and input them into the trained classifier;
[0081] Adjust the size of the image to be detected according to the preset requirements. In this embodiment, it is preferably adjusted to 224*224*3, which is the same as the preprocessing size in step S2, to generate the first detection image;
[0082] Send the first detection image to the backbone network for depth feature extraction and matching recognition to generate classification recognition information and the classification probability value corresponding to the classification recognition information;
[0083] Judge whether the classification probability value is greater than the preset classification probability threshold, such as 60%. If so, take the detection frame and the classification recognition information as the recognition classification result; if not, continue to compare the remaining classification probability values until the classification probability value is greater than the preset classification probability threshold to obtain the recognition result.
[0084] Reference Figure 3 , Figure 3This is the network structure diagram of this embodiment based on deep convolution and multi-head attention mechanism. The input is the original image of the database and the data augmentation function. Through the Trivial augmentation data augmentation module, the enhanced indoor scene image is obtained. Local feature extraction is performed through three convolutional layers with different strides, while more feature information is retained. The multi-stage training method is adopted, and deep feature extraction is performed by sequentially stacking the deep convolution and multi-head attention mechanism. After the final pooling layer and fully connected layer, the finally trained classifier is obtained for the recognition and detection of the final indoor scene.
[0085] As Figure 3 shown, the multi-stage deep network indoor scene recognition method based on the multi-head attention mechanism in this embodiment includes three convolutional layers for downsampling and retaining more feature information, and four stages of stacked deep convolutional layers and multi-head attention mechanism for deep feature extraction, one pooling layer, one fully connected layer, and one classification output layer. By combining the convolutional layer and the multi-head attention mechanism, referring to the R50 model structure, the multi-stage stacking of the multi-head attention mechanism is adopted for deep feature extraction, mining the local features and deep feature information of the indoor scene image, so as to obtain better detection results.
[0086] In this embodiment, a multi-stage deep network indoor scene recognition method based on the multi-head attention mechanism is proposed. Experiments on the emotional robot interaction environment scene recognition are carried out using a deep neural network, and the experimental results are analyzed and verified. The experimental results are shown in Table 1:
[0087] Table 1 Comparison of experimental results between the method of the present invention and the Vision Transformer method
[0088] IndoorCVPR_09 Scene15 Vision Transformer 82.1% 90.3% The method of the present invention 86.4% 93.5%
[0089] From the experimental results in Table 1, the method proposed by the present invention has an improvement of 4.3% in recognition accuracy compared with the Vision Transformer method on the IndoorCVPR_09 dataset, and an improvement of 3.2% in recognition accuracy on the Scene15 dataset. Therefore, the combination of the deep convolutional neural network and the multi-head attention mechanism effectively improves the recognition accuracy of indoor scene images.
[0090] Embodiment 2:
[0091] Referring to Figure 4 , this embodiment provides a multi-stage deep network indoor scene recognition device based on the multi-head attention mechanism, including the following modules:
[0092] The dataset acquisition module 1 is used to acquire the dataset of the indoor scene and divide the dataset into a training set and a test set according to a ratio;
[0093] A data augmentation module 2, which is used to preprocess and perform data augmentation on indoor scene images in the training set to obtain augmented images;
[0094] A downsampling module 3, which is used to input the augmented images into a convolutional layer network with 3 different strides for downsampling, reduce the size of the images, and at the same time retain more feature information and local information;
[0095] A deep feature extraction module 4, which is used to input the feature information extracted by the downsampling module into a backbone network, and use depth convolution and multi-head self-attention mechanisms to perform deep feature extraction in a multi-stage training manner to obtain deep feature information;
[0096] A classifier acquisition module 5, which is used to input the deep feature information into a pooling layer, a fully connected layer, and a classifier in sequence to obtain the final weights and the trained classifier;
[0097] A scene recognition module 6, which is used to test the indoor scene images in the test set by using the trained classifier and the final weights, so as to determine the indoor scene category.
[0098] Embodiment 3:
[0099] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the multi-stage deep network indoor scene recognition method described in Embodiment 1, and can achieve the same technical effects, which will not be elaborated here.
[0100] Embodiment 4:
[0101] This embodiment provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the multi-stage deep network indoor scene recognition method described in Embodiment 1, and can achieve the same technical effects, which will not be elaborated here.
[0102] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0103] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. Among the several apparatus claims listing several apparatuses, several of these apparatuses may be embodied by the same hardware item. The use of the terms "first", "second", and "third", etc. does not denote any order and these terms may be construed as identifiers.
[0104] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A multi-stage deep network indoor scene recognition method based on the multi-head attention mechanism, characterized in that It includes the following steps: S1: Obtain a dataset of indoor scenes and divide the dataset into a training set and a test set according to a ratio; S2: Preprocess and perform data augmentation on the indoor scene images in the training set to obtain augmented images; In step S2, the Trivial augmentation method is used to perform data augmentation on the preprocessed indoor scene images, specifically including: Add a set of data augmentation functions A As input, the set A of data augmentation functions includes rotation, translation, flipping, equalization, pixel value flipping, and brightness, and each data augmentation function has its corresponding augmentation range {0, 1, 2…, N}; Randomly sample a data augmentation function from A and uniformly sample a value from the augmentation range {0, 1, 2, …, N} as the intensity m , where N represents any positive integer. Perform data augmentation on the input image according to the intensity m and return the augmented image; S3: Input the augmented images into a convolutional layer network with 3 different strides for downsampling to reduce the size of the images while retaining more feature information and local information; S4: Input the feature information extracted in S3 into the backbone network, and use depth convolution and multi-head self-attention mechanisms to perform deep feature extraction in a multi-stage training manner to obtain deep feature information; In step S4, the backbone network is divided into four stages to generate feature maps of different scales. To produce hierarchical representations, a 2x2 convolutional layer with a stride of 2 is added before each stage to reduce the size of the intermediate features and project them into a larger dimension; In each stage, there is also a depthwise convolutional layer DW before the multi-head attention mechanism. The depthwise convolutional layer is used for local feature extraction while reducing the computational amount; the multi-head attention mechanism is used to dynamically adjust the weight values to obtain more local feature information and global feature information; S5: Input the deep feature information into a pooling layer, a fully connected layer, and a classifier in sequence to obtain the final weights and the trained classifier; S6: Use the trained classifier and the final weights to test the indoor scene images in the test set to determine the indoor scene category.
2. The indoor scene recognition method of the multi-stage deep network based on the multi-head attention mechanism according to claim 1, characterized in that In step S1, the step of obtaining the dataset of indoor scenes includes: Use a sentiment robot to collect scene image data of common indoor interaction environments and merge it with the IndoorCVPR_09 related dataset to produce a dataset of indoor scenes.
3. The indoor scene recognition method of the multi-stage deep network based on the multi-head attention mechanism according to claim 1, wherein In step S3, the step of inputting the augmented images into a convolutional layer network with 3 different strides for downsampling includes: Take the augmented images as input images and input them into a convolutional layer network with 3 different strides; The first convolutional layer uses a 3x3 convolutional kernel with a stride of 2 and an output channel of 32 to perform a downsampling operation on the input images to reduce the size of the input images and retain more feature information; Take the output of the previous convolutional layer as input and use two 3x3 convolutional kernels with a stride of 1 to obtain better local information.
4. The indoor scene recognition method of the multi-stage deep network based on the multi-head attention mechanism according to claim 1, characterized in that In step S5, it ends with a global average pooling layer, a fully connected layer, and a 1000-way classification layer with softmax to obtain the final weights and the trained classifier.
5. The indoor scene recognition method of the multi-stage deep network based on the multi-head attention mechanism according to claim 1, characterized in that, In step S6, the step of using the trained classifier and the final weights to test the indoor scene images in the test set includes: Take the indoor scene images in the test set as the images to be detected and input them into the trained classifier; Adjust the size of the images to be detected according to the preset requirements to generate the first detection images; The first detected image is transmitted to the backbone network for depth feature extraction and matching recognition, generating classification recognition information and a classification probability value corresponding to the classification recognition information; Determine whether the classification probability value is greater than a preset classification probability threshold. If so, the detection frame and the classification recognition information are used as the recognized classification result; if not, continue to compare the remaining classification probability values until the classification probability value is greater than the preset classification probability threshold to obtain the recognition result.
6. An indoor scene recognition device for a multi-stage deep network based on a multi-head attention mechanism, characterized in that, It includes the following modules: A dataset acquisition module, used to acquire a dataset of an indoor scene and divide the dataset into a training set and a test set according to a ratio; A data augmentation module, used to perform preprocessing and data augmentation on the indoor scene images in the training set to obtain enhanced images, specifically including: Using the Trivial augmentation method to perform data augmentation on the preprocessed indoor scene images, specifically including: Add a set of data augmentation functions A As input, the set A of data augmentation functions includes rotation, translation, flipping, equalization, pixel value flipping, and brightness, and each data augmentation function has its corresponding augmentation range {0, 1, 2…, N}; Randomly sample a data augmentation function from A and uniformly sample a value from the augmentation range {0, 1, 2, …, N} as the intensity m , where N represents any positive integer. Perform data augmentation on the input image according to the intensity m and return the augmented image; A downsampling module, used to input the enhanced images into a convolutional layer network with 3 different strides for downsampling, reducing the size of the images while retaining more feature information and local information; A depth feature extraction module, used to input the feature information extracted by the downsampling module into the backbone network, and use depth convolution and multi-head self-attention mechanisms to perform depth feature extraction in a multi-stage training manner to obtain depth feature information. The specific steps include: The backbone network is divided into four stages to generate feature maps of different scales. To generate hierarchical representations, a 2x2 convolutional layer with a stride of 2 is added before each stage to reduce the size of the intermediate features and project them into a larger dimension; In each stage, there is also a depth convolutional layer DW before the multi-head attention mechanism. The depth convolutional layer is used for local feature extraction while reducing the computational amount; the multi-head attention mechanism is used to dynamically adjust the weight values to obtain more local feature information and global feature information; A classifier acquisition module, used to sequentially input the depth feature information into a pooling layer, a fully connected layer, and a classifier to obtain the final weights and the trained classifier; A scene recognition module, used to test the indoor scene images in the test set using the trained classifier and the final weights to determine the indoor scene category.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multi-stage deep network indoor scene recognition method according to any one of claims 1-5.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-stage deep network indoor scene recognition method according to any one of claims 1-5.