Elevator key and floor number recognition method based on MobileNetV4-YOLOv8 deep learning algorithm

By introducing the MobileNetV4ConvSmall module and ELA attention mechanism in the YOLOv8 network model, the problem of inability to balance detection accuracy and calculation efficiency in elevator buttons and floor digital recognition is solved, and efficient and accurate recognition effect is achieved.

CN120182933APending Publication Date: 2025-06-20HEZE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510227811.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has problems in the identification of elevator buttons and floor numbers, especially when the elevator buttons are small in size, complex background, and affected by light changes, the recognition effect is poor.

Method used

The elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm are adopted. The MobileNetV4ConvSmall module is introduced into the backbone feature extraction network of the YOLOv8 network model and the ELA attention mechanism module is integrated to improve the efficiency and accuracy of feature extraction.

Benefits of technology

It realizes more efficient and accurate elevator buttons and floor digit recognition, improves detection accuracy and reduces calculation costs, and is suitable for small object detection in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182933A_ABST
    Figure CN120182933A_ABST
Patent Text Reader

Abstract

The invention provides an elevator key and floor number recognition method based on a MobileNetV4-YOLOv8 deep learning algorithm, and relates to a robot for elevator key and floor number recognition. The elevator key and floor number recognition method comprises the steps that a MobileNetV4-YOLOv8 network model is constructed, and an elevator key and floor number recognition data set is constructed; the method comprises the following steps of: training a MobileNetV4-YOLOv8 network model by adopting an ELOU loss function on the basis of an elevator key and floor digital identification data set to obtain a trained MobileNetV4-YOLOv8 network model; and carrying out elevator key and floor number identification by utilizing the trained MobileNetV4-YOLOv8 network model. According to the method, the problems of complex background, small key target size and unobvious key characteristics in elevator key and floor number recognition of a robot in an elevator can be solved, and the positions and types of the elevator keys and floor numbers can be accurately recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an efficient elevator button and floor number recognition method, specifically an elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm. Background Art

[0002] The rapid development of robot technology has brought great convenience to people's lives. Building robots can make the application of robots closer to people's actual needs. However, to work between different floors, they need to have the ability to move autonomously in buildings. Among them, taking the elevator has become a technical problem to be solved urgently.

[0003] Currently, the more mature solution in the market is to interact with the elevator control system through wireless communication technology to directly control the operation of the elevator. However, this solution has many pain points in actual applications. First of all, the elevator control systems of different buildings vary greatly. Robots need to communicate with elevators with various different protocols, which not only increases the technical difficulty but also leads to high adaptation costs. Secondly, this solution relying on elevator protocols lacks generality and is difficult to promote on a large scale. When robots are put into buildings with different systems, they need to be individually adapted and debugged for each elevator system, which greatly limits the wide application of this type of robot.

[0004] Therefore, in order to overcome the above technical problems, a new solution has gradually attracted attention, that is, using a robotic arm to directly control the elevator. By simulating the behavior of humans operating elevator buttons with the robotic arm and identifying the corresponding floors, robots can take the elevator autonomously without communicating with the elevator control system. However, to implement this solution, one of the core problems is how to accurately identify elevator buttons and floor numbers. Elevator buttons are usually small in size and have a complex background, and may be affected by various factors such as changes in light and differences in button materials, which makes button recognition a very challenging task. Existing object detection algorithms, although performing well in general object detection tasks, still have problems such as the inability to balance detection accuracy and computational efficiency in small object detection scenarios such as elevator buttons. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides an elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm, which can accurately recognize elevator buttons and floor numbers and solve the problems of the original YOLO algorithm with many parameters and low computational efficiency.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] (1) Construct the MobileNetV4-YOLOv8 network model, which introduces the MobileNetV4ConvSmall module into the backbone feature extraction network of the YOLOv8 network model to make the model more lightweight;

[0008] (2) Construct a dataset for elevator button and floor number recognition. The elevator button and floor number dataset is a publicly available dataset on Roboflow, and preprocess and augment the dataset;

[0009] (3) Based on the elevator button and floor number dataset, use the ELOU loss function to enhance the ability to extract elevator button and floor number features, train the MobileNetV4-YOLOv8 network model, and obtain a trained network model;

[0010] (4) Use the trained dataset for elevator button and floor number recognition.

[0011] Preferably, the method of introducing the MobileNetV4ConvSmall module into the backbone feature extraction network of the YOLOv8 network model is as follows:

[0012] Replace the Conv and C2f modules in the backbone network of the YOLOv8 network model with the MobileNetV4ConvSmall module.

[0013] Preferably, in the elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm, the MobileNetV4ConvSmall module consists of an initial convolutional layer and four hierarchical modules, and each hierarchical module contains multiple convolutional layers and Inverted Bottleneck modules.

[0014] Preferably, the ELOU loss function is:

[0015]

[0016] where L EIOU is the EIOU loss function, L IOU is the IOU loss, L dis is the distance loss, L asp is the height-width loss, IOU is the intersection over union of the predicted box and the ground truth box, b is the center of the predicted box, ω is the width of the predicted box, h is the height of the predicted box, b gt is the center point of the ground truth box, ω gt is the width of the ground truth box, h gtThe height of the ground truth box, c is the diagonal length of the minimum bounding rectangle of the predicted box and the ground truth box, c ω is the diagonal width of the minimum bounding rectangle of the predicted box and the ground truth box, c h is the diagonal height of the minimum bounding rectangle of the predicted box and the ground truth box, and ρ is the Euclidean distance between two points.

[0017] Preferably, the role of integrating the ELA attention mechanism module in the network is as follows:

[0018] Enhance the ability to extract features of elevator buttons and floor numbers, especially in the detection of small targets in complex backgrounds.

[0019] The beneficial effects of the present invention are: An elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm proposed by the present invention inherits the detection head network of YOLOv8, replaces the backbone network of YOLOv8 with the lightweight network mobileNetv4, and enhances the feature extraction ability by adding the ELA attention mechanism. A better elevator button and floor number recognition method is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solution of the present invention, the drawings used in the present invention will be briefly introduced below.

[0021] Figure 1 is a schematic diagram of the invention process;

[0022] Figure 2 is an improved YOLOv8 network structure diagram

[0023] Figure 3 is a structural diagram of the ELA attention mechanism module;

[0024] Figure 4 is a schematic diagram of model testing. DETAILED DESCRIPTION OF THE INVENTION

[0025] 1. Construct a MobileNetV4-YOLOv8 network model

[0026] As Figure 1 shown, the elevator button and floor number recognition method of the present invention first needs to construct a MobileNetV4-YOLOv8 network model, as follows Figure 2 . This network model realizes lightweight and efficient feature extraction by introducing the MobileNetV4ConvSmall module into the backbone feature extraction network of YOLOv8. At the same time, the ELA attention mechanism module is integrated into the backbone network to enhance the feature extraction ability.

[0027] The specific steps are as follows:

[0028] (1) Network structure design

[0029] The backbone network part of the YOLOv8 network model is mainly responsible for extracting features from the input image and providing basic information for subsequent object detection. To improve the efficiency and accuracy of feature extraction, the present invention replaces the Conv and C2f modules in the backbone network of the YOLOv8 network model with the MobileNetV4ConvSmall module. The MobileNetV4ConvSmall module is a lightweight feature extraction module that can significantly reduce the number of model parameters and computational complexity while maintaining a relatively high feature extraction accuracy, thereby improving the running efficiency of the model.

[0030] (2) Integrate the ELA attention mechanism module

[0031] To further enhance the feature extraction ability, the present invention integrates the ELA attention mechanism module into the network. As Figure 3 shown, ELA obtains feature vectors in the horizontal and vertical directions through strip pooling in the spatial dimension, maintains a narrow kernel shape to capture long-range dependencies, and prevents irrelevant regions from affecting label prediction, thereby generating rich target location features in their respective directions. In the second step, one-dimensional convolution is applied to perform local interaction with the two feature vectors, and the kernel size is adjusted to represent the local interaction range. Then the obtained feature vectors are processed with GN and non-linear activation functions to generate position attention predictions in the two directions respectively. The final position attention is obtained by multiplying the attention predictions in the two directions to obtain the position information of the region of interest.

[0032] 2. Construct an elevator button and floor number recognition dataset

[0033] Constructing a high-quality elevator button and floor number recognition dataset is one of the key steps to implement the method of the present invention. The quality of the dataset directly affects the training effect and detection performance of the model. The specific steps are as follows:

[0034] (1) Data collection

[0035] Obtain the publicly available elevator button and floor number recognition dataset on Roboflow, and obtain a large number of elevator button and floor number panel images through on-site shooting with a camera, network image search, etc. according to requirements, covering various common and special scenarios to improve the generalization ability of the model.

[0036] (2) Data preprocessing

[0037] Preprocess the collected images to increase the diversity of the dataset. These preprocessing operations can simulate elevator button and floor number panel images under different angles, scales, and lighting conditions, thereby improving the model's adaptability to various changes.

[0038] (3) Data annotation

[0039] Use the annotation tool Roboflow to annotate the elevator buttons and floor numbers in the images. The annotation content includes the category information (such as numbers, letters, function buttons) and coordinate information of the buttons. The category information indicates the category to which each button belongs, such as numeric buttons like "1", "2", "3", etc., and function buttons like "up", "down", etc.; the coordinate information annotates the position and size of each button in the image, usually represented in the form of a bounding box, including the upper-left and lower-right coordinates of the bounding box.

[0040] (4) Data division

[0041] Fuse the self-made dataset and the publicly available dataset downloaded from Roboflow and divide them into a training set, a validation set, and a test set according to the ratio of 6:2:2. Among them, 60% of the data is used as the training set to train the model, 20% of the data is used as the validation set to verify the model's performance and adjust the training parameters, and 20% of the data is used as the test set to evaluate the final performance of the model. Ensure that the model can fully learn the characteristics of the data during the training process, and can accurately evaluate the model's performance on the validation set and the test set, avoiding overfitting and underfitting phenomena.

[0042] 3. Network model training

[0043] Based on the constructed elevator button and floor number recognition dataset, use the ELOU loss function to train the MobileNetV4-YOLOv8 network model. The training process is a crucial link for the model to learn data features and optimize parameters, directly affecting the detection performance of the model. The specific steps are as follows:

[0044] (1) Model parameter setting

[0045] Parameter setting is crucial for the performance and efficiency of the model. The types and setting values of the parameters selected during training are shown in Table 1.

[0046] Table 1

[0047] Parameter type Set value Parameter type Set value epochs 120 optimizer SGD batch size 24 lr 0.01 workers 4 momentum 0.937

[0048] Among them, epochs represents the number of training rounds, batch size represents the number of samples used to calculate the gradient in each iteration, Workers represents the number of threads used during data loading, optimizer represents the optimizer, lr represents the learning rate, and momentum represents the momentum.

[0049] (2) Loss function calculation

[0050] In each training batch, the ELOU loss between the predicted bounding box and the ground truth bounding box is calculated. The ELOU loss function is an improved loss function that not only considers the intersection over union (IOU) between the predicted bounding box and the ground truth bounding box, but also considers factors such as the distance between the centers of the predicted bounding box and the ground truth bounding box, and the aspect ratio, etc., and can more comprehensively measure the difference between the predicted bounding box and the ground truth bounding box. The formula for the ELOU loss function is:

[0051]

[0052] Among them, L EIOU is the EIOU loss function, L IOU is the IOU loss, L dis is the distance loss, L asp is the height-width loss, IOU is the intersection over union between the predicted bounding box and the ground truth bounding box, b is the center of the predicted bounding box, ω is the width of the predicted bounding box, h is the height of the predicted bounding box, b gt is the center point of the ground truth bounding box, ω gt is the width of the ground truth bounding box, h gt is the height of the ground truth bounding box, c is the diagonal length of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, c ω is the diagonal width of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, c h is the diagonal height of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, ρ is the Euclidean distance between two points.

[0053] (3) Model validation and optimization

[0054] During the training process, the validation set is used to validate the model, and the training parameters are adjusted according to the validation results,

[0055] optimize the model structure, and improve the generalization ability of the model. The role of the validation set is to evaluate the performance of the model on unseen data, and through the validation set, it can be timely discovered whether the model has overfitting or underfitting phenomena.

[0056] 4. Elevator button and floor number recognition

[0057] Use the trained MobileNetV4 - YOLOv8 network model to identify elevator buttons and floor numbers. This is the final application stage of the method of the present invention. The model detects and identifies elevator buttons and floor numbers in the actual scene, providing accurate button positions and category information for the robot. The specific steps are as follows:

[0058] (1) Image input

[0059] Input the collected images of elevator buttons and floor number panels into the trained model to start forward propagation.

[0060] (2) Object detection

[0061] Through forward propagation, the model outputs the position information (bounding box coordinates) and category information of elevator buttons and floor numbers in the image. During forward propagation, the input image first passes through the backbone network to extract features, then through the neck network for feature fusion, and finally through the head network for object detection. The head network generates a series of bounding boxes, each of which contains a category prediction and a confidence score, indicating the probability that there is an object in the bounding box.

[0062] (3) Recognition output

[0063] Output the final detection results, including information such as the positions, categories, and confidences of the buttons and floor numbers. The output detection results are presented in the form of text data, providing a basis for subsequent robot control and decision - making.

[0064] 5. Experimental verification

[0065] To verify the effectiveness of the method of the present invention, the following experimental verification was carried out:

[0066] (1) Experimental environment

[0067] Use NVIDIA RTX 4060 GPU and run the PyTorch framework.

[0068] (2) Dataset

[0069] Use the above - constructed dataset for elevator button and floor number recognition for training and testing. The dataset contains a large number of images of elevator button and floor number panels in different scenarios, which can fully verify the performance of the model.

[0070] (3) Evaluation metrics

[0071] Adopt the mean average precision (mAP) metric to evaluate the model performance. mAP is used to measure the detection accuracy of the model. At the same time, to ensure the effectiveness of this algorithm, it is compared with some models of the YOLO series. The experimental results are shown in Table 2.

[0072] Table 2

[0073] Algorithm mAP(0.5) mAP(0.5:0.95) FPS YOLOv5n 0.935 0.69 239.4 YOLOv6s 0.932 0.701 149.8 YOLOv8-P2 0.773 0.573 151.3 The algorithm of the present invention 0.969 0.721 159.2

[0074] Among them, mAP(0.5) refers to the average precision calculated when the IoU (Intersection over Union) threshold is 0.5. mAP(0.5:0.95) refers to the average mAP calculated under multiple thresholds where the IoU threshold ranges from 0.5 to 0.95 (with a step size of 0.05). FPS, i.e., frames per second, is used to evaluate the processing speed. By comparing with other models, the model of the present invention has higher precision and faster processing speed.

[0075] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying elevator buttons and floor numbers based on the MobileNetV4-YOLOv8 deep learning algorithm, characterized in that: The elevator button and floor number recognition method comprises the following steps: (1) Constructing a MobileNetV4-YOLOv8 network model, wherein the MobileNetV4-YOLOv8 network model introduces the MobileNetV4ConvSmall module into the YOLOv8 network model backbone feature extraction network, and introduces the ELA attention mechanism module; (2) Constructing an elevator button and floor number recognition dataset, which is a combination of a public dataset on Roboflow and a self-made dataset; (3) Based on the elevator button and floor number dataset, the ELOU loss function is used to train the MobileNetV4-YOLOv8 network model to obtain the trained network model; (4) Use the trained data set to recognize elevator buttons and floor numbers.

2. The elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm according to claim 1 is characterized in that: The way to introduce the MobileNetV4ConvSmall module into the backbone feature extraction network of the YOLOv8 network model is: Replace the Conv and C2f modules of the backbone network of the YOLOv8 network model with the MobileNetV4ConvSmall module.

3. The elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm according to claims 1 and 2 is characterized in that: The MobileNetV4ConvSmall module consists of an initial convolutional layer and four hierarchical modules, each of which contains multiple convolutional layers and an Inverted Bottleneck module.

4. The elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm according to claim 1 is characterized in that: The ELOU loss function is: Among them, L EIOU is the EIOU loss function, L IOU is the IOU loss, L dis is the distance loss, L asp is the height-width loss, IOU is the intersection-over-union ratio between the predicted box and the real box, b is the center of the predicted box, ω is the width of the predicted box, h is the height of the predicted box, and b gt is the center point of the real frame, ω gt is the width of the real frame, h gt is the height of the real box, c is the diagonal length of the minimum circumscribed rectangular box of the predicted box and the real box, c ω is the diagonal width of the minimum bounding rectangle of the predicted box and the real box, c h is the diagonal height of the minimum circumscribed rectangular box of the predicted box and the true box, and ρ is the Euclidean distance between the two points.

5. The elevator button and floor number recognition method based on the MobileNetV4-YOLOv8 deep learning algorithm according to claim 1 is characterized in that: Integrate the ELA attention mechanism module in the network.