A detection method based on an adaptive improved YOLOv7 network carrier inspection vehicle

By improving the YOLOv7 network and combining it with SLAM technology, the intelligent inspection vehicle achieved autonomous navigation and accurate identification in the power plant, solving the problems of low identification accuracy and inflexible navigation in the existing technology, and improving the safety and efficiency of the plant.

CN117237909BActive Publication Date: 2025-11-28CHINA THREE GORGES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310966210.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-11-28
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

Existing intelligent inspection vehicles have unsatisfactory recognition accuracy and high error rate in power plants. They are unable to navigate autonomously and stop automatically in complex environments, lack flexibility, cannot provide proactive tracking processing, and the existing network models have poor generalization ability.

Method used

An adaptively improved YOLOv7 network is adopted, combined with image recognition and following technology, and utilizes a Haar classifier and ultrasonic ranging module to achieve accurate tracking and obstacle avoidance. The model's generalization is improved by modifying the anchor box and meta-ACON activation function, and autonomous navigation is achieved by combining SLAM technology and D*lite pathfinding algorithm.

Benefits of technology

It improves factory safety and work efficiency, reduces the complexity of factory operations, enables autonomous navigation and automatic parking, enhances recognition accuracy and flexibility in complex environments, and improves recognition accuracy and system generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237909B_ABST
    Figure CN117237909B_ABST
Patent Text Reader

Abstract

A detection method based on an adaptive improved YOLOv7 network carrier inspection vehicle, comprising the following steps: 1) using the construction intelligent platform and man-machine interaction, the camera input end is connected with the development version output end, the cloud server is connected with the mobile phone APP, the cloud server signal is detected and the data is recorded; 2) the operator adopts face unlocking inspection vehicle, the system collects and records data; 3) the inspection vehicle starts following mode, the distance acquisition device is arranged at the front of the vehicle body and is used for sensing distance and direction; 4) the improved YOLOv7 model is trained, the model is used for identifying illegal operation and product, and the output is displayed through the display and the mobile phone APP; 5) through the identification result, the inferior product is classified and autonomously navigated to the packaging point, and the packaging work is carried out by the special officer, and after the patrol vehicle completes the assigned task, the vehicle automatically cruises to the specified parking point. According to the application, by combining tracking technology with image recognition technology, the complexity of factory operation and the worker inspection time are effectively reduced, and the staff error behavior is efficiently corrected.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and intelligent detection, in particular, to a detection method of a carrier inspection vehicle based on an adaptive improved YOLOv7 network. BACKGROUND

[0002] In the current power plant pattern, the integration of transformative technologies including the Internet of Things, big data analysis, and mobile applications has triggered a new round of industrial revolution. Industrial transformation, including concepts such as digital factories, smart factories, and intelligent manufacturing, has shifted from theoretical discussions to practical implementations. Smart factory examples are built on the foundation of digital factories, utilizing the Internet of Things and monitoring technologies to enhance information management and services, thereby improving overall factory efficiency. This approach aims to improve the controllability of the production process, minimize human intervention on the production line, and optimize planning and scheduling activities. The core goal of a smart factory is to establish an efficient, energy-saving, green, and people-oriented manufacturing environment through the integration of intelligent tools and systems, ultimately enabling effective human-machine interaction. However, the current intelligent inspection vehicles used in power plants have inherent limitations, including suboptimal recognition accuracy, high error rates, and challenges in navigating complex environments. Additionally, these vehicles often lack flexibility in determining travel routes, failing to provide functions such as active tracking, autonomous navigation, and automatic parking. Patent document with publication number CN114937247A discloses a substation monitoring method and system based on deep learning, which adopts Mixup to construct a data-enhanced oil leakage image dataset, converts the dataset into three categories, and uses a PP-YOLO model for training to achieve real-time monitoring of substation oil leakage. Patent document with publication number CN111291691A discloses a substation secondary equipment instrument panel reading detection method based on deep learning, which uses the YOLOv3 algorithm to position and classify training of instrument panel images, uses OPENCV OCR to recognize images, and determines the instrument type based on the characters in the instrument image. The network model in the above-mentioned method is the original model, which greatly limits its real-time performance and accuracy in practical applications, and also has the problems of poor generalization ability, high error rate, and low recognition accuracy in complex environments.

[0003] In summary, there is an urgent need to introduce a new generation of intelligent automatic following vehicles to replace the traditional manual inspection mode and become a key technical solution for power plant operation transformation. SUMMARY

[0004] The purpose of the present application is to provide a detection method of a carrier inspection vehicle based on an adaptive improved YOLOv7 network, which expands and combines image recognition and following technology, effectively improving factory safety and work efficiency, and reducing the complexity of factory operation and worker inspection time.

[0005] To solve the above technical problems, the technical scheme adopted by the present application is:

[0006] A detection method based on an adaptive improved YOLOv7 network carrier inspection vehicle, which uses the following steps when in use:

[0007] Step 1 uses the intelligent platform and human-computer interaction to connect the camera input end to the development board output end, the cloud server to the mobile phone APP, detects the cloud server signal and records the data;

[0008] Step 2, the operator uses face unlocking inspection vehicle, the system collects and records data, and constantly tracks and follows the actions of the staff;

[0009] Step 3, the inspection vehicle starts the following mode, the distance acquisition device is arranged at the front of the vehicle body, which is used to sense the distance and direction;

[0010] Step 4 uses the improved YOLOv7 model to identify illegal operations and products, and outputs through the display and mobile phone APP;

[0011] Step 5, according to the identification result, the inferior products are classified and navigated to the packaging point, and the special officer carries out the packaging work, and the patrol vehicle automatically cruises to the designated parking point after completing the assigned task.

[0012] In step 1, the system of the inspection vehicle is composed of a development board, a camera input, a cloud server and a mobile application APP; the development board is used to build an intelligent platform and human-computer interaction, the camera input end is connected to the development board input end, and the development board signal interacts with the lower computer through the Bluetooth module; the cloud server connects the mobile phone APP, the mobile application APP controls whether the development board runs the related program through the cloud server, the cloud server combines with the WiFi module as the center hub of data transmission and reception, mainly realizes the functions including following mode, product search, product classification, etc., at the same time, the application program can also be used as a gateway to access information from the cloud server, allowing users to control the path of the inspection vehicle and open it to the designated parking position.

[0013] In step 2, the following steps are included:

[0014] Step 2-1: the camera performs specific identification on the user, uses face information, applies Haar classifier to cascade and analyze different individual face and limb symbol data, completes face unlocking, and the staff can conveniently unlock the inspection vehicle;

[0015] Step 2-2: the camera constantly tracks and follows the actions of the staff, ensures that the inspection vehicle accurately follows the actions of the staff, effectively prevents the situation of tracking the wrong target, and improves the overall safety and efficiency of the inspection process.

[0016] In step 3, the inspection vehicle uses an ultrasonic ranging module with an echo detection method to accurately measure the distance and prevent any potential collision between the inspection vehicle and the user. The system continuously receives distance information from the ultrasonic distance sensor and adjusts the speed of the inspection vehicle accordingly to maintain a certain safety distance from the user. If a significant distance difference is detected, the system will start the calibration process.

[0017] In step 4, the detection anchor box is improved to be more accurate, and an adaptive function is used to increase the model generalization, which includes the following steps:

[0018] Step 4-1, use K-means++ algorithm to calculate the initial anchor box, and improve the selection of initial anchor box to potentially improve the performance of target detection model;

[0019] The steps of K-means++ based anchor clustering are as follows:

[0020] 1) Randomly select a sample target box as the initial cluster center, calculate the minimum intersection over union (IOU) distance between the remaining sample boxes and the current cluster center as follows:

[0021] A(x) = 1 - I (x,c) (1)

[0022] Where I (x,c) represents the intersection ratio of x and c, x is the sub-target labeled sample box, and c is the cluster center;

[0023] 2) Calculate the probability O(x) of each insulator sample box being selected as the next cluster center, and use the roulette method to select the next cluster center;

[0024]

[0025] Where x is the total sample of the target labeled frame, and A(x) is the shortest distance between each sample and the existing cluster center;

[0026] 3) Repeat steps 1) and 2) until K cluster centers are selected;

[0027] 4) Calculate the distance of each sample x to the K cluster centers, and divide the sample into the class corresponding to the smallest distance cluster center, and recalculate the cluster center of each class c l , and update the classification and cluster center until the anchor box size is unchanged;

[0028]

[0029] In the formula: l = 1, …, K; K is the number of anchor boxes of different sizes, the value of which is determined by the number of anchor boxes of the detection model; the detection model includes 3 detection feature maps, each of which corresponds to three anchor boxes;

[0030] Step 4-2, using meta-ACON activation function to improve the generalization and transmission performance of the model; the meta-ACON activation function designs an adaptive function to calculate β, through two convolutional layers, so that all pixels in each channel share a weight, and finally the β is obtained through the Sigmoid function;

[0031]

[0032] In the formula, σ is the sigmoid activation function, W1 and W2 are convolutional layers, x c,h,w is an input vector containing three spaces of layers, channels and pixels;

[0033] A new architecture design space is provided in meta-ACON, because the switching factor determines the nonlinearity in activation, meta-ACON learns a more extensive distribution than ACON, and each sample has its own switch factor, rather than sharing one, which helps to improve the generalization and transmission performance of the activation behavior;

[0034] Step 4-3, resize the input picture to a specified size and input it into the improved YOLOv7 backbone network to extract high-level feature representations from the input picture for subsequent target detection tasks.

[0035] In step 4-3, the improved YOLOv7 backbone network is: after the first CBM module, the second CBM module, the third CBM module, and the fourth CBM module in the feature extraction module, the feature map F1 is obtained→ after the ELAN module composed of multiple CBS, the feature map F2 is obtained→ after the MP layer realizes spatial down-sampling, the feature map F3 is obtained→ after the ELAN module, the feature map F4 is obtained→ after the MP layer and the ELAN module, the feature map F5 is obtained→ after the MP layer and the ELAN module again, the feature map F6 is obtained;

[0036] For the 32 times down-sampling feature map output by the backbone, after the SPPCSPC module, the feature map F7 is obtained→ after the up-sampling UP operation→ the concat operation with the feature map F8 after the CBM→ the feature fusion after the ELANH module→ the up-sampling UP operation again→ the concat operation with the feature map F9 after the CBM→ the feature fusion after the ELANH module, the feature map F10 is obtained;

[0037] The feature map F10 is adjusted in channel number by RepConv, and then a Conv is used to predict small-sized objects;

[0038] The feature map F10 that has passed through the ELANH module is further processed by an MP layer, and then is concatenated with the feature map F11 that has passed through the ELANH module, and then is fused by the ELANH module to obtain a feature map F13, which is adjusted in channel number by RepConv, and then a Conv is used to predict medium-sized objects;

[0039] The feature map F13 that has passed through the ELANH module is further processed by an MP layer, and then is concatenated with the feature map F7 that has passed through the SPPCSPC module, and then is fused by the ELANH module to obtain a feature map F14, which is adjusted in channel number by RepConv, and then a Conv is used to predict large-sized objects.

[0040] In step 5, the system uses a development board to realize real-time image recognition of the captured camera image, and accurately determines the type of product existing in the factory, and the development board interacts with the lower computer through Bluetooth string, and once the inspection process is completed, the intelligent inspection vehicle automatically navigates to the initial point and prepares for the next task.

[0041] The mobile application APP provides basic information for the staff, such as the current inspection vehicle area, the product with errors, the phenomenon that the staff does not wear work clothes or safety hats, and the factory safety detection. Through the provision of such an interactive interface, the inspection vehicle aims to ensure the safety of the factory, and at the same time, it can also avoid the situation that the staff cannot start the inspection vehicle due to the temporary inability to use the mobile phone. The display interface is a means for the staff to interact with the inspection vehicle, and simplifies the complexity of the inspection.

[0042] In order to prevent the phenomenon that the inspection vehicle follows the target in confusion, the real-time video stream information collected by the image sensor is used, and the inspection vehicle uses the face and body landmark information of the staff to ensure accurate tracking. In order to achieve this, a Haar classifier is applied to cascade and analyze the face and body symbol data of different individuals to realize accurate recognition.

[0043] The intelligent inspection vehicle uses SLAM technology and D*lite pathfinding algorithm, which enables the inspection vehicle to determine its current pose, coordinates, and dynamically plan its local path. The intelligent inspection vehicle starts from an unknown position, gradually constructs a spatial navigation map, and constantly updates its own position to realize accurate and instant positioning.

[0044] When the worker unlocks the inspection vehicle using facial recognition, the laser radar system initiates automatic tracking, allowing the inspection vehicle to open its work interface. The integrated camera recognizes objects using the YOLOv7 algorithm, enabling the inspection vehicle to retrieve corresponding products and display the recognition results on the screen. The worker can place defective products into the inspection vehicle's car barrier, and through the inspection vehicle's autonomous navigation system, the process of collecting defective products is completed, making the inspection process more convenient.

[0045] Compared with the prior art, the present application has the following beneficial effects: image recognition and self-following technology are integrated, relying on a cloud server platform, an APP user login interface recognizes workers, a face is used to unlock the inspection vehicle, a camera detects and recognizes images of products and related workers, and after integration and calculation, a user interface displays illegal operations and unqualified product numbers. The worker uploads the information to the cloud server, the intelligent inspection vehicle broadcasts criticism, and at the same time, the unqualified products are sent to the designated location by self-navigation, and are packed by the worker. The intelligent inspection vehicle enhances the overall efficiency of the factory to a certain extent, has high completion degree and various functions. In the top-down design of the system, a Raspberry Pi controls an STM32 to drive a motor. A variety of sensors are used to complete the switching of different working modes of the unmanned inspection vehicle.

[0046] The present application introduces an unmanned inspection vehicle system that can be controlled through a mobile application. After completing the inspection, the inspection vehicle can be activated in follow-up mode through the application, allowing it to automatically navigate to the optimal route to deliver the substandard products to the packing point. This eliminates the need for manual trolley transportation, improving convenience and efficiency.

[0047] The system uses improved YOLOv7 for model training, providing robust and general target recognition capabilities. The YOLO model has strong generalization ability and can accurately identify various products, making it more adaptable to different scenarios.

[0048] It has strong practicality, and the combination of radar and ultrasonic technology provides accurate and comprehensive tracking capabilities. This multi-angle, close-range tracking system is superior to traditional infrared tracking methods and is suitable for complex and dynamic factory environments.

[0049] It has good real-time performance, and compared with traditional image recognition methods, the YOLO model training provides greater convenience and operational efficiency. It provides real-time results, fast response time, and by using a custom dataset for training, it can significantly improve the accuracy of product and worker identification.

[0050] The system's personalized service has high flexibility. The mobile application records data, which is stored in the cloud for easy access. Workers can view their historical records and receive records of unqualified products and worker violations through the application and the inspection vehicle. BRIEF DESCRIPTION OF DRAWINGS

[0051] The application will be further described below in conjunction with the accompanying drawings and embodiments:

[0052] Figure 1 It is the overall structure schematic diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0053] Figure 2 It is the working process schematic diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0054] Figure 3 It is the improved YOLO model process schematic diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0055] Figure 4 It is the product identification overall schematic diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0056] Figure 5 It is the image recognition demonstration diagram schematic diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0057] Figure 6 It is the YOLOv7 structure schematic diagram of the application;

[0058] Figure 7 It is the worker identification effect diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0059] Figure 8 It is the dial damage identification effect diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application;

[0060] Figure 9 It is the training and testing effect diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application for detecting workers without wearing safety helmets;

[0061] Figure 10 It is the training and testing effect diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application for detecting workers without wearing safety helmets;

[0062] Figure 11 It is the training and testing effect diagram of the intelligent inspection vehicle based on the adaptive improved YOLOv7 network carrier of the application for detecting workers without wearing safety helmets. DETAILED DESCRIPTION

[0063] A detection method based on an adaptive improved YOLOv7 network carrier inspection vehicle, which uses the following steps when in use:

[0064] Step 1 uses the intelligent platform to build human-computer interaction, the camera input end is connected to the development version output end, the cloud server is connected to the mobile phone APP, the cloud server signal is detected and data is recorded;

[0065] In step 2, the operator uses face unlocking to patrol the car, the system collects and records data, and constantly tracks and follows the actions of the staff;

[0066] In step 3, the patrol car starts the following mode, and the distance acquisition device is arranged at the front of the car body for sensing distance and direction;

[0067] Step 4 uses the improved YOLOv7 model to identify illegal operations and products, and the output is displayed through the display and the mobile phone APP;

[0068] Step 5, through the identification result, the inferior products are classified and navigated to the packaging point by the staff, and the patrol car automatically cruises to the designated parking point after completing the assigned task.

[0069] In step 1, the system of the patrol car is composed of a development board, a camera input, a cloud server and a mobile application APP; the development board is used to build an intelligent platform and human-computer interaction, the camera input end is connected to the development version input end, and the development board signal interacts with the lower computer through the Bluetooth module; the cloud server is connected to the mobile phone APP, the mobile application APP controls whether the development board runs the related program through the cloud server, the cloud server combines the WiFi module as the center hub of data transmission and reception, and mainly realizes the functions including following mode, product search and product classification, etc. At the same time, the application program can also be used as a gateway to access information from the cloud server, allowing users to control the path of the patrol car and open it to the designated parking position.

[0070] In step 2, the following steps are included:

[0071] Step 2-1: the camera performs specific identification on the user, uses face information, applies Haar classifier to cascade and analyze different individual face and limb symbol data, completes face unlocking, and the staff can conveniently unlock the patrol car;

[0072] Step 2-2: the camera constantly tracks and follows the actions of the staff, ensures that the patrol car accurately follows the actions of the staff, effectively prevents the situation of tracking the wrong target, and improves the overall safety and efficiency of the inspection process.

[0073] In step 3, the inspection vehicle uses an ultrasonic ranging module with an echo detection method to accurately measure the distance and prevent any potential collision between the inspection vehicle and the user. The system continuously receives distance information from the ultrasonic distance sensor and adjusts the speed of the inspection vehicle accordingly to maintain a certain safe distance from the user. If a significant distance difference is detected, the system will start the calibration process.

[0074] In step 4, the detection anchor frame is improved to be more accurate, and an adaptive function is used to increase the model generalization, including the following steps:

[0075] In step 4-1, the K-means++ algorithm is used to calculate the initial anchor frame, and the selection of the initial anchor frame is improved to potentially improve the performance of the target detection model.

[0076] The steps of the K-means++ based anchor clustering are as follows:

[0077] 1) Randomly select a sample target frame as the initial cluster center, and calculate the minimum intersection-over-union distance between the remaining sample frames and the current cluster center as follows:

[0078] A(x) = 1 - I (x,c) (5)

[0079] where I (x,c) represents the intersection-over-union distance between x and c, x is the sub-target labeled sample frame, and c is the cluster center.

[0080] 2) Calculate the probability O(x) of each insulator sample frame being selected as the next cluster center, and use the roulette method to select the next cluster center;

[0081]

[0082] where x is the total sample of the target labeled frame, and A(x) is the shortest distance between each sample and the existing cluster center;

[0083] 3) Repeat steps 1) and 2) until K cluster centers are selected;

[0084] 4) Calculate the distance of each sample x to the K cluster centers, and divide the sample into the class corresponding to the smallest distance cluster center, and recalculate the cluster center of each class c l , and continuously update the classification and cluster center until the anchor frame size is unchanged;

[0085]

[0086] In the formula: l = 1, …, K; K is the number of different size anchor boxes, the value is determined by the anchor box number of the detection model; the detection model contains 3 detection feature maps, each feature map corresponds to three anchor boxes, so K takes 9;

[0087] In step 4-2, the meta-ACON activation function is used to improve the generalization and transmission performance of the model; the meta-ACON activation function designs an adaptive function to calculate β, which passes through two convolution layers, so that all pixels in each channel share a weight, and finally calculates β through the Sigmoid function;

[0088]

[0089] In the formula, σ is the sigmoid activation function, W1 and W2 are convolution layers, x c,h,w is an input vector containing three spaces of layers, channels and pixels;

[0090] A new architecture design space is provided in meta-ACON, because the switching factor determines the nonlinearity in activation, meta-ACON learns a wider distribution than ACON, and each sample has its own switch factor, rather than sharing one, which helps to improve the generalization and transmission performance of the activation behavior;

[0091] In step 4-3, the input picture is resized to 640x640, and input into the backbone network of the improved YOLOv7, and after the first CBM module 1, the second CBM module 2, the third CBM module 3, and the fourth CBM module 4 in the feature extraction module, the feature map F1(160*160*128) is obtained → After the ELAN module 5, the ELAN is composed of multiple CBS, the feature map F2(160*160*256) is obtained → After the MP layer 6, the spatial down-sampling is realized, and the feature map F3(80*80*256) is obtained → After the ELAN module 7, the feature map F4(80*80*512) is obtained → After the MP layer 8 and the ELAN module 9, the feature map F5(40*40*1024) is obtained → After the MP layer 10 and the ELAN module 11, the feature map F6(20*20*1024) is obtained again;

[0092] The 32 times down-sampling feature map of the last output of the backbone passes through the SPPCSPC module 14 to obtain a feature map F7 (20*20*512) → an up-sampling UP15 operation is performed → a concat16 operation is performed with the feature map F8 (80*80*128) passing through the CBM12 → a feature fusion is performed through the ELANH module 17 → an up-sampling UP18 operation is performed again → a concat19 operation is performed with the feature map F9 (40*40*256) passing through the CBM13 → a feature fusion is performed through the ELANH module 20, to obtain a feature map F10 (20*20*512);

[0093] The feature map F10 above passes through the RepConv31 to adjust the channel number → the Conv32 is used to predict small-sized objects;

[0094] The feature map F10 passing through the ELANH module 20 further passes through the MP layer 21 → a concat22 operation is performed with the feature map F11 (40*40*256) passing through the ELANH module 17 → a feature fusion is performed through the ELANH module 23, to obtain a feature map F13 (40*40*256) → the RepConv29 is used to adjust the channel number → the Conv30 is used to predict medium-sized objects;

[0095] The feature map F13 passing through the ELANH module 23 further passes through the MP layer 24 → a concat25 operation is performed with the feature map F7 passing through the SPPCSPC module 14 → a feature fusion is performed through the ELANH module 26, to obtain a feature map F14 (80*80*128) → the RepConv27 is used to adjust the channel number → the Conv28 is used to predict large-sized objects.

[0096] In step 5, the system uses the development board to realize real-time image recognition of the captured camera image, accurately determines the type of product existing in the factory, and interacts with the lower computer through Bluetooth string. Once the inspection process is completed, the intelligent inspection vehicle will automatically navigate to the initial point and prepare for the next task.

[0097] The development board master chip is Raspberry Pi 4B, which provides processing capability and interface with various components; the system uses YOLOv7 algorithm for target recognition; the system uses BCM43455 chip as WiFi module to promote wireless data transmission between the inspection vehicle and external devices (such as cloud server or mobile application); in addition, the TB6612FNG chip is used as a driver to realize control of the motor and ensure smooth movement of the inspection vehicle.

[0098] The display interface provides basic information for the staff, such as the current inspection vehicle area, the product with errors, the phenomenon of employees not wearing work clothes or safety hats, and the factory safety detection. By providing such an interactive interface, the inspection vehicle aims to ensure the safety of the factory, while also avoiding the situation where the staff cannot start the inspection vehicle due to the temporary unavailability of the mobile phone. The display interface is a means for the staff to interact with the inspection vehicle, simplifying the complexity of the inspection.

[0099] To prevent the phenomenon of the inspection vehicle following the target confusion. Using real-time video stream information collected by the image sensor, the inspection vehicle uses the face and body sign information of the staff to ensure accurate tracking. To achieve this, the Haar classifier is applied to cascade and analyze the face and body symbol data of different individuals to achieve accurate recognition.

[0100] The intelligent inspection vehicle uses SLAM technology and D*lite pathfinding algorithm, which enables the inspection vehicle to determine its current pose, coordinates, and dynamically plan its local path. The intelligent inspection vehicle starts from an unknown location, gradually builds a spatial navigation map, and constantly updates its own position to achieve accurate and instant positioning.

[0101] When the staff uses face recognition to unlock the inspection vehicle, the laser radar system starts automatic tracking, allowing the inspection vehicle to open its work interface. The integrated camera recognizes objects through the YOLOv7 algorithm, enabling the inspection vehicle to retrieve the corresponding product and display the recognition result on the screen. The staff can put the defective products into the inspection vehicle's car barrier, and through the inspection vehicle's autonomous navigation system, complete the centralized collection of defective products, and conveniently complete the inspection process.

[0102] In order to better understand the present application for those skilled in the art, further explanation and description are as follows:

[0103] Reference Figures 1-8 A detection method of an inspection vehicle based on an adaptive improved YOLOv7 network carrier, comprising the following steps:

[0104] Step 1 utilizes the construction of an intelligent platform with human-computer interaction, a camera connected to a cloud server development board, a cloud server connected to a mobile phone APP, detection of cloud server signals and recording of data, a mobile application (APP) as a control interface for the development board, and determination of the execution of relevant programs through the cloud server. The cloud server, combined with a WiFi module as the center of data transmission and reception, not only securely stores employee data and product numbers, but also records this information to generate an error list. At the same time, the APP controls the host computer through the cloud server to determine whether to run the relevant program, and the operator can initiate start and stop commands through the mobile phone. After receiving the permission signal, the host computer continuously retrieves signals from the cloud. Once permission is obtained, the host computer starts the program and sends instructions to the lower computer. After permission, the inspection vehicle is activated and enters the self-following mode.

[0105] Step 2 includes two steps. In step 2-1, the operator uses face unlocking for the inspection vehicle, and the system collects and records data. The Haar classifier is applied to cascade analysis of facial and body feature data from different individuals. By capturing important information through the camera, the inspection vehicle can accurately identify the owner and correspondingly unlock. This innovative approach improves the safety and convenience of the inspection vehicle system; in step 2-2, to prevent the inspection vehicle from following the target in a chaotic manner, real-time video stream information collected by the image sensor is used, and the inspection vehicle uses the face and body landmark information of the staff to ensure accurate tracking.

[0106] In step 3, the inspection vehicle starts the following mode, and the distance acquisition device is set in front of the vehicle body to perceive distance and orientation. To ensure accurate control of the car's movement, the lower computer is connected to the motor drive module, which controls the TB6612FNG chip reduction motor. The integrated ultrasonic distance module allows the system to detect nearby obstacles and dynamically adjust the speed of the cart, ensuring the safety of users and the surrounding environment. The integrated laser radar and ultrasonic distance module can detect nearby obstacles and dynamically adjust the speed of the inspection vehicle, giving priority to the safety of staff and the environment. When evaluating the distance on the left and right, if the distance on the right exceeds the distance on the left, the system will increase the speed of the left wheel, causing a slight right turn, keeping the inspection vehicle aligned with the operator. This alignment ensures that the inspection vehicle remains on the track, minimizing significant deviations from the intended path. Conversely, if the difference in distance falls within an acceptable range, no adjustment is made, and the handcart continues to move in a straight line. This approach effectively reduces unnecessary sudden changes in the car's movement when the distance difference is not significant, making navigation smoother.

[0107] Step 4 uses the improved YOLOv7 model to identify the violation operation and product, and the output is displayed through the display and mobile phone APP. The camera integrated in the inspection vehicle can identify various products in the factory and display the identification results on the built-in display screen. Through feature point detection, extraction and comparison, accurate product identification is achieved. The identified products are compared with the comprehensive product line using the YOLOv7 algorithm to retrieve the product information with the highest similarity, and a confidence threshold is established to ensure the reliability of the identification. When the calculated similarity exceeds the set threshold, it indicates that the identification is successful. The improved YOLOv7 algorithm used in this design has the characteristics of high accuracy of the calculation model and is suitable for real-time detection. For the generation of anchors, the K-means algorithm is used to cluster the anchors of the data set, and K-means++ clustering is used to compare parameters. From the results (as shown in Table 1), when 9 anchors parameters are obtained using K-means++ clustering, all evaluation indicators are significantly increased.

[0108] Table 1 Comparison of improved performance of clustering anchors algorithm

[0109]

[0110] From Table 1, it can be seen that the Fitness and Recall indicators are improved by 0.88% and 1.1% respectively after K-means++ clustering optimization. The improved clustering method generates more refined anchors. The last column of the table gives the size of the 9 groups of clustered anchors, which are self-adapted to match the target size and corrected and fine-tuned during the optimization process. The optimized results of the original anchors are used as the new input parameters of YOLOv7.

[0111] ACON (Active or Not) is an adaptive activation function that determines whether to activate neurons. This activation function helps improve the generalization ability of the network model. The unique method of ACON is that it can adjust the smoothness of the standard maximum function, which usually has sharp transitions and discontinuities. The smoothness can be approximately described as:

[0112]

[0113] where x i is the input vector; n represents the number of data set samples; β represents the switch weight coefficient. When β→∞, S β →max, at this time the entire smooth activation function behaves nonlinearly, and when β→0, S β→ mean, at this time the activation function presents linear, neither activation state. In the meta-ACON activation function, an adaptive function of β is designed on the basis of the ACON activation function, through two convolutional layers, so that all pixels in each channel share a weight, and finally the β is obtained through the Sigmoid function.

[0114]

[0115] In the formula, σ is a sigmoid activation function, W1 and W2 are convolutional layers, x c,h,w is an input vector containing three spaces of layers, channels, and pixels.

[0116] A new architecture design space is provided in meta-ACON, because the switching factor determines the nonlinearity in activation, meta-ACON learns a more extensive distribution than ACON, and each sample has its own switch factor, rather than sharing one, which helps to improve the generalization and transmission performance.

[0117] Table 2 ablation experiment

[0118]

[0119] Compared with the original precision and mAP0.5:0.95, it is increased by 1.2% and 2.5% respectively. After the YOLO algorithm identifies the most relevant key points, the system compares the identified objects with the product library, and retrieves the product with the highest similarity. When the calculated similarity exceeds the predefined confidence threshold, it indicates that the corresponding product is successfully identified.

[0120] Step 5: Based on the recognition results, the system autonomously classifies the substandard products and navigates to the packaging point. The designated staff then packages the products. After completing the assigned tasks, the patrol vehicle automatically drives to the designated parking spot. The intelligent patrol vehicle utilizes SLAM technology and D*lite pathfinding algorithms to determine its current pose and coordinates, enabling precise positioning and autonomous navigation. The SLAM technology uses sensor data and pre-existing maps to estimate the current position and orientation of the patrol vehicle within the factory space. A particle filter-based Monte Carlo localization algorithm is employed to obtain global positioning information. The system continuously constructs and updates the environmental map using sensor data and positioning results. Considering the presence of dynamic obstacles in the factory, the D*lite pathfinding algorithm enables real-time path updating and obstacle avoidance. It adjusts the planned path based on the position and motion information of dynamic obstacles to ensure collision avoidance and optimal path selection. The combination of SLAM technology and D*lite algorithms enables the intelligent patrol vehicle to accurately locate itself, navigate the environment, and avoid obstacles.

[0121] The upper computer processing chip in the development version is Raspberry Pi 4B, the lower computer is STM32 chip, the WiFi module is BCM43455 chip built-in Raspberry Pi, and the driver chip is TB6612FNG chip. Raspberry Pi 4B serves as the main control chip, which is both an intelligent platform and a controller for the interactive intelligent patrol vehicle. The driver device is controlled by the instructions of Raspberry Pi 4B and effectively cooperates with the STM32 chip to achieve smooth switching of the intelligent patrol vehicle mode. These components collectively form a powerful integrated system that enhances efficient operability.

[0122] The intelligent patrol vehicle integrated with recognition and fault detection has a clear advantage in speed. In addition, for places where a large number of substandard products need to be transported, the patrol vehicle also offers a transportation service option. In this case, the patrol vehicle activates the self-patrol function and automatically navigates to the packaging point, where designated staff can transport the items to the designated location. This convenient feature enhances the overall work efficiency of the factory and provides a convenient solution for operators who need to identify a large number of products or tedious items.

[0123] The integrated patrol vehicle system combines high-precision cameras, sensors, facial recognition, computer vision, artificial intelligence recognition, and big data computing, enabling seamless switching between different working modes. The system independently captures a comprehensive data set and uses the YOLOv7 algorithm for model training. Compared to traditional image segmentation recognition, the YOLOv7 image recognition technology offers more convenient operability for operators. By utilizing this intelligent patrol vehicle system, the system provides convenient navigation, efficient product recognition, and simplified patrol processes for staff.

Claims

1. A detection method based on an adaptive improved YOLOv7 network carrier inspection vehicle, characterized by, It uses the following steps when in use: Step 1: Use the intelligent platform built to interact with humans, connect the camera input end to the development version output end, connect the cloud server to the mobile phone APP, detect the cloud server signal and record the data; Step 2: The operator uses face unlocking to patrol the vehicle, the system collects and records data, and constantly tracks and follows the actions of the staff; Step 3: The patrol vehicle starts the following mode, and the distance acquisition device is arranged at the front of the vehicle body for sensing distance and direction; Step 4: Use the improved YOLOv7 model to identify illegal operations and products, and output through the display and mobile phone APP; Step 5: Through the identification result, the inferior products are classified and navigated to the packaging point by the staff, and the patrol vehicle automatically cruises to the designated parking point after completing the assigned task; The improved YOLOv7 model uses the K-means++ algorithm to calculate the initial anchor frame; The improved YOLOv7 model uses meta-ACON as the activation function; In the improved YOLOv7 model, the backbone network is: after the first CBM module (1), the second CBM module (2), the third CBM module (3), and the fourth CBM module (4) in the feature extraction module, a feature map F1 is obtained→after the ELAN module (5), the ELAN is composed of multiple CBS, a feature map F2 is obtained→after the MP layer (6) realizes spatial down-sampling, a feature map F3 is obtained→after the ELAN module (7), a feature map F4 is obtained→after the MP layer (8) and the ELAN module (9), a feature map F5 is obtained→after the MP layer (10) and the ELAN module (11), a feature map F6 is obtained; For the 32 times down-sampling feature map output by the backbone, after the SPPCSPC module (14), a feature map F7 is obtained→after the up-sampling UP (15) operation→the feature map F8 after the CBM (12) is concatenated (16)→after the ELANH module (17) feature fusion→again after the up-sampling UP (18) operation→the feature map F9 after the CBM (13) is concatenated (19)→after the ELANH module (20) feature fusion, a feature map F10 is obtained; The above feature map F10 is adjusted in channel number by RepConv (31)→uses Conv (32) to predict smaller size objects; The feature map F10 after the ELANH module (20) is further processed by the MP layer (21)→the feature map F11 after the ELANH module (17) is concatenated (22)→after the ELANH module (23) feature fusion, a feature map F13 is obtained→after the RepConv (29) adjusts the channel number→uses Conv (30) to predict medium size objects; The feature map F13 passing through the ELANH module (23) and then passing through the MP layer (24) is subjected to a concat (25) operation with the feature map F7 passing through the SPPCSPC module (14) and then subjected to feature fusion through the ELANH module (26) to obtain a feature map F14, which is subjected to a RepConv (27) to adjust the number of channels and then subjected to a Conv (28) to predict larger objects.

2. The detection method according to claim 1, characterized in that, In step 1, the system of the inspection vehicle is composed of a development board, a camera input, a cloud server and a mobile application APP; the development board is used to build an intelligent platform and human-computer interaction, the camera input end is connected to the input end of the development board, and the development board signal interacts with the lower machine through the Bluetooth module therein; the cloud server is connected to the mobile APP, and the mobile application APP controls whether the development board runs the related program through the cloud server; the cloud server combines the WiFi module to serve as the center hub of data transmission and reception, and mainly realizes the functions of following mode, product search and product classification. At the same time, the application program can also serve as a gateway to access information from the cloud server, allowing users to control the path of the inspection vehicle and open it to the designated parking position.

3. The method of claim 1, wherein In step 2, the following steps are specifically included: Step 2-1: the camera identifies the user, uses facial information, applies a Haar classifier to cascade and analyze different individual facial and limb symbol data, completes face unlocking, and the staff can conveniently unlock the inspection vehicle; Step 2-2: the camera constantly tracks and follows the actions of the staff, ensures that the inspection vehicle accurately follows the actions of the staff, effectively prevents the situation of tracking the wrong target, and improves the overall safety and efficiency of the inspection process.

4. The method of claim 1, wherein, In step 3, the inspection vehicle uses an ultrasonic ranging module with an echo detection method to accurately measure the distance and prevent any potential collision between the inspection vehicle and the user. The system continuously receives distance information from the ultrasonic distance sensor and adjusts the speed of the inspection vehicle accordingly to maintain a certain safety distance from the user. If a significant distance difference is detected, the system will start the calibration process.

5. The method of claim 1, wherein In step 4, the detection anchor box is improved to be more accurate, and an adaptive function is used to increase the model generalization, which specifically includes the following steps: Step 4-1, the K-means++ algorithm is used to calculate the initial anchor box, and the selection of the initial anchor box is improved to potentially improve the performance of the target detection model; The steps of the K-means++ based anchor clustering method are as follows: 1) randomly select a sample target box as the initial clustering center, and calculate the minimum intersection over union (IOU) distance between the remaining sample boxes and the current clustering center as follows: (1); In the formula denotes x , c the intersection of the two rectangular frames is parallel to the ratio, x is a sub-target label sample frame, c is a clustering center; 2) Calculate the probability of each insulator sample frame being selected as the next cluster center , the next cluster center is selected using roulette method; (2); wherein x is the total samples of the target marked frame, is the shortest distance between each sample and the current existing cluster center; 3) repeat steps 1) and 2) until a cluster center is selected K ; 4) Calculate each sample x arrive K The distance between each cluster center is calculated, and the sample is assigned to the cluster corresponding to the cluster center with the smallest distance. Each cluster is then recalculated. The cluster centers are continuously updated, and the classification and cluster centers are updated until the anchor box size remains unchanged; (3); In the formula: ; K is the number of anchor boxes of different sizes, the value of which is determined by the number of anchor boxes of the detection model; the detection model contains 3 detection feature maps, and each feature map corresponds to three anchor boxes; Step 4-2 employs the meta-ACON activation function to improve the model's generalization and transfer performance; the meta-ACON activation function is designed with a computational... β The adaptive function, through two convolutional layers, ensures that all pixels in each channel share a single weight, and is finally calculated using the Sigmoid function. β ; (4); wherein sigmoid activation function, W 1, W 2 is a convolutional layer, is an input vector of three spaces: channel, lane, pixel. A new architecture design space is provided in meta-ACON, as the switching factor determines the nonlinearity in activation, meta-ACON learns a wider distribution than ACON, and each sample has its own switching factor instead of sharing one, which helps to improve the generalization and transmission performance of the activation behavior. Step 4-3, resize the input image to the specified size and input it into the backbone network of the improved YOLOv7 to extract high-level feature representations from the input image for subsequent target detection tasks.

6. The method of claim 1, wherein, In step 5, the system uses a development board to realize real-time image recognition of the captured camera image, accurately determines the type of product present in the factory, and interacts with the lower computer through Bluetooth string, once the inspection process is completed, the intelligent inspection vehicle will automatically navigate to the initial point and prepare for the next task.

7. The method of claim 2, wherein The mobile application APP provides basic information for the staff, such as the current inspection vehicle area, the presence of defective products, the phenomenon of employees not wearing work clothes or safety hats, and factory safety detection. By providing this interactive interface, the inspection vehicle aims to ensure factory safety while avoiding situations where staff cannot start the inspection vehicle due to temporary inability to use their mobile phones. The display interface is a means for staff to interact with the inspection vehicle, simplifying the complexity of the inspection process.

8. The method of claim 3, wherein, In step 2-2, to prevent the inspection vehicle from following the target in a chaotic manner, the real-time video stream information collected by the image sensor is used, and the inspection vehicle uses the face and body landmark information of the staff to ensure accurate tracking. To achieve this, the Haar classifier is applied to cascade and analyze different individual face and body symbol data to achieve accurate recognition.

9. The detection method according to claim 1 or 2 or 3 or 8, characterized in that, When the staff use face recognition to unlock the inspection vehicle, the laser radar system starts automatic tracking, allowing the inspection vehicle to open its work interface. The integrated camera identifies objects through the YOLOv7 algorithm, allowing the inspection vehicle to retrieve the corresponding product and display the recognition result on the screen. The staff can put the defective products into the inspection vehicle's car barrier, and through the inspection vehicle's autonomous navigation system, the defective products can be collected in a centralized manner, making the inspection process more convenient.

Citation Information

Patent Citations

  • Transformer substation secondary equipment instrument panel reading detection method based on deep learning

    CN111291691A

  • A deep learning-based substation monitoring method, system, and electronic equipment.

    CN114937247A