A Multi-Scene Traffic Target Detection Method and System Based on Lightweight Networks
By using a lightweight YOLO V3-Tiny network in the traffic signal assistance system for real-time analysis and multi-scenario applicability, the problem of poor hardware and software integration in existing systems is solved, achieving high-precision multi-scenario traffic monitoring and accident reduction.
Patent Information
- Application Number
- CN202311324506.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-12
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-10-12
Smart Images

Figure CN117274955B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing and target detection technology, and in particular relates to a multi-scene traffic target detection method and system based on lightweight networks. Background Technology
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Artificial intelligence-based embedded traffic signal assistance systems have become an important indicator of smart city development. They can lay the foundation for the Internet of Things (IoT) in smart cities, combine with emerging IoT technologies to ensure basic traffic safety in cities, create technologically innovative smart cities, and enhance the city's image.
[0004] Currently, traffic signal safety assistance systems are developing well in various cities and different traffic environments, such as radar-based road sentinels.
[0005] However, these traditional devices have significant limitations. Traditional traffic signal assistance systems typically monitor traffic using fixed-function hardware to aid in traffic flow. Furthermore, traditional embedded traffic signal assistance systems are designed solely through hardware and software integration, making program modification impossible and limiting their ability to meet specific traffic signal requirements. Their hardware is relatively simple, usually consisting of only a power supply, light source, and specific hardware components, and their level of intelligence is low, making them inadequate for the increasingly complex scenarios in the traffic field. They also suffer from low monitoring efficiency, incomplete functionality, inability to upload traffic data, and inability to dynamically adjust based on actual road conditions. These shortcomings contribute to traffic congestion and safety issues. Therefore, intelligent processing of traditional traffic assistance systems is a better solution to these problems.
[0006] In recent years, the rise of deep learning has made AI-based traffic assistance systems a hot topic again. However, due to limitations in the integration of algorithms and devices, promoting the intelligent application of embedded AI devices in traffic safety assistance systems faces significant challenges. The compatibility of most deep learning algorithms with embedded systems is the biggest obstacle to the practical application of such systems, mainly due to two factors: firstly, the small size of embedded systems leads to low computing power; secondly, the trained AI algorithms are typically large, require significant storage space, and demand high computing power, making it difficult to embed these various trained models into smaller embedded systems. These conflicting factors prevent the perfect integration of algorithms and embedded systems.
[0007] In summary, existing traffic signal assistance systems lack network processing capabilities, resulting in poor hardware-software integration. The networks they use require high computing power, leading to excessively large, complex, and heavy hardware components. Consequently, they are unable to adapt to a limited number of scenarios and cannot modify their functions according to actual traffic conditions, resulting in frequent accidents. Summary of the Invention
[0008] To overcome the shortcomings of the prior art, this invention provides a multi-scenario traffic target detection method and system based on a lightweight network. It features high hardware and software integration, small size, applicability to multiple scenarios, and high-precision real-time wireless traffic monitoring.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] The first aspect of the present invention is a multi-scenario traffic target detection method based on a lightweight network, comprising: real-time acquisition of traffic condition video data in a traffic scene and transmission to an embedded device;
[0011] The embedded device performs real-time analysis and processing of traffic video data. Based on the analysis and processing results, it determines whether the target is detected at the intersection at the same time. If the target is detected at the same time, it controls the feedback and reminder module to issue a reminder. At the same time, the analysis and processing results are encapsulated and transmitted to the server via push streaming.
[0012] The embedded device is equipped with a lightweight target detection model, which is an improved YOLO V3-Tiny network. The improved YOLO V3-Tiny network is obtained by pruning and sparse training the YOLO V3-Tiny network to reduce the number of layers and channels. The improved YOLO V3-Tiny network is then deployed on the embedded device after format conversion and verification.
[0013] A second aspect of the present invention provides a multi-scenario traffic target detection system based on a lightweight network, comprising:
[0014] The data acquisition module is configured to: acquire real-time traffic video information in traffic scenarios and transmit it to the embedded device;
[0015] The embedded device is configured to: perform real-time analysis and processing of traffic video data; determine whether the target is detected at the intersection based on the analysis and processing results; if the target is detected at the intersection, control the feedback and reminder module to issue a reminder; and simultaneously encapsulate the analysis and processing results and transmit them to the server via push streaming.
[0016] The embedded device is equipped with a lightweight target detection model, which is an improved YOLO V3-Tiny network. The improved YOLO V3-Tiny network is obtained by pruning and sparse training the YOLO V3-Tiny network to reduce the number of layers and channels. The improved YOLO V3-Tiny network is then deployed on the embedded device after format conversion and verification.
[0017] The above one or more technical solutions have the following beneficial effects:
[0018] (1) This invention uses lightweight technology to lightweight the deep learning algorithm, and then converts the format of the lightweight model before deploying it to the embedded device. This enables the deep learning algorithm to be more adaptable to the embedded device, solves the problem of poor integration between software and hardware, and realizes real-time detection of vehicles, pedestrians, etc. The lightweight processing includes using Tiny pruning and sparse training to process the model, reducing the number of channels and network layers of the model. This can reduce the model size while ensuring recognition accuracy, making this invention applicable to various monitoring environments and improving the applicability of the system.
[0019] (2) This invention is applicable to areas with limited visibility where accidents frequently occur. By reminding vehicles and pedestrians at intersections, it can simultaneously use both visual and auditory reminders to effectively reduce the occurrence of traffic accidents at intersections.
[0020] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 The flowchart is a multi-scenario traffic target detection method based on lightweight networks, as shown in the first embodiment.
[0023] Figure 2 This is a structural diagram of the lightweight target detection model in the first embodiment.
[0024] Figure 3 This is a flowchart of the lightweight target detection model format conversion for the first embodiment.
[0025] Figure 4 This is a control flowchart for the feedback reminder module in the first embodiment.
[0026] Figure 5This is a structural diagram of a multi-scenario traffic target detection system based on a lightweight network, as shown in the second embodiment. Detailed Implementation
[0027] Example 1
[0028] like Figure 1 As shown, this embodiment discloses a multi-scenario traffic target detection method based on lightweight networks, including:
[0029] Step 1: The data acquisition module deployed at the traffic intersection collects real-time traffic video data in the traffic scene and transmits it to the embedded device;
[0030] Step 2: The embedded device performs real-time analysis and processing of traffic video data. Based on the analysis and processing results, it determines whether the target is detected at the intersection simultaneously. If the target is detected at the intersection simultaneously, it controls the feedback reminder module to issue a reminder. At the same time, the analysis and processing results are encapsulated and transmitted to the server via push streaming.
[0031] In step 1, the data acquisition module includes a camera used to collect traffic video data. This invention primarily targets traffic scenarios with limited visibility, deploying the system on key traffic arteries with restricted visibility to monitor vehicle speed and location in real time. Through multiple feedback mechanisms, alerts are provided to vehicles, pedestrians, and other targets, effectively reducing the risk of accidents. Simultaneously, the system uploads data in real time to record vehicle violations and other actions.
[0032] In step 2, a lightweight target detection network is deployed within the embedded device. This network is an improved YOLO V3-Tiny network. The embedded device uses the RV1126 as its core processing unit and acquires data using an IMX415 camera. The data acquisition module transmits the data to the embedded device. The embedded device encapsulates the processed video data and pushes it to the server via a push stream. Pulling the stream involves retrieving information from the server at a specific address for viewing. On the software side, a traditional network with multiple lightweight processing steps is intertwined with the RV1126 using a lightweight model. The RV1126 receives the data stream from the camera, performs real-time video frame processing using its internal lightweight model and calculations, and then provides information to relevant personnel in the traffic scene through the display and audio output feedback mechanism provided by the RV1126.
[0033] This invention uses a small embedded device with the RV1126 chip as its motherboard, which features low cost, small size, low power consumption, and high computing power. This motherboard effectively reduces system development costs, is simple to deploy, and is easy to apply.
[0034] Step 2 includes: Step 201: Obtaining an improved YOLO V3-Tiny network through pruning and sparse training; this invention prunes the YOLO V3-Tiny network to achieve better and more stable compression results. This pruning method only requires cutting the convolutional layers located above the only upsampling layers, based on normal pruning. Such a pruning strategy can effectively reduce the network parameters and computational load while maintaining model performance, thereby achieving more efficient deployment and operation.
[0035] This invention uses YOLO V3-Tiny to segment detected images into multiple regions and predicts the bounding boxes and probabilities of each region. YOLO V3-Tiny employs a backbone network consisting of 7 convolutional and max-pooling layers to extract features, similar to the structure of Darknet19. The YOLO V3-Tiny detection head of this invention uses 13x13 and 26x26 resolutions for object detection.
[0036] like Figure 2 As shown, the improved YOLOv3-Tiny small model has a total of 22 layers, including five different network layers: convolutional layers (13), pooling layers (6), routing layers (2), upsampling layers (1), and output YOLO layers (2). In YOLOv3-Tiny, except for the convolutional layers before the YOLO layers, each convolutional layer is followed by a Batch Normalization (BN) layer. The convolutional layers are used to extract features, the pooling layers are used to select features extracted by the convolutional layers, the routing layers are used to extract features obtained from previous convolutions in the current layer, the upsampling layers are used to enlarge the image so that the image can be displayed at high resolution, and the output layers are used for output. This invention removes the convolutional layers above the upsampling layers, making the YOLOv3-Tiny model more lightweight.
[0037] Specifically, the YOLOv3-Tiny model undergoes sparse training, meaning it is trained from scratch without using a pre-trained model. During sparse training, a scaling factor is introduced for each channel, and this scaling factor is multiplied by the channel's output. Then, joint training is performed, simultaneously optimizing the network weights and the scaling factor.
[0038] Formula (1) includes the following elements: (x, y) represent the training input and target, and W represents the trainable weights. The first summation term corresponds to the regular training loss of a convolutional neural network (CNN), while g(·) represents the sparsity penalty on the scaling factor. The parameter λ is introduced to balance these two terms.
[0039]
[0040] To improve training speed and enhance the network's generalization ability, a normalization process is performed on the intermediate layer data of the network. Let z... in and z out These are the input and output of the Batch Normalization (BN) layer, respectively, where B represents the current mini-batch. The BN layer performs the following transformations:
[0041]
[0042] where μ B and σ B γ and β are the mean and standard deviation of the input activations computed on Batch Normalization (BN), respectively. γ and β are trainable scaling and translation parameters that provide the possibility of linearly transforming the normalized activations back to any scale. The gradient is quantized to a low bit width using the quantization method from DoReFa-Net. Note that the gradient is unbounded and may have a much larger range of values than the activations.
[0043] The gradient is then quantized to k bits using the following function, where r and k are the number of quantization bits:
[0044]
[0045] In the formula, dr is the backpropagation gradient of the output r of a certain layer, and the maximum value is taken as all axes of the gradient tensor dr, as shown in formula (4):
[0046]
[0047] Furthermore, each instance in each mini-batch has its own scaling factor. The above function first performs an affine transformation on the gradient, mapping it to [0,1], and then performs an inverse transformation after quantization. To further compensate for the potential bias caused by gradient quantization, an additional noise function N(k) is introduced, where N(k) is Equation (5), and σ is uniformly distributed in (-0.5, 0.5).
[0048]
[0049] Finally, equation (6) is used to quantize the gradient at the k-bit position, as shown below:
[0050]
[0051] It is important to note that channels with small scaling factors (scaling factors tending to 0 or equal to 0) need to be removed to eliminate redundant channels, thereby obtaining a fine-tuned and pruned network. This can help reduce unnecessary channels in the network, thus achieving model compression and acceleration.
[0052] Step 202: After format conversion and verification, the improved Tiny-YOLOv3 network is deployed on an embedded device;
[0053] In the task of training models, the system combines Figure 2 The framework shown uses the rknn-toolkit for model conversion and validation. The final model is deployed on hardware using C++ programming, requiring the pt model to be converted to a format acceptable to the RK platform. Specifically, this includes:
[0054] Step 2021: Processing the trained model
[0055] The system trains the pt model best.pt, then uses the export.py tool to convert the model to ONNX format, obtaining the best.onnx file. Next, in the Python environment, the system uses a conversion program to convert the .onnx model file to a .rknn model file. After the model conversion is complete, the system performs verification to ensure that no critical information is lost during the conversion process, guaranteeing that the converted model can run normally on hardware. This verification process is crucial because it ensures the reliability and accuracy of the converted model.
[0056] Step 2022: Hardware Porting
[0057] After verifying and ensuring the integrity of the model, the system uses a cross-compilation method under the Ubuntu operating environment to link the model file ending with the .rknn suffix with the required demonstration programs and necessary software packages to generate the final executable file. This executable file is then embedded into the RV1126 embedded processor, which has been pre-flashed with firmware. By accessing the RV1126's internal system, the system can directly run this executable program, achieving seamless execution of the model and its demonstrations.
[0058] Step 2023: Debugging and Verification
[0059] To simplify the testing process, the system employs a push-pull streaming method for real-time debugging, which offers advantages such as strong compatibility and ease of operation. Specifically, using the RTSP protocol, real-time video streams from the local device or camera are pushed, transmitting video frames to a pre-configured embedded device. Subsequently, by pulling the stream, the processed real-time video can be viewed. The advantage of this debugging method lies in its significant reduction of testing complexity. By using RTSP push-pull streaming, the system can directly simulate and observe the processed real-time video effects in a real embedded environment without involving complex data transmission and file processing.
[0060] Step 203: The embedded device analyzes and processes the collected video data, detects target categories such as pedestrians and vehicles, and then controls the feedback reminder module to issue reminders. By feeding back the data obtained after real-time processing to pedestrians and vehicles in different traffic environments, and using multiple feedback methods to remind traffic personnel, the safety of the traffic environment is improved.
[0061] The feedback and reminder module can be a display screen, a voice alarm, or an indicator light. In this embodiment, the traffic scenario selected is a T-shaped intersection. A T-shaped intersection includes a main lane and a merging lane. Generally, national highways, provincial highways, and other expressways usually have guardrails placed between the two lanes. Therefore, vehicles traveling in the merging lane can only merge into the main lane when merging into the main lane. Therefore, a display module facing the direction of travel of the main lane can be placed on the side of the main lane connecting to the merging lane at the T-shaped intersection, and a display module can be placed in the merging lane.
[0062] like Figure 4 As shown, initially, all display modules default to displaying yellow text such as "Intersection ahead, please be careful," and the indicator lights on the display modules default to green. The embedded device analyzes and processes the collected video. When a detected target appears simultaneously in the main lane and the merging lane, the embedded device controls the indicator lights on the display modules to turn red and flash, and the display modules display "Vehicles / pedestrians are crossing the intersection ahead, please be careful." When the target objects in both the main lane and the merging lane have passed the intersection, the feedback reminder module returns to the default state.
[0063] Furthermore, a voice alarm module and a light sensor can be added. Both the voice alarm module and the light sensor are connected to an embedded device. When a target is detected simultaneously in the main lane and the merging lane, and the light sensor detects that the traffic environment is daytime, the embedded device controls the voice alarm module to issue a voice reminder: "Vehicle / pedestrian is approaching ahead, please be careful." When the light sensor detects that the traffic environment is nighttime, the voice alarm module is turned off to prevent nighttime noise pollution. In this embodiment, a display screen is used for visual feedback, and the displayed content can be changed according to different situations, providing different reminders for different situations, effectively reducing noise pollution and the probability of intersection accidents.
[0064] Step 204: The embedded device encapsulates the analysis and processing results and transmits them to the server via push streaming. The server parses the encapsulated data and stores it in the cloud database. Relevant traffic management personnel can use network devices to retrieve and display the data stored on the server via pull streaming. The front-end interface can monitor traffic data in real time and display it visually, and monitor changes in traffic data in real time.
[0065] This invention utilizes Internet of Things (IoT) technology to transmit collected vehicle and pedestrian data to a cloud server and store it in a cloud-based MySQL database. This allows traffic personnel to view the monitoring data in real time via network-based devices, and also facilitates subsequent data processing, such as data analysis, providing data support for optimizing the monitored traffic environment. Furthermore, the visualization design method provided by this invention enables the integrated, real-time, and aesthetically pleasing display of all data.
[0066] The real-time display of embedded device online / offline status in this invention uses a publish-subscribe pattern, which offers advantages such as high scalability, loose coupling, and simple communication. In the publish-subscribe pattern, the message sender (publisher) and message receiver (subscriber) are decoupled, meaning they do not need to directly contact or be aware of each other's existence.
[0067] The core idea of the publish / subscribe pattern lies in introducing an intermediary role called a broker, which is responsible for all message routing and distribution. In this pattern, publishers send messages with specific topics to the broker, while subscribers receive messages of interest by subscribing to specific topics from the broker. The broker plays a crucial role, effectively coordinating and managing message delivery. In this invention, the specific topic $SYS / provided by EMQ is used to implement the functionality. By subscribing to these topics, the server can obtain device status information; for example, after a device connects to the server, it will push a status message like $SYS / **** / connected. This subscription method enables the acquisition of device online / offline information.
[0068] This invention uses the Hisilicon 3861 as a Wi-Fi module, combined with the RV1126, as a message sending source. During the system debugging and development simulation phase, the MQTT graphical client MQTT X is used for message transmission. Its intuitive UI presents a chat interface, simplifying the page operation process. The system can more easily perform rapid testing of MQTT / MQTTS connections, as well as subscribe to and publish MQTT messages.
[0069] The embedded device uses 4G, 5G, or Wi-Fi modules to upload data to the server. The client retrieves and displays the data stored on the server via streaming. For ease of management, the client also provides a built-in dashboard interface for relevant management operations. The client's main display interface includes various aspects such as vehicle flow, pedestrian flow, violation vehicle records, and warning counts, integrating each data point and clearly presenting it to traffic management personnel through graphs. The client also performs integrated data analysis to determine whether the road conditions require the system and what functions the intersection needs, providing feedback to the system for functional modifications.
[0070] Example 2
[0071] This embodiment discloses a multi-scenario traffic target detection system based on lightweight networks, including:
[0072] The data acquisition module is configured to: acquire real-time traffic video information in traffic scenarios and transmit it to the embedded device;
[0073] The embedded device is configured to: perform real-time analysis and processing of traffic video data; determine whether the target is detected at the intersection based on the analysis and processing results; if the target is detected at the intersection, control the feedback and reminder module to issue a reminder; and simultaneously encapsulate the analysis and processing results and transmit them to the server via push streaming.
[0074] The embedded device is equipped with a lightweight target detection model, which is an improved YOLO V3-Tiny network. The improved YOLO V3-Tiny network is obtained by pruning and sparse training the YOLO V3-Tiny network to reduce the number of layers and channels. The improved YOLO V3-Tiny network is then deployed on the embedded device after format conversion and verification.
[0075] The server is configured to parse the encapsulated data and store it in a cloud database.
[0076] The feedback and alert module is configured to alert vehicles and pedestrians at intersections. After data analysis and processing, feedback processing is incorporated to provide real-time updates to pedestrians and vehicles in different traffic environments. Multiple feedback methods are used to alert traffic personnel, thereby improving traffic safety.
[0077] The client is configured to retrieve and display data stored on the server via streaming. Traffic management personnel can monitor and visualize traffic data in real time through network-based devices on the front-end interface, tracking changes in traffic data, including pedestrian and vehicle flow.
[0078] like Figure 5As shown, this application's system employs a front-end / back-end separation strategy and utilizes an MVVM architecture, where the View and Model interact through the ViewModel. In the MVVM pattern, the View and Model do not have a direct connection; instead, data is transferred through the ViewModel. The interaction between the Model and ViewModel is bidirectional, meaning that changes in data within the View are synchronized to the Model, and changes in data within the Model also affect the View's presentation in real time. Through bidirectional data binding, the ViewModel effectively connects the View and Model layers, enabling synchronization between the View and Model without manual intervention.
[0079] Therefore, developers can focus on business logic, avoiding direct DOM manipulation and the need to worry about data state synchronization. MVVM handles the unified management of complex data state maintenance, making pages more concise and efficient. Leveraging Vue's two-way data binding and technological maturity, the system implements hot reloading functionality. This allows changes to page code to be reflected immediately in the browser without refreshing, facilitating faster debugging and optimization.
[0080] The client application incorporates a data monitoring module that integrates acquired data to monitor vehicle and pedestrian traffic at intersections. Detailed data analysis allows for the determination of whether the monitoring system is necessary and whether its functionality needs adjustment and optimization. First, the system gathers information from multiple data sources, including traffic flow and pedestrian crossing patterns. By integrating this heterogeneous data, the system constructs a comprehensive and accurate picture of the intersection's traffic conditions. Second, by analyzing vehicle and pedestrian data, this technology can provide substantial support for road safety and traffic planning. In-depth analysis of vehicle flow and pedestrian crossing patterns identifies potential safety hazards, providing recommendations for necessary preventative measures. Based on the conclusions drawn from data analysis, the system can decide whether to install the monitoring system at the intersection. If the data indicates traffic congestion, frequent accidents, or pedestrian difficulties, then installing the monitoring system may be a viable solution. However, if the data shows that the current traffic conditions are good, the system may deem the monitoring system unnecessary.
[0081] Ultimately, the data analysis results can guide modifications to the system's functionality. By analyzing the data provided by the monitoring module, the system can detect potential functional limitations or deficiencies. In summary, the data monitoring module provides an effective means of integrating and analyzing data to monitor vehicle and pedestrian data at intersections. Through in-depth data analysis, this module can provide a basis for decision-making, support intersection traffic safety and planning, and stimulate continuous system improvement to meet evolving needs.
[0082] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0083] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A multi-scenario traffic target detection method based on lightweight networks, characterized in that, include: Real-time acquisition of traffic condition video data in traffic scenarios and transmission to embedded devices; The embedded device performs real-time analysis and processing of traffic video data. Based on the analysis and processing results, it determines whether the target is detected at the intersection at the same time. If the target is detected at the same time, it controls the feedback and reminder module to issue a reminder. At the same time, the analysis and processing results are encapsulated and transmitted to the server via push streaming. The embedded device is equipped with a lightweight target detection model, which is an improved YOLOV3-Tiny network. The improved YOLOV3-Tiny network is obtained by pruning and sparse training the YOLO V3-Tiny network to reduce the number of layers and channels. The improved YOLO V3-Tiny network is then deployed on the embedded device after format conversion and verification. The pruning and sparse training of the YOLO V3-Tiny network includes: during sparse training, introducing a scaling factor for each channel in the BN layer and multiplying the scaling factor by the channel output; Then, joint training is performed, while optimizing the network weights and scaling factors, removing redundant channels and convolutional layers, to obtain the improved YOLO V3-Tiny network. The improved YOLO V3-Tiny network consists of five different network layers: convolutional layer, pooling layer, routing layer, upsampling layer, and output layer; The convolutional layer is used to extract features, the pooling layer is used to select features extracted from the convolutional layer, the routing layer is used to extract features obtained from previous convolutions in the current layer, the upsampling layer is used to enlarge the image, and the output layer is used for output.
2. The multi-scenario traffic target detection method based on lightweight networks as described in claim 1, characterized in that, The process of converting the improved YOLO V3-Tiny network to a new format includes: After sparsely training the improved YOLO V 3-Tiny network, the best.pt file is obtained. The export.py tool is used to convert the network into the .onnx format to obtain the best.onnx model file. Convert the best.onnx model file to a rknn model file in a Python environment.
3. The multi-scenario traffic target detection method based on lightweight networks as described in claim 2, characterized in that, After verifying the network after format conversion, the cross-compilation method is used in the Ubuntu operating environment to link the RKNN model file with the required demonstration programs and software packages to generate the final executable file. The executable file is then embedded into the RV1126 embedded processor with pre-programmed firmware in the embedded device to complete the network deployment.
4. The multi-scenario traffic target detection method based on lightweight networks as described in claim 1, characterized in that, The feedback reminder module includes a display module, which is set on one side of the main road connecting to the merging lane at the T-junction, and on the other side of the merging lane. When the embedded device detects a target appearing simultaneously in both lanes of a T-junction, the indicator light on the control display module turns red and flashes, and the display module simultaneously displays "Vehicle / pedestrian ahead, please be careful"; Once all detected targets in the lane have passed the intersection, the feedback and reminder module returns to its default state.
5. The multi-scenario traffic target detection method based on lightweight networks as described in claim 1, characterized in that, The server parses and stores the encapsulated data in a cloud database, and the client retrieves and displays the data stored on the server by pulling the stream.
6. A multi-scenario traffic target detection system based on lightweight networks, characterized in that, include: The data acquisition module is configured to: acquire real-time traffic video information in traffic scenarios and transmit it to the embedded device; The embedded device is configured to perform real-time analysis and processing of traffic video data, determine whether the target is detected at the intersection based on the analysis and processing results, and if the target is detected at the intersection, control the feedback reminder module to issue a reminder. Simultaneously, the analysis and processing results are packaged and transmitted to the server via push streaming. The embedded device is equipped with a lightweight target detection model, which is an improved YOLOV3-Tiny network. The improved YOLOV3-Tiny network is obtained by pruning and sparse training the YOLO V3-Tiny network to reduce the number of layers and channels. The improved YOLO V3-Tiny network is then deployed on the embedded device after format conversion and verification. The pruning and sparse training of the YOLO V3-Tiny network includes: during sparse training, introducing a scaling factor for each channel in the BN layer and multiplying the scaling factor with the channel output; then, performing joint training, simultaneously optimizing the network weights and scaling factor, removing redundant channels and convolutional layers, to obtain an improved YOLO V3-Tiny network. The improved YOLO V3-Tiny network consists of five different network layers: convolutional layer, pooling layer, routing layer, upsampling layer, and output layer; The convolutional layer is used to extract features, the pooling layer is used to select features extracted by the convolutional layer, the routing layer is used to draw out the feature layers obtained by the previous convolution in the current layer, the upsampling layer is used to enlarge the image, and the output layer is used to output the image. The server is configured to parse the encapsulated data and store it in a cloud database. The feedback and reminder module is configured to provide reminders for the detected targets at the intersection.
7. The multi-scenario traffic target detection system based on lightweight networks as described in claim 6, characterized in that, It also includes a client, which retrieves and displays data stored on the server via a streaming method.
8. The multi-scenario traffic target detection system based on lightweight networks as described in claim 6, characterized in that, The YOLO V3-Tiny network is pruned and sparsely trained, including: during sparse training, a scaling factor is introduced for each channel in the BN layer, and the scaling factor is multiplied by the channel output; then, joint training is performed to optimize the network weights and scaling factor, remove redundant channels and convolutional layers, and obtain an improved YOLO V3-Tiny network.
Citation Information
Patent Citations
Lightweight small target detection method based on improved YOLOv7
CN116206185A
Vehicle-mounted video target detection method based on deep learning
WO2020181685A1
Cited By
Method for detecting objects in car window based on deep learning
CN122223665A