Multi-channel video big data real-time analysis method and system based on cloud edge-end cooperation

By employing a cloud-edge-device collaborative multi-channel video big data analysis method in the rail train monitoring system, utilizing edge computing devices for real-time video stream decoding and target detection, and combining deep learning algorithms with real-time transmission from a Kafka server, the latency and network pressure issues in multi-channel video data analysis were resolved, achieving efficient video data processing and improved real-time performance.

CN117372929BActive Publication Date: 2026-08-04SHANDONG JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG JIAOTONG UNIV
Filing Date
2023-10-23
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing rail train monitoring systems suffer from high latency in video analysis and processing and heavy network aggregation pressure in multi-channel video data analysis, failing to meet real-time and efficiency requirements.

Method used

A real-time analysis method for multi-channel video big data based on cloud-edge-device collaboration is adopted. Edge computing devices are used to decode and detect targets in multi-channel video streams, and deep learning algorithms are combined to perform image classification and text recognition. The analysis results are sent to the client terminal in real time through a Kafka server.

Benefits of technology

While reducing the computing pressure on the cloud service center, it improves the real-time performance and efficiency of video data processing, reduces network latency, and enhances the effective utilization rate of video data applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117372929B_ABST
    Figure CN117372929B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of video big data intelligent analysis and application, in particular to a multi-channel video big data real-time analysis method and system based on cloud edge-end cooperation, the method comprising: acquiring multi-channel video stream of a rail train; based on an edge computing device, automatically decoding the multi-channel video stream, after target detection and classification of the decoded multi-channel video stream, encapsulating the target detection result and video frame data into a json file, and sending the json file as a message to a kafka server in real time; after the kafka server processes the json file sent by the edge computing device, sending the json file to a customer terminal for display. Based on the secondary development of Gsteamer and Deepstream, the present application reasonably encodes and decodes video stream big data, and the present application summarizes the collected multi-channel video to the edge computing device for one-time target detection, thereby improving the real-time performance of video data processing under the condition of reducing the computing pressure of the cloud service center.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent analysis technology and application of video big data, and in particular relates to a method and system for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous development of artificial intelligence, big data analytics, and the maturity of multimedia technologies, the amount of video data generated by cameras is growing rapidly. However, processing and analyzing such large amounts of video data remains a significant challenge. Traditional manual analysis methods are inefficient and lack precision, failing to meet the demands of today's large-scale, multi-channel video data analysis. Therefore, how to apply modern artificial intelligence, big data analytics, and deep learning algorithms to improve the analysis and processing capabilities of video big data has become a hot research topic.

[0004] Multi-channel video big data analytics involves real-time analysis and processing of multiple sets of video data. It organizes and manages video big data according to rules to facilitate subsequent data analysis and application, and is a crucial step in video big data processing. By introducing streaming media, artificial intelligence, and big data analytics methods into multi-channel video streams, more rational and structured real-time processing of multi-channel video stream data is achieved, improving the usability and application value of video big data.

[0005] Existing rail train monitoring systems essentially employ multi-channel video big data analysis methods to analyze real-time video data. Typically, multiple surveillance cameras are installed along the track, and the data collected by these cameras is transmitted in real time to a processing center for analysis and processing. The results are then displayed at the monitoring center. However, existing rail train monitoring methods still have significant technical shortcomings. For example, aggregating multiple video streams into a single terminal for analysis and processing results in high latency in video analysis and processing, causing excessive pressure on the network aggregation. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, the present invention provides a method and system for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] The first aspect of this invention provides a method for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration, comprising:

[0009] Acquire multi-channel video streams of rail trains;

[0010] The edge computing device automatically decodes multi-channel video streams, performs target detection and classification on the decoded multi-channel video streams, and encapsulates the target detection results and video frame data into a JSON file, which is then sent to the Kafka server as a message in real time.

[0011] The Kafka server processes the JSON files sent by the edge computing devices and then sends them to the client terminal for display.

[0012] A second aspect of the present invention provides a real-time analysis system for multi-channel video big data based on cloud-edge-device collaboration, comprising:

[0013] A video acquisition system is used to acquire multi-channel video streams from a rail train.

[0014] An edge computing device is connected to a video surveillance system. The edge computing device automatically decodes multi-channel video streams, performs target detection and classification on the decoded multi-channel video streams, and encapsulates the target detection results and video frame data into a JSON file, which is then sent to the Kafka server as a message in real time.

[0015] The Kafka server processes the JSON files sent by the edge computing device and then sends them to the client terminal for display.

[0016] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in the first aspect of the present invention.

[0017] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in the first aspect of the present invention.

[0018] The above one or more technical solutions have the following beneficial effects:

[0019] (1) This invention studies a real-time analysis method for multi-channel video big data under cloud-edge-device collaboration. Based on secondary development of Gsteamer and Deepstream, it performs reasonable encoding and decoding of video stream big data, adopts deep learning object detection algorithms in artificial intelligence, and selects appropriate regions of interest (ROIs) to achieve functions such as image classification, text recognition, contour extraction, and comparison. This invention aggregates multiple video streams to an edge computing device for object detection, improving the real-time performance of video data processing while reducing the computational pressure on the cloud service center.

[0020] (2) This invention combines multi-threaded resource allocation and other mechanisms to perform real-time structured analysis on multi-channel camera video streams. It utilizes NVIDIA's GPU acceleration engine to accelerate neural network inference, and enables multi-threaded communication between multiple monitoring devices and edge computing devices. Each edge device simultaneously receives and processes 16 camera video stream data. The inference results (JSON format) and the corresponding analysis effect images (base64 format) are sent to the cloud server in real time through Kafka in big data technology to achieve multi-message concurrency, thereby improving the effective utilization and real-time performance of the application.

[0021] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0023] Figure 1 The flowchart is a method for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration, as shown in the first embodiment.

[0024] Figure 2 The deepstream workflow diagram for the first instance.

[0025] Figure 3 This is a diagram of the target detection model structure for the first instance.

[0026] Figure 4 The diagram shows the attention mechanism model for the first instance.

[0027] Figure 5 This is a screenshot of the OCR text recognition result for the first example.

[0028] Figure 6 This is a flowchart of the OCR text recognition technology in the first embodiment.

[0029] Figure 7The image classification result for the first instance is shown. Detailed Implementation

[0030] Example 1

[0031] like Figure 1 As shown, this embodiment discloses a real-time analysis method for multi-channel video big data based on cloud-edge-device collaboration, including:

[0032] Step 1: Configure the required hardware and software environment for the system;

[0033] Step 2: Acquire multi-channel video streams of the rail train;

[0034] Step 3: Automatically decode multi-channel video streams based on edge computing devices. After target detection and classification of the decoded multi-channel video streams, encapsulate the target detection results and video frame data into a JSON file and send it as a message to the Kafka server in real time.

[0035] Step 4: The Kafka server interacts with the client terminal. The Kafka server processes the files sent by the edge computing device and then sends them to the client terminal for display.

[0036] In step 1, the required hardware and software environment of the system is configured, as shown in Table 1. Edge devices are based on... The Jetson Xavier NX series modules utilize the Jetpack software package. Based on specific requirements, Ubuntu 18.04.5 and Jetpack 4.6.1 were chosen. JetPack 4.6.1 is a version of NVIDIA Jetson embedded systems, containing a suite of software and tools to support the development and deployment of deep learning and computer vision applications. This includes deep learning acceleration libraries such as CUDA (Compute Unified Device Architecture), cuDNN, and TensorRT, as well as computer vision libraries such as OpenCV and VisionWorks.

[0037] Table 1

[0038]

[0039]

[0040] In step 2, multiple cameras are installed along the track to collect video data from the train's operating environment. By connecting the cameras to an edge computing device, the edge computing device captures the streams from multiple cameras in real time to obtain the original RTSP video stream.

[0041] The edge computing device includes a GStreamer component, an NVIDIA hardware decoding layer, an NVIDIA encoding layer, an NVStream container, and an object detection and classification module; the object detection and classification module includes an improved YOLO v5 network, a text recognition module, and an image classification module;

[0042] Step 3 includes: Step 301: Decoding the multi-channel video stream based on the edge computing device to obtain NV12 format video frame data, and converting the color channels of the NV12 format video frame data from YUV format to RGB format; specifically including:

[0043] Step 3011: Read the RTSP video stream from the multi-channel camera through the h264parse layer in the gstreamer component to obtain h.264 format video data V1;

[0044] Step 3012: Input V1 into the NVIDIA hardware decoding layer and obtain NV12 format video frame data V2 through GPU-accelerated decoding;

[0045] Step 3013: Convert the color channels of V2 from YUV format to RGB format, and aggregate the frame data of the multi-channel video stream at the same time into the nvstream container;

[0046] Among them, multi-channel video stream refers to video stream data collected by multiple cameras. One camera corresponds to one channel of video stream, and multiple cameras correspond to multiple channels of video stream.

[0047] A target detection is performed on the decoded multi-channel video stream based on the improved YOLOv5 network, specifically including:

[0048] Step 302: After aggregating the data into the nvstream container, a daemon process based on the sh script and Crontab timer inputs the color-channel converted NV12 format video frame data into the improved YOLO v5 network, obtaining a single object detection result for the current frame; specifically:

[0049] The video stream is fed into the YOLO v5 object detection layer by skipping frames, and the TensorRT inference engine is used to accelerate inference to obtain target information such as vehicles and pedestrians.

[0050] This invention selects YOLO v5 as the basic object detection model, such as Figure 3 Based on the characteristics of the target dataset in the actual scenario, some improvements were made to YOLO v5, mainly including parameter adjustment and network structure optimization. The adjusted parameters are shown in Table 2.

[0051] Table 2

[0052]

[0053]

[0054] The YOLO v5 neural network has an input size of [640, 640, 3]. The self-trained torch format network model needs to be converted into a weight file in wts format to further adapt to edge devices with Arm architecture. FP16 half-precision inference is selected to improve the efficiency of video structured analysis and increase the maximum number of video streams that the device can analyze in real time.

[0055] The network weight file in WTS format is converted into a network model in Engine format, and TensorRT is used to accelerate inference, enabling real-time object detection on multi-channel video frame data. Furthermore, based on this network structure, an SENet (Squeeze-and-Excitation Network) attention mechanism is added to the convolutional layers. This mechanism can prioritize different features of the target, increasing the weight of important features and suppressing unimportant features, such as... Figure 4 As shown.

[0056] Step 303: After performing a first target detection on the multi-channel video stream using the improved YOLO v5 neural network, a second target detection is performed on the multi-channel video stream containing vehicle targets based on OCR text recognition and image classification algorithms. Finally, the NVIDIA encoding layer encapsulates the two target detection results and the corresponding video frames into a JSON file; specifically including:

[0057] Step 3031: Perform OCR text recognition on the vehicle containing license plate information to obtain the vehicle's license plate number information. The result is as follows: Figure 5 As shown.

[0058] Step 3032: Analyze the status of pantograph, disconnector, etc. using image classification technology;

[0059] OCR text recognition is performed on vehicles containing license plate information to obtain the vehicle's license plate number. Image classification technology is then used to analyze the pantograph's status, etc. The OCR processing flow is as follows: Figure 6 As shown. To determine whether the target detection results include targets requiring secondary detection, such as license plates and disconnect switches, OCR technology is used to identify license plate numbers in the scene, and image classification technology is used to classify the status of disconnect switches awaiting detection. The results are as follows. Figure 7 As shown.

[0060] Step 3033: Encode the video frames in JPG format using Base64, and encapsulate them together with the relevant information obtained from the analysis in steps 3031 and 3032 into a JSON message for real-time transmission.

[0061] The video frame images are encoded from JPEG format to base64 format, and the channels are converted from RGBA four channels to RGB three channels.

[0062] The JSON message contains the camera ID, time, target detection result, license plate number, status of the target to be detected, and image base64 encoding, etc. The message content is encapsulated in the following format:

[0063]

[0064] Here, `object` represents the bounding box coordinates and object category of the detected target. `Image` is the base64 encoded image of the relevant image.

[0065] Step 4: Send the JSON message to the specified Kafka server for the front end to consume.

[0066] This invention utilizes Kafka technology to receive and send analysis results in real time through multiple channels, ensuring the real-time performance and stability of the overall system for big data.

[0067] Step 5: Design of system daemon process and automatic restart upon startup.

[0068] This invention employs a daemon process implementation scheme based on sh scripts and Crontab timers. The specific implementation scheme includes the following steps:

[0069] (1) Write a sh script to control the execution of the target process. The script contains code to detect whether the target process is running, as well as code to start, stop and restart the target process.

[0070] (2) Use the Crontab timer to set up a scheduled task so that the sh script can be run periodically.

[0071] The above scheme implements a daemon process, which can control the normal operation of the target process and restart it when needed, ensuring system stability and efficiency. The existence of the process is checked using `pgrep -f`; if the count is greater than 0, the program is running. If the process does not exist, it means OCR has not started, and the startup command is re-executed. Continuous, timed monitoring ensures the program runs normally and allows for recovery in case of emergencies or power outages.

[0072] The Sh script configuration is as follows:

[0073]

[0074] The Crontab timer configuration is as follows:

[0075]

[0076] Example 2

[0077] This embodiment discloses a multi-channel video big data real-time analysis system based on cloud-edge-device collaboration, including:

[0078] Video acquisition system, used to acquire multi-channel video streams of rail trains;

[0079] Edge computing devices connect to video surveillance systems and automatically decode multi-channel video streams. After performing target detection and classification on the decoded multi-channel video streams, the target detection results and video frame data are encapsulated into JSON files and sent as messages to the Kafka server in real time.

[0080] The Kafka server processes JSON files sent by edge computing devices and then sends them to the client terminal for display.

[0081] Example 3

[0082] The purpose of this embodiment is to provide a computer-readable storage medium.

[0083] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in Embodiment 1 of this disclosure.

[0084] Example 4

[0085] The purpose of this embodiment is to provide an electronic device.

[0086] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in a multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in Embodiment 1 of this disclosure.

[0087] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0088] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0089] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A real-time analysis method for multi-channel video big data based on cloud-edge-device collaboration, characterized in that, include: Acquire multi-channel video streams of rail trains; The edge computing device automatically decodes multi-channel video streams, performs target detection and classification on the decoded multi-channel video streams, and encapsulates the target detection results and video frame data into a JSON file, which is then sent to the Kafka server as a message in real time. The Kafka server processes the JSON files sent by the edge computing devices and then sends them to the client terminal for display. The edge computing device also includes a target detection and classification module, which includes: an improved YOLO v5 network, an OCR text recognition module, and an image classification module; A first target detection is performed on the multi-channel video stream based on the improved YOLO v5 network, and a second target detection is performed on the multi-channel video stream containing vehicle targets based on the OCR text recognition module and the image classification module. The method for vehicle target detection based on the decoded multi-channel video stream using the improved YOLOv5 network includes: inputting the video stream frame-by-frame into the target detection layer of the improved YOLOv5 network, accelerating inference using the TensorRT inference engine, using an attention mechanism to classify the importance of different features of the target, increasing the weight of important features, suppressing unimportant features, and obtaining vehicle targets and pedestrian targets. The secondary target detection of multi-channel video streams containing vehicle targets, based on the OCR text recognition module and image classification module, includes: The detected vehicle targets are subjected to OCR text recognition to obtain the vehicle license plate information; Image classification algorithms are used to obtain the status information of the disconnect switch and pantograph of the vehicle target.

2. The method for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration as described in claim 1, characterized in that, The edge computing device includes a GStreamer component, an NVIDIA hardware decoding layer, and an NVStream container; The automatic decoding of multi-channel video streams based on the edge computing device includes: The h.264parse layer in the gstreamer component is used to read the RTSP video stream from the multi-channel camera and obtain the h.264 format video data V1. The video data V1 is input into the NVIDIA hardware decoding layer, and the NV12 format video frame data V2 is obtained through GPU-accelerated decoding. The color channels of video frame data V2 are converted from YUV format to RGB format, and the frame data of multiple video streams at the same time are aggregated into the nvstream container.

3. The method for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration as described in claim 1, characterized in that, The target process is controlled by a sh script. The sh script contains code to detect whether the target process is running, as well as code to start, stop, and restart the target process. Use the Crontab timer to set up a scheduled task to run the sh script periodically.

4. The method for real-time analysis of multi-channel video big data based on cloud-edge-device collaboration as described in claim 1, characterized in that, The JSON file includes: camera ID, time, target detection results, and image base64 encoding.

5. A real-time analysis system for multi-channel video big data based on cloud-edge-device collaboration, characterized in that, include: A video acquisition system is used to acquire multi-channel video streams from a rail train. An edge computing device is connected to a video surveillance system. The edge computing device automatically decodes multi-channel video streams, performs target detection and classification on the decoded multi-channel video streams, and encapsulates the target detection results and video frame data into a JSON file, which is then sent to the Kafka server as a message in real time. The edge computing device also includes a target detection and classification module, which includes: an improved YOLO v5 network, an OCR text recognition module, and an image classification module; A first target detection is performed on the multi-channel video stream based on the improved YOLO v5 network, and a second target detection is performed on the multi-channel video stream containing vehicle targets based on the OCR text recognition module and the image classification module. The method for vehicle target detection based on the decoded multi-channel video stream using the improved YOLOv5 network includes: inputting the video stream frame-by-frame into the target detection layer of the improved YOLOv5 network, accelerating inference using the TensorRT inference engine, using an attention mechanism to classify the importance of different features of the target, increasing the weight of important features, suppressing unimportant features, and obtaining vehicle targets and pedestrian targets. The secondary target detection of multi-channel video streams containing vehicle targets, based on the OCR text recognition module and image classification module, includes: The detected vehicle targets are subjected to OCR text recognition to obtain the vehicle license plate information; Image classification algorithms are used to obtain the status information of the disconnect switch and pantograph of the vehicle target; The Kafka server processes the JSON files sent by the edge computing device and then sends them to the client terminal for display.

6. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in any one of claims 1-4.

7. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multi-channel video big data real-time analysis method based on cloud-edge-device collaboration as described in any one of claims 1-4.