An edge-side monitoring video structured storage method based on Ascend processors

By using domestic Ascend processors in the edge-end hardware system, key information in surveillance videos is extracted and stored, and the problems of high concurrency demand and low retrieval efficiency of traditional video surveillance data centers are solved, and more efficient surveillance video processing and storage are achieved.

CN114821411BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210399070.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-06-13
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Traditional video surveillance data centers have high concurrency demand, high data processing pressure, and low retrieval efficiency, so they cannot effectively process massive urban monitoring data.

Method used

The edge hardware system based on the domestic Ascend processor is adopted to extract key video segments and pedestrian appearance information in the surveillance video in real time, and provide local storage and retrieval services to reduce dependence on the server.

Benefits of technology

The computing pressure on the server is significantly reduced through edge-end computing, improves the real-time processing capability and data storage efficiency of surveillance videos, and reduces network bandwidth load and privacy leakage risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821411B_ABST
    Figure CN114821411B_ABST
Patent Text Reader

Abstract

An edge-side monitoring video structured storage method based on Ascend processors, comprising: building an edge-side hardware system based on domestic Ascend processors; updating the background image in real time according to the monitoring camera video stream; performing salt-and-pepper denoising on the binary matrix Bi obtained in the previous step, finding the total number c of pixel points with a value of 1, setting a threshold Tc, if c > Tc, it indicates that there is a moving target, using the YOLOV3 algorithm for target detection, segmenting the selected pedestrian images, and sending them to the pedestrian appearance extraction network to extract appearance labels; for the target detection network, if a pedestrian target is extracted, the frame is merged into the H264 file to generate label information about the key video segments and pedestrian targets; placing the H264 in the SD card, and when the cached data reaches the storage limit of the SD card, it will be regularly backed up to the server. The present invention speeds up the real-time performance and response speed of monitoring video emergency handling, and reduces the computing pressure on the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of embedded, deep learning, database, etc., and particularly relates to a method for structured storage of edge monitoring videos based on domestic Ascend processors, which extracts key video segments, the appearance and quantity of pedestrians in the video segments at the edge, and provides local storage and retrieval services for the key video segments. Background Art

[0002] Fast-changing traffic or security events require the city's monitoring system to have highly automated and intelligent capabilities to complete tasks such as event warning and trajectory tracking. Although machine learning has developed rapidly in recent years, how to truly apply it to the specific scenario of video behavior semantic detection of a large amount of urban monitoring data is an engineering problem that needs to be broken through currently.

[0003] In urban monitoring applications, a more efficient edge computing platform has become a future research trend.

[0004] In the past, due to the limited performance of the embedded processor at the edge, the previous monitoring cameras could not run algorithms with high computing power requirements, so only simple storage was performed at the edge. If these video data were sent to the server for processing, it would occupy a lot of network bandwidth. Moreover, the more cameras there are, the higher the concurrent demand for the server, the greater the processing pressure, and there is a large delay in processing. The current solution is to strengthen the computing power of the edge nodes, that is, the camera performs real-time processing when the video is input, and the identified key pictures, video segments, and extracted semantic information are sent to the server after being stored in a structured manner. In this way, the computing power can be migrated from the centralized cloud center to the camera edge closer to the user end. Executing part or all of the calculations at the edge can reduce latency, provide real-time response, reduce network bandwidth load, weaken the risk of privacy leakage, and improve data security. The emerging deep learning algorithms have promoted the rapid development of AI chips. The Ascend 310 AI processor of Huawei HiSilicon has 2 AI cores built into its chip, supports LPDDR4x with a width of 128 bits. Different from the parallel computing of the GPU, its chip internally accelerates matrix operations by integrating ASIC hardware circuits, can achieve a computing power of 22 TOPS INT8, and can be widely used in edge inference scenarios such as intelligent monitoring and drones. Strengthening the computing power of the edge through AI chips is a current hot issue, which also symbolizes the real start of artificial intelligence from theory to application.

[0005] In summary, the structured storage device for monitoring videos deployed at the edge through domestic Ascend processors in this project is of great significance to the development of the monitoring system. Summary of the Invention

[0006] To overcome the disadvantages of high concurrent demand, high data processing pressure, and low retrieval efficiency in traditional video surveillance data centers, the present invention provides an edge - end monitoring video structured storage method based on Ascend processors. At the edge - end, key video segments and the appearance and quantity of pedestrians in the video segments are extracted, and local storage and retrieval services are provided for the key video segments.

[0007] An edge - end monitoring video structured storage method based on Ascend processors, comprising the following steps:

[0008] 1) Build an edge - end hardware system based on domestic Ascend processors, which consists of a main control chip (108), an EMMC module (101), an RTC module (102), a debugging interface (103), a MIPI interface (107), a gigabit Ethernet interface (106), a USB3.0 interface (105), and an SDHC module (104);

[0009] 2) Update the background image in real - time according to the video stream of the monitoring camera. The specific update formula for the background image is shown in Equation (1):

[0010]

[0011] Matrix B represents the background image, matrix Bi represents the binary decision matrix. If the point on this matrix is 0, the corresponding pixel point is updated; if it is 1, it is not updated. F represents the original frame at adjacent moments;

[0012] 3) Perform salt - and - pepper denoising on the binary matrix Bi obtained in the previous step, find the total number of points c with pixel point value 1, set a threshold Tc. If c > Tc, it indicates the existence of a moving target. Use the YOLOV3 algorithm for target detection, segment the selected pedestrian images, and send them to the pedestrian appearance extraction network to extract appearance labels;

[0013] 4) For the target detection network in step (3), if a pedestrian target is extracted, merge this frame into the H264 file and generate label information about the key video segment and the pedestrian target;

[0014] P = {p 1 , p 2 , …, p n} (2)

[0015] T = {t 1 , t 2 , …, t n} (3)

[0016] C = {c 1 , c 2 , …, c 35} (4)

[0017] P represents the camera IDs distributed at different locations, T represents different moments, and C represents the appearance tags of pedestrians. The storage information for the key video segments can be represented by the camera ID, start time, and end time. Among them, we change the end time to the total number of frames from the start time to the end time:

[0018]

[0019] where p i and t i represent the camera ID and start time respectively. These two can determine the file name of the stored H264 file. fps represents the frame rate when encoding the H264 file. Then represents the total number of frames stored in this key video segment, and thus the position of the key video segment vi in the H264 file can be located:

[0020]

[0021] end_fream = start_fream + f(n) (7)

[0022] f represents the number of frames in each video segment. For pedestrian information, it can be represented by feature tags and key video segments:

[0023] pedestrian_info = {c i , vi} (8) Design the table structure according to the above formulas (5) and (7) and store it in the database;

[0024] 5) Place the H264 on the SD card. When the cached data reaches the storage limit of the SD card, it will be periodically backed up to the server.

[0025] Preferably, the construction of the edge hardware system based on domestic Ascend processors specifically includes:

[0026] Main control chip, using the altas200 module of Huawei HiSilicon, which integrates the Ascend 310 processor. Utilize its powerful computing power to run logic code and deploy deep learning algorithms, and use various interfaces to drive peripherals;

[0027] EMMC module, which solidifies the startup files of the embedded platform and automatically loads the linux file system after power-on;

[0028] RTC module, which provides an accurate time reference for the embedded platform. Even after accidental power-off and system crash and restart, it still has a time reference;

[0029] Debug interface: The UART module of the main control chip uses RS232 to output the system startup information, uses LCD to display the UI interface of the embedded platform, and decides whether to turn off the LCD to maintain low-power operation through the debug button. The working indicator light shows different working states of the system;

[0030] The SDIO interface is connected to the SD card, and the SD card caches the structured data of the monitored video;

[0031] The MIPI interface and the USB3.0 interface can be used to connect different types of monitored video cameras;

[0032] The RGMII interface and the MDIO interface of the main control chip are connected to the PHY chip. The PHY chip provides a gigabit network port externally, and the gigabit network port can read the input network video stream or upload data to the server.

[0033] Preferably, the specific method for updating the background image in step (2) is as follows:

[0034] Select an initial frame to initialize the background image B(x, y, z), and subtract the background image from each pixel point on the original frames F at different adjacent times to obtain a differential image matrix:

[0035] MD(D,t)={D(x,y,z,t 1 ),D(x,y,z,t 2 ),…,D(x,y,z,t n )} (9)

[0036] The matrix MD represents the differential image matrix, D represents the differential frame obtained by subtracting the background image from each frame, and convert the differential image matrix into a grayscale image matrix:

[0037] MG(G,t)={G(x,y,t 1 ),G(x,y,t 2 ),…,G(x,y,t n )} (10)

[0038] The matrix MG represents the grayscale image matrix, G represents the grayscale frame converted from the differential frame. In order to reduce the calculation amount, the conversion formula for each pixel point value is:

[0039] V=(76R+150G+30B)>>8 (11)

[0040] V is the value of each pixel point after conversion to grayscale, R, G, and B respectively represent the pixel values on the RGB three channels. Set the threshold Tb to convert the grayscale image matrix into a binary matrix, and its conversion formula is:

[0041]

[0042] If the value of a pixel point on the binary matrix Bi is 1, it means that the point may be a moving target, and the background image does not update the point; if it is 0, the background image updates the point. 4. The structured storage method for edge surveillance video based on Ascend processor as described in claim 1 is characterized in that the appearance label described in step 3) specifically includes age, gender, and clothing color, with a total of 35 categories; the appearance label P has a total of 35.

[0043] Preferably, the appearance labels described in step 3) specifically include age, gender, clothing color, and a total of 35 categories; there are a total of 35 appearance labels P.

[0044] The edge hardware system is mainly composed of a main control chip, USB interface module, MIPI interface module, Gigabit Ethernet module, EMMC module, SD card module, RTC module, debugging interface and power module.

[0045] The main control chip needs to have a rich peripheral interface and super computing power, and it needs to run a tailored Linux system. Rich peripheral interfaces can expand various peripheral interfaces. High computing power is required to meet the deployment conditions of deep learning algorithms and run our logic code and complete software encoding and decoding of video images. Receive input of various surveillance camera data through USB interface or MIPI interface. Gigabit Ethernet provides network communication, which can send data to the server and receive data from network video streams. EMMC can solidify the operating system on the board to ensure that even if there are uncontrollable factors such as power failure, the system can still run automatically when it is powered off and restarted again. SD card can store structured data and can be expanded according to the needs of different camera scenarios. RTC module is used to provide real-time clock for the system. Storing structured data will use database or timestamp, and this information needs to use the onboard real-time clock. The debugging interface can not only display the working status of the system in real time, but also locate the problem in the first time when the system crashes. The power module supplies power to the main control chip and various peripheral modules.

[0046] The edge hardware system obtains the camera video input through the USB interface, processes the camera input data through the V4L2 library of the Linux system and converts it into YUV image frames. In order to maintain the low power operation of the system, the deep learning algorithm cannot be used for real-time video detection. It is necessary to use the traditional digital image processing method to detect moving targets. The main steps are as follows:

[0047] Select an initial frame to initialize the background image B(x, y, z), and subtract the background image from each pixel on the original frame F at different adjacent moments to obtain the differential image matrix:

[0048] MD(D,t) = {D(x,y,z,t 1 ), D(x,y,z,t 2 ), …, D(x,y,z,t n )} (9)

[0049] The matrix MD represents the difference image matrix, D represents the difference frame obtained by subtracting the background image from each frame, and the difference image matrix is converted into a grayscale image matrix:

[0050] MG(G,t) = {G(x,y,t 1 ), G(x,y,t 2 ), …, G(x,y,t n )} (10)

[0051] The matrix MG represents the grayscale image matrix, G represents the grayscale frame converted from the difference frame. To reduce the computational complexity, the conversion formula for each pixel value is:

[0052] V = (76R + 150G + 30B) >> 8 (11)

[0053] V is the value of each pixel after conversion to grayscale. R, G, and B represent the pixel values on the RGB three channels respectively. Set the threshold Tb to convert the grayscale image matrix into a binary matrix, and its conversion formula is:

[0054]

[0055] If the pixel value of the binary matrix Bi is 1, it means that the point may be a moving target, and the background image does not update this point. If it is 0, the background image updates this point. Then the update formula for the background image is:

[0056]

[0057] The matrix B represents the background image, the matrix Bi represents the binary decision matrix. If the point on this matrix is 0, update the corresponding pixel point; if it is 1, do not update. F represents the original frame at adjacent times.

[0058] Perform salt-and-pepper noise reduction on the binary matrix Bi obtained in the previous step, find the total number c of pixel points with a value of 1, set the threshold Tc. If c > Tc, it means that there is a moving target. Use the YOLOV3 algorithm for target detection, segment the selected pedestrian images, and send them to the pedestrian appearance extraction network to extract appearance labels. If a pedestrian target is detected, merge this frame into the H264 file and generate label information about the key video segment and the pedestrian target.

[0059] P = {p 1 , p 2 , …, pn} (2)

[0060] T = {t 1 , t 2 , …, t n} (3)

[0061] C = {c 1 , c 2 , …, c 35} (4)

[0062] P represents the camera IDs distributed at different locations, T represents different moments, and C represents the appearance labels of pedestrians, with a total of 35. The storage information for the key video segments can be represented by the camera ID, start time, and end time. Among them, we change the end time to the total number of frames from the start time to the end time:

[0063]

[0064] Among them, p i and t i represent the camera ID and start time respectively. These two can determine the file name of the stored H264 file. fps represents the frame rate when encoding the H264 file. Then represents the total number of frames stored in this key video segment, and the position of the key video segment vi in the H264 file can be located:

[0065]

[0066] end_fream = start_fream + f(n) (7)

[0067] f represents the number of frames in each video segment. For pedestrian information, it can be represented by feature labels and key video segments:

[0068] pedestrian_info = {c i , vi} (8)

[0069] Design the table structure according to the above formulas (5) and (7) and store it in the database.

[0070] Place the H264 on the SD card. When the cached data reaches the storage limit of the SD card, it will be backed up to the server regularly.

[0071] The server in the data center serves as the server end. When the monitoring video structured storage device installed at the edge end uploads structured data at regular intervals, it will first initiate a connection request to the server end as the client end. The server end needs to receive a large amount of data uploaded by the edge devices. Limited by network bandwidth and processing efficiency, the number of connections at the same time is limited. After the server end replies that the connection is allowed, the client end will then upload the structured data. The H264 files are stored in the distributed file system, while the structured data such as the tag information of the key video ends is stored in the distributed database. At the same time, the server also provides a Web retrieval service. By accessing the Web page through the browser on any client machine, it supports multi-dimensional input of time, camera id, and person characteristics for retrieval. While retrieving locally, the server sends the retrieved person to each edge device. After receiving the retrieval task, each edge device will also perform local retrieval and then send the query result to the server for display.

[0072] The advantages of the present invention are as follows: An edge hardware system based on domestic Ascend processors is built to obtain the video stream of the monitoring camera to update the background image in real time, and a binary decision matrix is obtained through the difference frames of multiple current images and the background image to determine whether there are moving targets in the video. The AI acceleration unit of the Ascend processor is used to run the target detection network and the pedestrian appearance prediction network to extract the pedestrian targets and the appearance characteristics of pedestrians in the moving targets. The key video segments where pedestrian targets appear in the monitoring video, the number of pedestrians in the video segments, and the appearance characteristics are extracted at the edge end. A structured storage method is established through pedestrian attributes to provide local storage and retrieval services for the key video segments. The main computing tasks are completed by the edge nodes, which greatly speeds up the real-time performance and response speed of handling emergencies in the monitoring video, reduces the computing pressure on the server, and the structured data storage method improves the efficiency of data storage and retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a schematic diagram of the topological structure of the edge hardware system of the present invention.

[0074] Figure 2 It is a schematic diagram of the structure of the edge hardware system of the present invention.

[0075] Figure 3 It is a flowchart of the edge software working process of the present invention.

[0076] Figure 4 It is a flowchart of the overall working process when the edge end and the data center are connected in the present invention.

[0077] Figure 5 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] The present invention will be introduced in detail in combination with the accompanying drawings:

[0079] Figure 1 It is the topological structure of the edge-side hardware system. Structured data is generated at the edge-side and sent to the server in the data center for backup storage.

[0080] Figure 2 It is the edge-side hardware system structure of this device. It consists of a main control chip (108), an EMMC module (101), an RTC module (102), a debugging interface (103), a MIPI interface (107), a gigabit Ethernet interface (106), a USB3.0 interface (105), and an SDHC module (104).

[0081] The main control chip (108) is the core of the entire embedded hardware system. It not only needs to run the logical code at the edge-side but also requires rich peripheral interfaces. The altas200 module of Huawei Ascend is adopted, which integrates 8 Cortex-A55 cores. This module also integrates the Ascend 310 AI processor of Hisilicon. There are 2 AI cores built into its chip, which is a hardware module for accelerating the calculation of deep learning neural networks and can achieve a computing power of 22 TOPS INT8, enabling it to easily run some small and medium-sized deep learning algorithms. At the same time, it also supports 128-bit wide LPDDR4x. Using the SDK development package provided by Huawei, the linux kernel and drivers are reconstructed to run the linux operating system, and peripheral interfaces such as pcie and network protocol stacks can be used. Deep learning algorithms are deployed on it while running our logical code.

[0082] The EMMC (101) module can solidify the linux kernel and file system of the embedded platform. After power-on, the main control chip automatically loads the startup file, linux kernel, and file system in the EMMC according to its hardware strobe pins. Compared with an SD card, it has a longer lifespan, lower power consumption, and faster data read and write rates.

[0083] The RTC module (102) can provide an accurate time reference for the embedded platform. A crystal oscillator with relatively high precision is used as the clock source. It provides a real-time clock for the embedded platform to ensure that the local database will not have data confusion due to the lack of timestamps. It can still provide a time reference even when the device accidentally loses power and the system crashes and restarts.

[0084] The debugging interface (103) includes an RS232 interface, an LCD display interface, a working status indicator light and a debugging button. The main control chip outputs the system startup information through the UART pin, including the ROM code, the Linux kernel and the startup information of the boot-up service solidified in the atlas200 module, and provides a serial terminal to operate the Linux system. The TTL level of the UART pin of the main control chip is converted to the RS232 level through the RS232 chip and then communicates with the computer. The processed key video can be directly displayed through the LCD screen, and the debugging button can be used to decide whether to turn off the LCD to maintain low power consumption operation. The working indicator light displays the different working states of the system.

[0085] The main control chip is connected to the SD card via the SDIO interface (104), and the SD card is used to cache the structured data of the surveillance video.

[0086] The MIPI interface (107) and the USB 3.0 interface (105) can be used to connect different types of surveillance video cameras.

[0087] The RGMII interface and MDIO interface of the main control chip are connected to the PHY chip, and the PHY chip provides a gigabit network port (106) to the outside. The gigabit network port can read the input network video stream or upload data to the server.

[0088] The edge hardware system obtains the camera video input through the USB interface, processes the camera input data through the V4L2 library of the Linux system and converts it into YUV image frames. In order to maintain the low power operation of the system, the deep learning algorithm cannot be used for real-time video detection. It is necessary to use the traditional digital image processing method to detect moving targets. The main steps are as follows:

[0089] Select an initial frame to initialize the background image B(x, y, z), and subtract the background image from each pixel on the original frame F at different adjacent moments to obtain the differential image matrix:

[0090] MD(D,t)={D(x,y,z,t 1 ),D(x,y,z,t 2 ),…,D(x,y,z,t n )} (9)

[0091] The matrix MD represents the difference image matrix, and D represents the difference frame obtained by subtracting the background image from each frame. The difference image matrix is ​​converted into a grayscale image matrix:

[0092] MG(G,t)={G(x,y,t 1 ),G(x,y,t 2 ),…,G(x,y,tn )} (10)

[0093] The matrix MG represents the grayscale image matrix, and G represents the grayscale frame converted from the differential frame. To reduce the computational complexity, the conversion formula for each pixel value is as follows:

[0094] V = (76R + 150G + 30B) >> 8 (11)

[0095] V is the value of each pixel after conversion to grayscale. R, G, and B represent the pixel values on the RGB three channels respectively. A threshold Tb is set to convert the grayscale image matrix into a binary matrix, and its conversion formula is:

[0096]

[0097] If the value of a pixel on the binary matrix Bi is 1, it indicates that the point may be a moving target, and the background image does not update this point. If it is 0, the background image updates this point. Then the update formula for the background image is:

[0098]

[0099] The matrix B represents the background image, the matrix Bi represents the binary decision matrix. If the point on this matrix is 0, the corresponding pixel point is updated, and if it is 1, it is not updated. F represents the original frame at adjacent times.

[0100] Perform salt-and-pepper noise reduction on the binary matrix Bi obtained in the previous step, find the total number c of pixel points with a value of 1, set a threshold Tc. If c > Tc, it indicates the existence of a moving target. Use the YOLOV3 algorithm for target detection, segment the selected pedestrian image, and send it to the pedestrian appearance extraction network to extract the appearance label. If a pedestrian target is detected, merge this frame into the H264 file and generate label information about the key video segment and the pedestrian target.

[0101] P = {p 1 , p 2 , …, p n} (2)

[0102] T = {t 1 , t 2 , …, t n} (3)

[0103] C = {c 1 , c 2 , …, c 35} (4)

[0104] Let \(P\) represent the camera IDs distributed at different locations, \(T\) represent different moments, and \(C\) represent the appearance labels of pedestrians, with a total of 35. The storage information for the key video segments can be represented by the camera ID, start time, and end time. Here, we change the end time to the total number of frames from the start time to the end time:

[0105]

[0106] where \(p\) i and \(t\) i represent the camera ID and start time respectively. These two can determine the file name of the stored H264 file. \(fps\) represents the frame rate when encoding the H264 file. Then represents the total number of frames stored in this key video segment, and thus the position of the key video segment \(v_i\) in the H264 file can be located:

[0107]

[0108] end_fream = start_fream + f(n) (7)

[0109] \(f\) represents the number of frames in each video segment. For pedestrian information, it can be represented by feature labels and key video segments:

[0110] pedestrian_info = {c i , v_i} (8)

[0111] Design the table structure according to the above formulas (5) and (7) and store it in the database.

[0112] Place the H264 on the SD card. When the cached data reaches the storage limit of the SD card, it will be backed up to the server regularly.

[0113] Figure 4It is the overall workflow when the edge device is connected to the data center. The servers in the data center act as the server side. When the monitoring video structured storage device installed at the edge uploads structured data at regular intervals, it will first initiate a connection request to the server side as the client. The server side needs to receive a large amount of data uploaded by the edge devices. Limited by network bandwidth and processing efficiency, the number of connections at the same time is limited. After the server side replies that the connection is allowed, the client will then upload the structured data. The H264 files are stored in the distributed file system, while the structured data such as the tag information of the key video ends will be stored in the distributed database. At the same time, the server will also provide a Web retrieval service. By accessing the Web page through the browser on any client machine, it supports multi-dimensional input of time, camera ID, and human features for retrieval. While retrieving locally, the server will send the retrieved person to each edge device. After receiving the retrieval task, each edge device will also perform local retrieval and then send the query results to the server for display.

[0114] The specific implementation manners described above have elaborated on the technical solutions of the present invention. It should be understood that the above-described implementation manners are the preferred implementation manners of the present invention, but the implementation manners of the present invention are not limited by the described embodiments. Any other modifications, substitutions, combinations, and tailoring made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.

Claims

1. An edge - end monitoring video structured storage method based on Ascend processors, comprising the following steps: 1) Build an edge - end hardware system based on domestic Ascend processors, which consists of a main control chip (108), an EMMC module (101), an RTC module (102), a debugging interface (103), a MIPI interface (107), a gigabit Ethernet interface (106), a USB3.0 hub interface (105), and an SDHC module (104); 2) Update the background image in real - time according to the monitoring camera video stream. The specific background image update formula is shown in Equation (1): The matrix B represents the background image, the matrix Bi represents the binary decision matrix. If the point on the matrix is 0, the corresponding pixel point is updated; if it is 1, it is not updated. F represents the original frame at adjacent moments; 3) Perform salt - and - pepper denoising on the binary matrix Bi obtained in the previous step, find the total number c of pixel points with a value of 1, set a threshold Tc. If c > Tc, it means there is a moving target. Use the YOLOV3 algorithm for target detection, segment the selected pedestrian images, and send them to the pedestrian appearance extraction network to extract appearance labels; 4) For the target detection network in step (3), if a pedestrian target is extracted, merge this frame into the H264 file and generate label information about the key video segment and the pedestrian target; P = {p 1 , p 2 , …, p n} (2) T = {t 1 , t 2 , …, t n} (3) C = {c 1 , c 2 , …, c 35} (4) P represents the camera id distributed at different locations, T represents different moments, C represents the appearance label of the pedestrian. The storage information of the key video segment can be represented by the camera id, start time, and end time. Here, we change the end time to the total number of frames from the start time to the end time: where p i and t i represent the camera ID and the start time respectively. These two can determine the file name of the stored H264 file. fps represents the frame rate when encoding the H264 file. Then represents the total number of frames stored in this key video segment. Then the position of the key video segment vi in the H264 file can be located: end_fream = start_fream + f(n) (7) f represents the number of frames in each video segment. For pedestrian information, it can be represented by the feature label and the key video segment: pedestrian_info={c i ,v i} (8) Design the table structure according to the above formulas (5) and (7) and store it in the database; 5) Place the H264 on the SD card. When the cached data reaches the storage limit of the SD card, it will be regularly backed up to the server.

2. The edge - end monitoring video structured storage method based on Ascend processors according to claim 1, characterized in that the building of the edge - end hardware system based on domestic Ascend processors specifically includes: The main control chip uses the Altas200 module of Huawei HiSilicon, which integrates the Ascend 310 processor. Utilize its powerful computing power to run logic code and deploy deep - learning algorithms, and use various interfaces to drive peripherals; The EMMC module solidifies the startup files of the embedded platform and automatically loads the linux file system after power - on; The RTC module provides an accurate time reference for the embedded platform, and still has a time reference even after accidental power - off and system crash and restart; The debugging interface. The uart module of the main control chip outputs the system startup information using RS232, displays the UI interface of the embedded platform using the LCD, and decides whether to turn off the LCD to maintain low - power operation through the debugging button. The working indicator shows different working states of the system; The SDIO interface is connected to the SD card, and the SD card caches the structured data of the monitoring video; The MIPI interface and the USB3.0 hub interface can be used to connect different types of surveillance video cameras; The RGMII interface and the MDIO interface of the main control chip are connected to the PHY chip. The PHY chip provides a gigabit network port externally. The gigabit network port can read the input network video stream or upload data to the server.

3. A method for edge-side surveillance video structured storage based on the Ascend processor according to claim 1, characterized in that, The specific method for updating the background image in step (2) is: Select an initial frame to initialize the background image B(x, y, z). Subtract the background image from each pixel point on the original frames F at different adjacent times to obtain a differential image matrix: MD(D,t) = {D(x,y,z,t 1 ), D(x,y,z,t 2 ), …, D(x,y,z,t n )} (9) Matrix MD represents the differential image matrix, D represents the differential frame obtained by subtracting the background image from each frame. Convert the differential image matrix into a grayscale image matrix: MG(G,t) = {G(x,y,t 1 ), G(x,y,t 2 ), …, G(x,y,t n )} (10) Matrix MG represents the grayscale image matrix, G represents the grayscale frame converted from the differential frame. In order to reduce the amount of calculation, the conversion formula for each pixel point value is: V=(76R + 150G + 30B)>>8 (11) V is the value of each pixel point after conversion to grayscale. R, G, and B represent the pixel values on the RGB three channels respectively. Set a threshold Tb to convert the grayscale image matrix into a binary matrix. The conversion formula is: Among them, if the value of the pixel point on the binary matrix Bi is 1, it means that this point may be a moving target, and the background image does not update this point. If it is 0, the background image updates this point.

4. A method for edge-side surveillance video structured storage based on the Ascend processor according to claim 1, characterized in that, The appearance labels described in step 3) specifically include age, gender, and clothing color, with a total of 35 categories; there are 35 appearance labels P in total.

Citation Information

Patent Citations

  • Video structured storage method and device based on edge computing, equipment and medium

    CN110795595A

  • Mercuric chloride processor-based knuckle print and finger vein identity recognition device

    CN113052072A