Embedded video recognition safety integrated control platform including an AI-based video recognition system
The video recognition system addresses inaccuracies in existing systems by employing a CNN-based object recognition and real-time re-learning via YOLO, ensuring accurate detection and prevention of accidents near heavy equipment.
Patent Information
- Application Number
- JP2024525215
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-10-26
- Publication Date
- 2025-09-10
- Estimated Expiration
- 2043-10-26
AI Technical Summary
Existing video recognition systems for construction and industrial sites suffer from inaccurate object detection due to harsh environmental factors and the difficulty in updating learning datasets, leading to frequent collisions and entrapment accidents involving heavy equipment.
A video recognition system utilizing a CNN-based object recognition unit, an event information generating unit, and a deep learning-based YOLO algorithm for real-time re-learning and classification of detected events, integrated with an LTE communication unit for updating learning models and ensuring accurate detection of human objects near heavy equipment.
The system achieves over 88% object detection accuracy, a 7m maximum recognition distance, 360-degree coverage, and rapid risk recognition, effectively preventing collisions and entrapment accidents through continuous learning and data updates.
Smart Images

Figure 0007737119000001 
Figure 0007737119000002 
Figure 0007737119000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an embedded video recognition safety integrated control platform including an artificial intelligence-based video recognition system, and more particularly, to a video recognition system including: a video data receiving unit that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN (Convolutional Neural Network)-based object recognition unit that recognizes human objects in the video data received from the video data receiving unit and boxes and detects the recognized human objects according to the recognized human object data; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius and situation, including the vicinity of heavy equipment; a false detection / undetected classification unit that receives event information detected from the CNN-based object recognition unit and classifies it into true detection, false detection, or undetected; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Back), that classifies the false detection or undetected classification received from the false detection / undetected classification unit. the AI video recognition platform including an LTE communication unit for communicating with an external device including an external server or an administrator terminal; a re-learning unit for re-learning in real time through a CNN-based object recognition unit; a learning model update unit for uploading the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit online in real time; and an AI video recognition platform including an LTE communication unit for communicating with the video recognition system and an external device including an external server or an administrator terminal. [Background technology]
[0002] According to the OECD National Industrial Accident Fatality Comparison and Analysis Report published by the Korea Construction Industry Institute in September 2020, it was confirmed that in 2017, there were approximately 3.61 industrial accident fatalities for every 100,000 workers in South Korea.
[0003] In particular, the average industrial death rate among OECD member countries was 2.43 per 100,000 people, with South Korea (3.61) ranking fifth highest after Canada (5.84), Turkey (5.17, as of 2016), Chile (4.04), and Luxembourg (3.69). Furthermore, South Korea ranked first among the "3050 club"—countries with a population of over 50 million and a per capita national income of over $30,000—following South Korea (3.61), Japan (1.50), the US (3.36), the UK (0.88), France (2.18), Germany (1.03), and Italy (2.10). While the industrial death rate is generally on a downward trend, South Korea's industrial death rate remains higher than other countries.
[0004] Additionally, according to the current status of industrial accidents published by the Ministry of Employment and Labor in 2020, the total number of accidents was confirmed to be 92,383. When examining accident victims by industry, excluding other businesses, the construction industry had the highest number of accidents at 24,617 (26.6% of the total), followed by the manufacturing industry at 23,127 (25.0% of the total), accounting for 51.6% of all industrial accidents.
[0005] The highest number of deaths occurred in the construction industry (567, 27.5%) and manufacturing industry (469, 22.7%), with the highest number of deaths occurring in the manufacturing industry (249) and construction workplaces (225), both of which fall within the 5-49 age bracket.
[0006] In addition, it was confirmed that more than 51% of all industrial accidents such as those mentioned above were conventional accidents (falls / tumbles / being pinched), which occurred due to non-compliance with basic safety rules and safe work methods.
[0007] For reference, among the current status of fatal accidents by type of accident, examples of collisions and entrapment accidents involving construction machinery at industrial sites are typically collisions and entrapment with heavy equipment such as excavators, dump trucks, forklifts, and containers.
[0008] In addition, the direct losses (industrial accident compensation payments) due to the above-mentioned industrial accidents amounted to 5,996,819 million won in 2021, an increase of 8.45% compared to the previous year. The estimated economic losses, including direct and indirect losses, amounted to 29,984,095 million won, an increase of 8.45% compared to the previous year, which is higher than Gyeonggi Province's entire budget for 2021 (28,792.5 billion won) and amounts to 1.5% of the national GDP.
[0009] As mentioned above, the government is also implementing various policies to prevent the continued occurrence of industrial disasters and safety accidents. The revised Occupational Safety and Health Act, which significantly strengthened safety regulations at industrial sites, including preventing the "outsourcing of risks," came into effect on January 27, 2022. The Act also provides for imprisonment for management in the event of the death or accident of a worker, introduces a punitive damages system, and makes employers and corporations liable for compensation of up to five times the amount of damages if they intentionally or through gross negligence violate their safety and health obligations and cause a serious disaster or damage.
[0010] However, because construction and industrial sites use a wide variety of construction machinery, there are potential disaster-causing factors everywhere. Furthermore, because numerous processes are carried out in conjunction with one another at construction sites, if a previous process is not carried out properly, it will have an immediate impact on the next process. The combination of potential dangerous factors can lead to large-scale disasters occurring all at once, making it extremely difficult to reduce industrial accidents at construction sites.
[0011] In particular, various construction machines, such as forklifts, excavators, and dump trucks, are used at various industrial and construction sites to improve the ease and productivity of work such as moving materials and excavating. However, since these construction machines are often used in cooperation with workers, accidents resulting in deaths frequently occur due to inexperienced or careless operation of the machines or failure to recognize blind spots. Therefore, it is necessary to develop technology to prevent accidents by developing safety systems for construction machines and equipment.
[0012] As a result, current Korean industrial sites are making efforts to prevent collisions and entrapment accidents by attaching rear cameras and collision prevention bars to construction equipment and by stationing traffic lights, as shown in [Figure 1]. However, workers and traffic lights are constantly exposed to danger as they work simultaneously around heavy equipment, and in most cases traffic lights and surveyors still suffer collisions and entrapment accidents.
[0013] Therefore, various proximity warning devices have been developed and used recently to prevent the risk of collision safety accidents involving workers due to heavy equipment at industrial sites. These are broadly divided into tag-based and non-tag-based technologies that use various sensors, and products incorporating various technologies are being developed, commercialized, and sold.
[0014] The tag-based technology is a method in which the vehicle and worker carry sensor tags that can directly transmit and receive radio waves to measure the distance between heavy equipment and the worker. Tags are divided into unidirectional and bidirectional tag types depending on the tag's radio wave transmission and reception method. The sensors mainly used in bidirectional tag types are RF and UWB, which measure the radio wave strength indicator (RSSI) and time of arrival (TOA) received by the sensor and convert them into distance. While RF-based proximity alarm products are currently mainstream, UWB-based products, which offer superior performance in terms of distance measurement accuracy and uniformity, have been developed and commercialized.
[0015] In addition, the unidirectional tag system includes passive products that use RFID tags, but this system involves installing an RFID reader on the vehicle and carrying a directional passive RFID tag as a tag for pedestrians. This type of product has a larger distance recognition error than bidirectional tag systems, which reduces the reliability of the product.
[0016] In contrast, the non-tag-based approach warnings use simple sensors such as non-tag-based cameras and ultrasound, and due to the characteristics of the non-tag-based sensors themselves and their limited functionality in terms of operation, radar products using radio wave signals and radar products using optical technology have been developed, and recently, with the development of AI technology, object image recognition products using cameras have also been developed.
[0017] In particular, when considering the prior art for managing and controlling dangerous situations for construction equipment and workers based on the AI image recognition system, Korean Patent No. 10-1808587 (registration date: December 7, 2017) describes a video input unit comprising a PTZ (Pan / Tilt / Zoom) camera or a fixed camera that can rotate 360 degrees and has built-in up / down / left / right and zoom functions, an abnormal situation detection unit that detects whether the video captured by the video input unit corresponds to one of abnormal situations selected from intrusion, crowding, loitering, abandonment, emergence, entering and exiting, climbing walls, people counting, falling, going the wrong way, and number recognition using a preset abnormal situation algorithm, and when an abnormal situation is detected by the abnormal situation detection unit, object recognition is performed through an image pre-processing process, an object extraction image generation process, and an object analysis process, and the object analysis process is performed by analyzing the object extracted in the object extraction image generation process. An intelligent integrated surveillance and control system using object recognition, tracking and abnormal situation detection technologies has been developed, which includes an object recognition unit that extracts edge patterns from detected objects using a Haar algorithm, a HOG algorithm or a SURF algorithm depending on the abnormal situation and recognizes and discriminates objects through pattern matching with learned data accumulated and stored through a deep learning algorithm; an object tracking unit that analyzes changes in coordinates of objects recognized by the object recognition unit, predicts the moving path or moving direction of the object in the captured image, and tracks the object so that it is located at the center of the captured image; and an integrated control unit that displays and monitors images captured by the image input unit, and sets and controls the image input unit, abnormal situation detection unit, object recognition unit and object tracking unit.
[0018] In addition, Korean Patent No. 10-2185859 (registered date: November 26, 2020) relates to an object tracking device using deep learning that recognizes human objects from video data and tracks the corresponding human objects on a frame-by-frame basis, and includes: a video data receiving unit that receives the video data from a camera module; a pre-processing unit that resizes the received video data and reduces the influence of light; an object recognition unit that recognizes human objects from video data that has completed pre-processing through deep learning-based object recognition learning and boxes the recognized human objects according to the recognized human object data; a calculation unit that calculates whether or not the boxed human object data in the first frame portion of the video data matches the boxed human object data in the second frame portion following the first frame, and recognizes boxes that show a degree of match equal to or higher than a set degree of match as the same human object; and a box recognized as the same human object by the calculation unit. a coordinate information generating unit that accumulates and collects the movement directions of the human objects measured by the movement direction measuring unit, extracts a movement path of the human object, and stores coordinate information for the shape of a road through the extracted movement path; a coordinate information receiving unit that receives coordinate information for measuring a moving population from an administrator terminal; and a moving population calculating unit that calculates the moving population as 0 if the coordinate information received from the coordinate information receiving unit is not included in the coordinate information for the road, and calculates the moving population based on each box that passes through the corresponding coordinate and sets a straight line perpendicular to the extension direction of the road at the corresponding coordinate if the coordinate information received from the coordinate information receiving unit is included in the coordinate information for the road, wherein the deep learning-based object recognition learning is performed in real time through YOLO (You Only Look Once).
[0019] In addition, Korean Patent No. 10-2206662 (Registration date: January 18, 2021) is a patent for a port container terminal that uses a port gate, yard / block, and block entrance. The system includes a number of cameras installed in each area of the Port Container Terminal (Port Container Center), ARMGC, and QC; and an FPGA-based embedded vision system (TLEM) that receives video data from the cameras and uses a deep learning module to detect each object in the camera video according to the learning data, recognize characters, and perform lane recognition, vehicle license plate recognition, character recognition of container number (ISO code), container damage recognition, recognition of vehicles entering block entrances, detection of vehicles and workers entering dangerous areas, vehicles driving in the wrong direction, detection of whether yard workers are wearing safety protective gear / safety vests, and detection of the location of loading and unloading equipment, and transmits the learned object extraction events (text) and object extraction video data marked with a square box to the middleware. The FPGA-based embedded vision system (TLEM) is equipped with a non-GPU-based deep learning module that detects each object in the video and recognizes the characters of vehicle license plates and container numbers (ISO code). A vision camera system for vehicle entrance and exit control and object recognition at a port container terminal has been developed.
[0020] In addition, Korean Patent No. 10-2263512 (registration date: June 4, 2021) describes an IoT integrated intelligent video analysis platform system that integrates and analyzes video data and non-video data, including a video data acquisition unit that acquires at least one piece of video data; a non-video data acquisition unit that acquires at least one piece of non-video data; a video data processing unit that analyzes the video data; a non-video data processing unit that analyzes the non-video data; and an integrated data determination unit that finally determines the abnormal situation when the video data processing unit or the non-video data processing unit determines that the abnormal situation exists from the video data or the non-video data; and the video data processing unit recognizes an object from the acquired video data, estimates the state of the object, estimates the authenticity of the object, and determines the action of the object. the non-video data processing unit, in analyzing the non-video data, defines a case where a measurement value of the non-video data deviates from a data range of a normal situation as an abnormal event, and determines an abnormal situation by considering whether the abnormal event has occurred, the time of occurrence, and the number of occurrences per a predefined unit time; the video data processing unit further includes an object processing unit that processes a function of recognizing objects from the acquired video data; and a user learning setting unit that provides a user with a function related to machine learning of video data; the object processing unit further includes an object authenticity identification unit that extracts objects from the video data and determines whether they are forged; an object state recognition unit that estimates a state of an object from the video data; and an object action recognition unit that estimates an action event of an object from the video data;The object authenticity identification unit extracts an image from the video data, analyzes the hue of each pixel constituting the extracted image, extracts a desired hue from the analyzed hue, and then derives a probability that the object is authentic through an authenticity determination algorithm. In order to extract a hue ratio through a K-means Clustering algorithm, the unit minimizes variance between clusters using a data similarity-based clustering algorithm. The clustered hues are used to identify the hue ratio within the item. The hue ratio can be extracted through OpenCV. A hue magnification factor for the authentic product is learned, and the difference between the authentic product and the counterfeit product image can be distinguished based on the extracted hue ratio. The unit also uses a DCGAN (Deep Convolution Generative Adversarial An IoT integrated intelligent image analysis platform system capable of smart object recognition has been developed, which generates a copy image using a 3D Network algorithm, and derives the probability that the object is genuine by applying an illegal copy detection algorithm that corresponds to a learning model for genuine product detection through feedback and learning between genuine and imitation product models using the difference in surface material.
[0021] However, although the above-mentioned conventional technologies are effective in terms of vehicle access control and worker safety management through object recognition and tracking from video data, they have fatal problems such as insufficient accuracy in object detection due to issues such as incorrect recognition due to various objects (pillars) that are characteristic of industrial sites, the occurrence of undetected or incorrect detection of objects due to harsh external environmental factors that are characteristic of construction sites, and the difficulty of collecting and updating automatic learning datasets in various environments.
[0022] As a result, the inventors have overcome the limitations of existing video recognition proximity warning systems, improved the visibility of construction equipment operators, ensured durability suitable for industrial sites so that blind spots around equipment can be accurately monitored in real time, and developed an AI-based video recognition platform that is capable of AI-based video object detection that is capable of omnidirectional detection while ensuring real-timeness and accuracy, communication that enables dangerous situations and equipment fleet management, and the creation of AI learning datasets, thereby completing the present invention. [Prior art documents] [Patent documents]
[0023] [Patent Document 1] Korean Patent Registration No. 10-1808587 (Registration Date: December 7, 2017) [Patent Document 2] Republic of Korea Registered Patent No. 10-2185859 (Registration Date: November 26, 2020) [Patent Document 3] Republic of Korea Patent Registered No. 10-2206662 (Registration Date: January 18, 2021) [Patent Document 4] Republic of Korea Registered Patent No. 10-2263512 (Registration Date: June 4, 2021) Summary of the Invention [Problem to be solved by the invention]
[0024] The present invention is directed to solving the above-mentioned problems, and provides a video recognition system including: a video data receiving unit that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN-based object recognition unit that recognizes a human object in the video data received from the video data receiving unit and detects the recognized human object by boxing it according to the recognized human object data; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius or situation, including the vicinity of heavy equipment; a false detection / undetected classification unit that receives the event information detected from the CNN-based object recognition unit and classifies it into true detection, false detection, or undetected; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Older) that classifies the false detection or undetected classification received from the false detection / undetected classification unit. a learning model update unit that uploads the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit in online real time; and an AI video recognition platform that includes an LTE communication unit for communicating with the video recognition system and an external device including an external server or an administrator terminal. [Means for solving the problem]
[0025] In order to solve the above technical problems, the present invention provides a video recognition system including: a video data receiving unit that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN-based object recognition unit that recognizes a human object in the video data received from the video data receiving unit and box-detects the recognized human object according to the data of the recognized human object; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius or situation, including the vicinity of heavy equipment; a false detection / undetected classification unit that receives the event information detected from the CNN-based object recognition unit and classifies it as a true detection, a false detection, or an undetected event; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Older) that classifies the false detection or undetected event classification received from the false detection / undetected classification unit. The technical solution is an embedded video recognition safety integrated control platform including an AI-based video recognition system, which includes: a re-learning unit that re-learns in real time through a CNN-based object recognition system (CNN-based object recognition system), a learning model update unit that uploads the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit online in real time; and an AI video recognition platform including an LTE communication unit for communication between the video recognition system and an external device including an external server or an administrator terminal.
[0026] The preset danger radius and situation including the periphery of the heavy equipment in the CNN-based object recognition unit includes setting a virtual boundary or a virtual area in the video shooting data received from the video data receiving unit, and appearing, exiting, or falling within the virtual boundary or virtual area.
[0027] The re-learning dataset is stored together with the video data after labeling the video data according to the false detection or undetected classification using an auto-labeling tool, and is uploaded online in real time.
[0028] The external device communicating with the LTE communication unit further includes an equipment fleet management system including equipment operating time, movement location tracking, and equipment downtime, and the equipment fleet management system is linked to the CNN-based object recognition unit to update and reflect the pre-set danger radius and situation including the periphery of the heavy equipment.
[0029] The external device communicating with the LTE communication unit further includes a site danger map display unit that collects event information detected from the CNN-based object recognition unit and displays danger zones on a map.
[0030] The external device communicating with the LTE communication unit further includes a dangerous event monitoring display unit for each equipment and each worker for tracking the equipment work route and the worker movement line based on the event information detected by the CNN-based object recognition unit.
[0031] The LTE communication unit is equipped with a GPS for determining and tracking the location of equipment or workers.
[0032] The external server or administrator terminal receives and uploads event information including event images or clips in real time when a worker approaches within the danger radius detected by the CNN-based object recognition unit.
[0033] The embedded video recognition safety integrated control platform, which includes the AI-based video recognition system, has an object detection accuracy of over 88% (people, F1 Score standard), a maximum recognition distance of over 7m, a 360-degree range, high temperature reliability up to 60℃, and a risk factor recognition speed of less than 0.5 seconds. [Effects of the Invention]
[0034] The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of the present invention includes a video recognition system including: a video data receiving unit that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN-based object recognition unit that recognizes a human object in the video data received from the video data receiving unit and box-detects the recognized human object according to the recognized human object data; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius or situation, including the vicinity of heavy equipment; a false detection / undetected classification unit that receives the event information detected from the CNN-based object recognition unit and classifies it into a true detection, a false detection, or an undetected event; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Back), that classifies the false detection or undetected event classification received from the false detection / undetected classification unit. The AI video recognition platform includes an AI image recognition platform that includes a re-learning unit that performs real-time re-learning via a learning model updater (CNN-based object recognition unit), a learning model updater that uploads the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit online in real time, and an LTE communication unit for communication between the image recognition system and external devices including an external server or an administrator terminal. The AI image recognition platform ensures object recognition accuracy in various industrial or construction environments, improves the reliability of dangerous situation events, and has groundbreaking effects in preventing collisions and entrapment accidents through sustainable learning data collection and re-learning. [Brief explanation of the drawings]
[0035] [Figure 1] This is an example of preventing collisions and pinch accidents involving construction equipment. [Figure 2] 1 is an overall schematic diagram of an embedded video recognition safety integrated control platform including an artificial intelligence-based video recognition system of the present invention; [Figure 3] 1 is a detailed schematic diagram of an embedded video recognition safety integrated control platform including an artificial intelligence-based video recognition system according to the present invention; [Figure 4]1 is an overall flowchart of an embedded video recognition safety integrated control platform including an artificial intelligence-based video recognition system of the present invention; [Figure 5] FIG. 1 is a process diagram of the re-learning part of the embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of the present invention. [Figure 6] FIG. 1 is a configuration diagram of the re-learning unit of an embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of the present invention. [Figure 7] 1 is a diagram illustrating an example of positive detection, false detection, or undetected classification according to the present invention. [Figure 8] This is an example diagram of a re-learning dataset created using the video shooting data auto-labeling tool. [Figure 9] This shows the structure of YOLO. [Figure 10] FIG. 1 is a diagram illustrating an example of an equipment fleet management system added to the present invention. [Figure 11] This is a diagram illustrating an example of dangerous event monitoring by equipment and worker. [Figure 12] FIG. 2 is a test and performance specification diagram of the video recognition safety integrated control platform of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] The present invention will be described in detail below with reference to examples and / or drawings so that those skilled in the art can easily understand the present invention. However, the present invention may be embodied in various different forms and is not limited to the examples and / or drawings described herein.
[0037] First, as shown in FIGS. 2 and 3, the embedded video recognition safety integrated control platform including the AI-based video recognition system of the present invention includes: a video data receiving unit 101 that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN-based object recognition unit that recognizes a human object in the video data received from the video data receiving unit and boxes and detects the recognized human object according to the recognized human object data; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius and situation including the vicinity of heavy equipment; a false detection / undetected classification unit that receives the event information detected from the CNN-based object recognition unit and classifies it into true detection, false detection, or undetected; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Back), that classifies the false detection or undetected classification received from the false detection / undetected classification unit. the AI video recognition platform 103 including a re-learning unit that performs real-time re-learning via a CNN-based object recognition unit (CNN-based object recognition unit), a learning model update unit that uploads the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit online in real time, and an LTE communication unit for communication between the video recognition system and an external device including an external server or an administrator terminal.
[0038] In this case, the preset danger radius and situation including the periphery of the heavy equipment in the CNN-based object recognition unit may include setting a virtual boundary or a virtual area in the video shooting data received from the video data receiving unit 101, and including appearance, entrance, and fall within the virtual boundary or virtual area, and the setting of the virtual boundary or virtual area may be changed or set at any time by the CNN-based object recognition unit.
[0039] More specifically, as shown in FIGS. 2 to 4, the operation flowchart of the embedded video recognition safety integrated control platform including the AI-based video recognition system of the present invention includes a video recognition system 102 that 1) recognizes a human object from video data received from the video data receiving unit 101, 2) recognizes a human object using a CNN-based object recognition unit, and detects the recognized human object by boxing it according to the recognized human object data, and 3) generates a warning light or a warning alarm when the detected human object is exposed to a predetermined danger radius and situation, including around heavy equipment. Before generating a warning light or a warning alarm, the system receives event information detected from the CNN-based object recognition unit and classifies it into true detection, false detection, or undetected. The false detection / undetected classification is re-learned in real time using YOLO (You Only Look Once), a deep learning-based object recognition learning algorithm, and then transmitted to the event information generating unit, thereby enabling more accurate detection of event information through video recognition.
[0040] In particular, the real-time re-learning process through YOLO (You Only Look Once) is performed as shown in FIGS. 5 and 6, through the steps of: image resizing and pre-processing (S11) the video data received from the video data receiving unit 101; object detection (S12) performing individual recognition on the image resized in the pre-processing (S11) using YOLO4 and TensorRT, generating a Bbox, and tracking the object; primary object recognition (S13) inferring whether the object detected in the object detection (S12) corresponds to a dangerous situation; generating event information from an event information generating unit if the object inferred in the primary object recognition (S13) corresponds to a dangerous situation; and secondary object recognition (S15) inferring that the object inferred in the primary object recognition (S13) does not correspond to a dangerous situation.
[0041] Meanwhile, the re-learning dataset is stored together with the corresponding video data after labeling the video data according to the false detection or undetected classification using an auto-labeling tool, and then uploaded online in real time.
[0042] That is, as shown in [FIGS. 7] and [FIG. 8], there is a false positive / missing detection classification unit that receives detected event information from the CNN-based object recognition unit and classifies it into true positive, false positive, or false negative; a re-learning unit that re-learns the false positive or false negative classification received from the false positive / missing detection classification unit in real time using YOLO (You Only Look Once), a deep learning-based object recognition learning algorithm; and a re-learning data set generated by the re-learning unit that is labeled into video shooting data according to the false positive or false negative classification using an auto-labeling tool, and then stored together with the corresponding video shooting data and uploaded online in real time.
[0043] YOLO is a deep learning-based map learning algorithm for object detection. It stands for You Only Look Once, and it determines the classification and location of an object through a single regression of the image. YOLO is based on a CNN structure, and its network architecture is based on the GoogLeNet model and includes 24 convolutional layers and two fully connected layers.
[0044] Figure 9 shows the structure of YOLO. YOLO processes images by scaling the input image, running a convolutional network on the image, and then thresholding the result based on the model's confidence level. A bounding box consists of five elements: x, y, w, h, and a confidence score. (x, y) are the center coordinates of the bounding box relative to the grid cell boundary. (w, h) are the width and height of the bounding box. The confidence score indicates the IOU between the predicted bounding box and all ground truth bounding boxes. Each grid cell also predicts a conditional class probability, Pr(Classi|Object), C. These probabilities are conditioned on the grid cell containing the object. Only one set of class probabilities is predicted per grid cell, regardless of the number of bounding boxes, B. At test time, the conditional class probabilities are multiplied to provide a class-specific confidence score for each box.
[0045] In Figure 9, the score encodes the probability that the class appears in the box and how well the predicted box fits the individual. YOLO's system models detection as a regression problem. It divides the image into an S × S grid and predicts, for each grid cell, B bounding boxes, C confidence for the box, and C class probability. These predictions are encoded as an S × S × (B ≦ 5 + C) tensor.
[0046] In the present invention, we use YOLOv4, the fourth version of YOLO, which has the advantages of high speed, real-time detection, greatly improved accuracy, and excellent performance.
[0047] In addition, the external device communicating with the LTE communication unit further includes an equipment fleet management system including equipment operating time, movement location tracking, and equipment downtime, as shown in FIG. 10, and the equipment fleet management system may be linked to the CNN-based object recognition unit to update and reflect a preset danger radius and situation including the periphery of the heavy equipment.
[0048] In addition, the external device communicating with the LTE communication unit may further include a dangerous event monitoring display unit for each equipment and each worker for tracking the equipment work path and the worker's movement line based on the event information detected by the CNN-based object recognition unit, as shown in FIG. 11.
[0049] Furthermore, the LTE communication unit is equipped with a GPS for determining and tracking the location of equipment or workers.
[0050] In addition, the external server or administrator terminal can be configured to receive and upload event information including event images or clip videos in real time when a worker approaches within the danger radius detected by the CNN-based object recognition unit.
[0051] In particular, the embedded video recognition safety integrated control platform including the AI-based video recognition system according to the present invention has an object detection accuracy of 88% or more (human, F1 Score standard), a maximum recognition distance of 7m or more, a range of 360 degrees, high temperature reliability of 60°C, and a risk factor recognition speed of 0.5s or less, as shown in [Figure 12]. [Industrial Applicability]
[0052] The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of the present invention includes a video recognition system including: a video data receiving unit that receives video data of the vicinity of a construction or industrial site from a camera module; a CNN-based object recognition unit that recognizes a human object in the video data received from the video data receiving unit and box-detects the recognized human object according to the recognized human object data; and an event information generating unit that generates a warning light or a warning alarm when the human object detected by the CNN-based object recognition unit is exposed to a predetermined danger radius or situation, including the vicinity of heavy equipment; a false detection / undetected classification unit that receives the event information detected from the CNN-based object recognition unit and classifies it into a true detection, a false detection, or an undetected event; and a deep learning-based object recognition learning algorithm, YOLO (You Only Look Back), that classifies the false detection or undetected event classification received from the false detection / undetected classification unit. The AI video recognition system includes an AI video recognition platform including a re-learning unit that performs real-time re-learning via a learning model updater (CNN-based object recognition unit), a learning model updater that uploads the re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit online in real time, and an LTE communication unit for communicating with the video recognition system and external devices including an external server or an administrator terminal. The system has industrial applicability as it has a groundbreaking effect in ensuring object recognition accuracy in various industrial or construction environments, improving the reliability of dangerous situation events, and preventing collisions and entrapment accidents through sustainable collection and re-learning of learning data.
Claims
1. a CNN (Convolutional Neural Network)-based object recognition unit that recognizes worker objects around the construction or industrial site from the video data received from the video data receiving unit and boxes and detects the recognized worker objects in accordance with the data of the recognized worker objects; and an event information generation unit that generates a warning light or a warning alarm when the worker object detected by the CNN-based object recognition unit is exposed to a situation including appearance, entry, or fall within a predetermined danger radius including around heavy equipment; a false detection / undetected classification unit that receives event information detected from the event information generation unit and classifies the event information into true detection, false detection, or undetected; and a YOLO (You Only Look) classifier that classifies the false detection or undetected classification received from the false detection / undetected classification unit into a false detection or undetected classification, which is a deep learning-based object recognition learning algorithm. an AI video recognition platform including an LTE communication unit for communication between the video recognition system and an external device including an external server or an administrator terminal; a re-learning unit for re-learning in real time through a CNN-based object recognition unit; a learning model update unit for uploading a re-learned dataset generated by the re-learning unit to the CNN-based object recognition unit in online real time; and an AI video recognition platform including an LTE communication unit for communication between the video recognition system and an external device including an external server or an administrator terminal.
2. An embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system described in claim 1, characterized in that a situation including appearance, entry, or fall within a predetermined danger radius including the area around the heavy equipment in the event information generating unit includes setting a virtual boundary line or virtual area with the danger radius in the video shooting data received from the video data receiving unit, and appearance, entry, exit, or fall within the virtual boundary line or virtual area.
3. 10. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of claim 1, wherein the re-learning data set is stored together with the video capture data after labeling the video capture data according to the false detection or undetected classification using an auto-labeling tool, and is uploaded online in real time.
4. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of claim 1, wherein the external device communicating with the LTE communication unit further includes an equipment fleet management system including equipment operating time, movement location tracking, and equipment downtime.
5. 10. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of claim 1, wherein the external device communicating with the LTE communication unit further includes a site danger map display unit that collects event information detected from the CNN-based object recognition unit and displays danger zones on a map.
6. 10. The embedded video recognition safety integrated control platform including the artificial intelligence based video recognition system of claim 1, wherein the external device communicating with the LTE communication unit further includes a dangerous event monitoring display unit for each equipment and each worker that displays an equipment work route and a worker movement line based on the event information detected from the event information generating unit.
7. 10. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of claim 1, wherein the LTE communication unit is equipped with a GPS for determining the location of equipment or workers.
8. 2. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system of claim 1, wherein the external server or the administrator terminal receives and uploads event information including an event image or a clip video in real time when a worker approaches within a danger radius detected by the event information generating unit.
9. The embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system according to any one of claims 1 to 8, characterized in that the embedded video recognition safety integrated control platform including the artificial intelligence-based video recognition system has an object detection accuracy of 88% or more (human, F1 Score standard), a maximum recognition distance of 7m or more, a range of 360 degrees, high temperature reliability of 60°C, and a hazard factor recognition speed of 0.5s or less.
Citation Information
Patent Citations
Utility vehicle and corresponding apparatus, method and computer program for a utility vehicle
EP4064118A1
Monitoring system
JP2022016197A
Intelligent integration visual surveillance control system by object detection and tracking and detecting abnormal behaviors
KR101808587B1
Apparatus and method for tracking object
KR102185859B1
Vision camera system to manage an entrance and exit management of vehicles and to recognize objects of camera video data in a port container terminal and method thereof
KR102206662B1