Artificial intelligence-based load and heavy equipment collision prevention device and control method thereof
The AI-based collision avoidance system addresses blind spots by using unsupervised and supervised learning to detect loads and people, measuring distances, and outputting warnings, effectively reducing construction site accidents.
Patent Information
- Application Number
- PCT/KR2025/000598
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-01-10
- Publication Date
- 2025-08-28
AI Technical Summary
Construction sites face frequent safety accidents due to collisions between construction equipment, workers, and loads, primarily because conventional systems provide images from the equipment operator's perspective, leading to blind spots and increased collision risks.
An AI-based collision avoidance system that uses unsupervised learning to detect loads of variable size and supervised learning to identify people or fixed equipment, measuring distances and outputting warnings when collision risks are detected, with cameras installed for a third-person view to cover blind spots.
Accurately detects hazardous objects and people, providing real-time blind spot images and warnings to prevent accidents, improving user convenience and safety by anticipating potential collisions.
Smart Images

Figure KR2025000598_28082025_PF_FP_ABST
Abstract
Description
AI-based load and heavy equipment collision avoidance device and its control method
[0001] The present disclosure relates to a collision avoidance device for loads and heavy equipment. More specifically, it relates to a collision avoidance device for loads and heavy equipment using artificial intelligence and a control method thereof.
[0002] Recently, construction sites have been operating with a large number of workers and construction equipment in situations where there is a lot of load, and safety accidents such as collisions between construction equipment, between construction equipment and workers, and between loads and workers or construction equipment are occurring very frequently.
[0003] In particular, construction equipment such as excavators have a problem in that although it is easy to observe the front from the driver's seat, it is difficult to observe the rear and sides due to the equipment and loads loaded on the vehicle body, and thus safety accidents such as collisions occur at points that are difficult for the driver to observe.
[0004] In 2023, the overall number of deaths decreased compared to the previous year, but the number of construction site fatalities actually increased by 10. Furthermore, safety issues such as heavy equipment operators' visibility being obstructed by loads and situations where peripheral vision is not secured are a frequent cause of safety accidents.
[0005] However, in the case of conventional technology, images can be created based on images taken by cameras mounted on heavy equipment to check the surroundings of the equipment, but there was a problem that users felt uncomfortable because the images were taken from the perspective of the heavy equipment operator, so there were blind spots, and there was a high possibility of a safety accident occurring in which the load collides with the user.
[0006] The embodiment disclosed in the present disclosure aims to provide an artificial intelligence-based load and heavy equipment collision avoidance device that detects collisions with surrounding objects and people based on images captured from an objective point of view when load and heavy equipment are moving, and outputs a warning message when there is a high possibility of collision risk, thereby preventing safety accidents in advance.
[0007] The embodiment disclosed in the present disclosure aims to provide an artificial intelligence-based load and heavy equipment collision avoidance device that finds a load of variable size from a camera-captured image and extracts the load image using unsupervised learning-based segmentation artificial intelligence.
[0008] The embodiment disclosed in the present disclosure aims to provide an artificial intelligence-based load and heavy equipment collision avoidance device that extracts a load image by performing unsupervised learning in the case of a load with variable size, and extracts a heavy equipment image and a person image by performing supervised learning in the case of a person or heavy equipment with fixed size.
[0009] The problems to be solved by the present disclosure are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0010] In order to achieve the above-described technical task, the present disclosure provides an artificial intelligence-based device for preventing collisions between loads and heavy equipment, comprising: a camera for capturing images of the front; a communication module for transmitting and receiving data with an external device; a memory for storing at least one process for preventing collisions between loads and heavy equipment based on the artificial intelligence; And a processor that performs an operation according to the process, wherein the processor extracts an image of a hazardous material including the load and the heavy equipment and an image of a person from the image using an artificial intelligence model, measures a distance between the extracted image of the hazardous material and the extracted image of the person, determines a collision risk based on the measured distance, and if there is a collision risk identified, outputs a warning message, recognizes a section in which a load is loaded for each piece of heavy equipment from the image using an artificial intelligence model, selects 1 to N pixel coordinates around the recognized section, trains the artificial intelligence model based on unsupervised learning using the selected pixel coordinates as prompt information, and operates the training 1 to N times to select an artificial intelligence model with the highest reliability, predicts a shape of the load based on the selected artificial intelligence model, and extracts an image of the load from the predicted result.
[0011] At this time, the processor extracts an image of a hazardous material including the load from the image using an artificial intelligence model learned based on the unsupervised learning, and the size of the load may have a variable characteristic.
[0012] Additionally, the processor can recognize the heavy equipment from the image using segmentation, and provide at least one pixel point based on the recognized heavy equipment to the artificial intelligence model to predict the shape of the load.
[0013] Additionally, the camera may be installed in a position capable of capturing at least one of the heavy equipment, the load, surrounding objects, and people from a third-person viewpoint.
[0014] Additionally, the processor can generate a blind spot image from the image using the artificial intelligence model, and transmit the generated blind spot image to the external device through the communication module.
[0015] Additionally, the processor can measure the relative speed between the extracted hazardous material image and the extracted human image, determine the risk of collision based on the measured relative speed, and output the warning message if there is the determined risk of collision.
[0016] Additionally, the processor can extract images of hazardous materials including heavy equipment and images of people from the images using the artificial intelligence model.
[0017] In addition, an artificial intelligence-based load and heavy equipment collision avoidance method performed by a processor of a device for achieving the above-described technical task includes the steps of: capturing a front image; extracting an image of a hazardous material including the load and the heavy equipment and an image of a person from the image using an artificial intelligence model; measuring a distance between the extracted image of the hazardous material and the extracted image of the person; identifying a collision risk based on the measured distance; and outputting a warning message if the identified collision risk exists, wherein the method may further include the steps of: recognizing a section in which a load is loaded for each piece of heavy equipment from the image using the artificial intelligence model; selecting 1 to N pixel coordinates around the identified section; training the artificial intelligence model based on unsupervised learning using the selected pixel coordinates as prompt information; selecting the artificial intelligence model with the highest reliability by performing the training 1 to N times; predicting a shape of the load based on the selected artificial intelligence model; and extracting an image of the load from the predicted result.
[0018] According to the present disclosure, it is possible to accurately detect hazardous objects and people, transmit blind spot images and warning messages to workers and drivers in real time in case of a hazardous situation, and accurately track the movement of loads to prevent safety accidents, thereby improving user convenience.
[0019] According to the present disclosure, when a load and heavy equipment are moving, a collision with surrounding objects and people can be detected, and if there is a high possibility of collision, a warning message can be output to prevent safety accidents in advance, thereby improving user convenience.
[0020] According to the present disclosure, it is possible to find a load of variable size from a camera-captured image and more accurately extract the load image using unsupervised learning-based segmentation artificial intelligence, thereby improving user convenience.
[0021] According to the present disclosure, a real-time monitoring system can be provided to provide blind spot images for areas where a heavy equipment operator's field of vision is not secured, and a warning message can be output by detecting a potential collision risk that the heavy equipment operator is unaware of in advance, thereby reducing the possibility of a safety accident.
[0022] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0023] Figure 1 is a configuration diagram of an artificial intelligence-based load and heavy equipment collision avoidance device according to the present disclosure.
[0024] Figure 2 is a diagram illustrating a learning method of an artificial intelligence learning model (NN) according to the present disclosure.
[0025] FIG. 3 is a diagram illustrating a flowchart of a learning method of an artificial intelligence learning model (NN of FIG. 2) according to the present disclosure.
[0026] FIG. 4 is a flowchart illustrating an artificial intelligence-based load and heavy equipment collision avoidance method according to the present disclosure.
[0027] FIG. 5 is a diagram illustrating an embodiment of the operation of artificial intelligence-based load and heavy equipment collision avoidance according to the present disclosure.
[0028] FIG. 6 is a diagram illustrating an embodiment of the operation of artificial intelligence-based load and heavy equipment collision avoidance according to the present disclosure.
[0029] FIG. 7 is a diagram illustrating an embodiment of the operation of artificial intelligence-based load and heavy equipment collision avoidance according to the present disclosure.
[0030] FIG. 8 is a diagram illustrating a configuration of an artificial intelligence-based load and heavy equipment collision avoidance device according to the present disclosure.
[0031] FIG. 9 is a flowchart illustrating an artificial intelligence-based load and heavy equipment collision avoidance method according to the present disclosure.
[0032] FIG. 10 is a drawing illustrating the concept of the present disclosure according to the present disclosure.
[0033] FIG. 11 is a diagram illustrating an embodiment of unsupervised learning according to the present disclosure.
[0034] Figure 12 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0035] Figure 13 is a diagram showing a system configuration diagram in which a signal number according to the present disclosure is worn according to the present invention.
[0036] Figure 14 is a drawing showing a system configuration diagram in which the present invention according to the present disclosure is installed on heavy equipment.
[0037] FIG. 15 is a drawing showing an example of comparing the prior art and the present invention in a tower crane according to the present disclosure.
[0038] FIG. 16 is a drawing showing an example of comparing the prior art and the present invention in a forklift according to the present disclosure.
[0039] Figure 17 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0040] FIG. 18 is a drawing illustrating an embodiment in which a third-party viewpoint camera is installed according to the present disclosure.
[0041] Figure 19 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0042] Figure 20 is a drawing illustrating an example of a load shape extraction embodiment in the present disclosure.
[0043] FIG. 21 is a diagram illustrating a load shape extraction algorithm according to the present disclosure.
[0044] FIG. 22 is a diagram illustrating a configuration of an artificial intelligence-based load and heavy equipment collision avoidance device according to the present disclosure.
[0045] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and any content that is common in the technical field to which this disclosure pertains or that overlaps between embodiments is omitted. The terms "part, module, element, block" used in the specification may be implemented in software or hardware, and depending on the embodiments, multiple "parts, modules, elements, blocks" may be implemented as a single component, or a single "part, module, element, block" may include multiple components.
[0046] Throughout the specification, when a part is said to be "connected" to another part, this includes not only direct connection but also indirect connection, and indirect connection includes connection via a wireless communication network.
[0047] Additionally, when a part is said to "include" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.
[0048] Throughout the specification, when we say that an element is "on" another element, this includes not only cases where the element is in contact with the other element, but also cases where another element exists between the two elements.
[0049] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
[0050] Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0051] The identification codes for each step are used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.
[0052] The operating principle and embodiments of the present disclosure are described below with reference to the attached drawings.
[0053] The present disclosure herein can be implemented not only in a server system but also in various devices capable of performing computational processing and providing results to a user. For example, the present disclosure can include a computer, a server device, and a mobile terminal, or can be implemented in any one of these forms.
[0054] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
[0055] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0056] The above portable terminal may include, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD).
[0057] The artificial intelligence-related functions according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU or a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0058] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is learned by a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed in the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0059] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0060] The processor can create a neural network, train (or learn) a neural network, perform computations based on received input data, and generate information signals based on the results of the computations, or retrain the neural network.
[0061] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), Generative Adversarial Network (GAN), Liquid State Machine (LSM), Extreme Learning Machine (ELM), It will be understood by those skilled in the art that any neural network may be included, including but not limited to ESN (Echo State Network), DRN (Deep Residual Network), DNC (Differentiable Neural Computer), NTM (Neural Turning Machine), CN (Capsule Network), KN (Kohonen Network), and AN (Attention Network).
[0062] According to an exemplary embodiment of the present disclosure, the processor may be configured to perform a process for generating a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, Region with Convolution Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restrcted Boltzman Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for natural language processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting, Various artificial intelligence structures and algorithms, including optimization, recommendation, and data creation, can be utilized, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0063] FIG. 1 is a block diagram conceptually illustrating an artificial intelligence-based load and heavy equipment collision avoidance device (hereinafter, electronic device (100)) according to an exemplary embodiment of the present disclosure.
[0064] The electronic device (100) may include a memory (110), a processor (120) including an artificial neural network processing module, a camera (130), a communication module (140), and a LiDAR (150).
[0065] According to one embodiment of the present disclosure, the memory (110) may include various types of storage media including at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a resistive memory cell such as a resistive RAM (ReRAM), a phase change RAM (PRAM), a magnetic RAM (MRAM), a spin-transfer torque MRAM (MRAM), a conductive bridging RAM (CBRAM), and a ferroelectric RAM (FeRAM).
[0066] The processor (120) may be configured with one or more cores and may include a processor for data analysis and deep learning, such as a neural processing unit (NPU), a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU) of a computing device. The processor may read a computer program stored in the memory (110) and perform data processing for machine learning according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the processor may perform operations for learning a neural network. The processor may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from input data, calculating errors, and updating weights of a neural network using backpropagation. At least one of the NPU, CPU, GPGPU, and TPU of the processor may process learning of a network function. For example, a CPU and a GPGPU can be used together to process network function learning and data classification using network functions. Furthermore, in one embodiment of the present disclosure, processors of multiple computing devices can be used together to process network function learning and data classification using network functions. Furthermore, a computer program executed on a computing device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.
[0067] The communication module (140) can enable communication between multiple computing devices, thereby enabling distributed execution of operations for determining a user-controlled operation range or learning a model on each of the multiple computing devices. The network unit can enable communication between multiple computing devices, thereby enabling distributed processing of operations for learning a model using a network function.
[0068] A communication module (140) according to one embodiment of the present disclosure can operate based on any form of wired and wireless communication technology currently in use and implemented, such as short-range, long-range, wired, and wireless, and can also be used in other networks.
[0069] The camera (130) and the processor (120) can communicate with each other through a network. Here, communication through the network, i.e., transmission and reception of data, can be performed by wire or wirelessly. To this end, each communication unit (not shown) can be configured as a wired communication module that connects to the Internet, etc., through a LAN (Local Area Network), a mobile communication module that connects to a mobile communication network via a mobile communication base station to transmit and receive data, a short-range communication module that uses a WLAN (Wireless Local Area Network) series communication method such as Wi-Fi or a WPAN (Wireless Personal Area Network) series communication method such as Bluetooth or Zigbee, a satellite communication module that uses a GNSS (Global Navigation Satellite System) such as a GPS (Global Positioning System), or a combination thereof.
[0070] The electronic device (100) can transmit and receive wireless signals with at least one of a base station, an external terminal, and an external server on a mobile communication network constructed according to technical standards or communication methods (e.g., GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), etc.) through a network.
[0071] Wireless technologies of the present disclosure include, for example, Wireless LAN (WLAN), Wireless-Fidelity (Wi-Fi), Wireless Fidelity (Wi-Fi) Direct, Digital Living Network Alliance (DLNA), Wireless Broadband (WiBro), World Interoperability for Microwave Access (WiMAX), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), and Long Term (LTE). Evolution), LTE-A (Long Term Evolution-Advanced), etc.
[0072] In addition, the communication technology of the present disclosure is Bluetooth TM), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), TTL (Transistor-Transistor Logic), USB, IEEE1394, Ethernet, MIDI (Musical Instrument Digital Interface), RS232, RS422, RS485, Optical Communication, Coaxial Cable Communication.
[0073] An output unit according to one embodiment of the present disclosure may display a user interface (UI) for determining a user-controlled operation range and providing judgment results. The output unit may output any form of information generated or determined by the processor and any form of information received by the network unit.
[0074] In one embodiment of the present disclosure, the output unit may include at least one of a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, and a three-dimensional display (3D display). Some of these display modules may be configured as transparent or light-transmitting so that the outside can be viewed therethrough. This may be referred to as a transparent display module, and representative examples of the transparent display module include, but are not limited to, TOLED (Transparent OLED).
[0075] An input unit according to one embodiment of the present disclosure can receive user input. The input unit may include keys and / or buttons on a user interface for receiving user input, or physical keys and / or buttons. A computer program for controlling a display according to embodiments of the present disclosure can be executed based on user input through the input unit.
[0076] The input unit according to embodiments of the present disclosure may receive a signal by detecting a user's button operation or touch input, or may receive a user's voice or motion through a camera or microphone and convert it into an input signal. For this purpose, speech recognition technology or motion recognition technology may be used.
[0077] The input unit according to embodiments of the present disclosure may be implemented as an external input device connected to an external system. For example, the input device may be at least one of a touchpad, a touch pen, a keyboard, or a mouse for receiving user input, but this is merely an example and is not limited thereto.
[0078] An input unit according to one embodiment of the present disclosure can recognize a user touch input. The input unit according to one embodiment of the present disclosure may have the same configuration as an output unit. The input unit may be configured as a touch screen implemented to receive a user's selection input. The touch screen may use any one of a contact-type electrostatic capacitance method, an infrared photodetector method, a surface-aws (SAW) method, a piezoelectric method, and a resistive film method. The detailed description of the touch screen described above is merely an example according to one embodiment of the present disclosure, and various touch screen panels may be employed in a computing device. The input unit configured as a touch screen may include a touch sensor. The touch sensor may be configured to convert a change in pressure applied to a specific portion of the input unit or an electrostatic capacitance generated at a specific portion of the input unit into an electrical input signal. The touch sensor may be configured to detect not only the position and area of a touch, but also the pressure at the time of the touch. When a touch input is received on the touch sensor, a corresponding signal(s) is sent to the touch controller. The touch controller can process the signal(s) and then transmit corresponding data to the processor. This allows the processor to recognize which area of the input unit has been touched, etc.
[0079] In one embodiment of the present disclosure, the server may include other components for implementing a server environment. The server may include any type of device. The server may be a digital device equipped with a processor, memory, and computing power, such as a laptop computer, notebook computer, desktop computer, web pad, or mobile phone.
[0080] Figure 2 is a diagram illustrating a learning method of an artificial intelligence learning model (NN) according to the present disclosure.
[0081] In detail, FIG. 2 shows the structure of a convolutional neural network (CNN) as an example of a neural network structure (210).
[0082] Referring to FIG. 2, the artificial intelligence learning model (NN) may include a plurality of layers (L1 to Ln). Each of the plurality of layers (L1 to Ln) may be a linear layer or a non-linear layer, and in one embodiment, at least one linear layer and at least one non-linear layer may be combined and referred to as one layer. For example, the linear layer may include a convolution layer and a fully connected layer, and the non-linear layer may include a pooling layer and an activation layer.
[0083] For example, the first layer (L1) may be a convolutional layer, the second layer (L2) may be a pooling layer, and the nth layer (Ln) may be a fully connected layer serving as an output layer. An artificial intelligence learning model (NN) may further include activation layers and layers that perform other types of operations.
[0084] Each of the plurality of layers (L1 to Ln) can receive input data (e.g., an image frame) or a feature map generated from a previous layer as an input feature map, and generate output data (QR) by operating on the input feature map. In one embodiment, the output data (QR) can be various network parameters (mean value, probability weight, standard deviation).
[0085] A feature map is data that expresses various features of input data.
[0086] The feature maps (FM1, FM2, FMn) may have, for example, a two-dimensional matrix or a three-dimensional matrix (or tensor) form. In one embodiment, the input first feature map (FM1) may be data corresponding to a current state. The feature maps (FM1, FM2, FMn) have a width (W) (or column), a height (H) (or row), and a depth (D), which may correspond to the x-axis, y-axis, and z-axis on the coordinate system, respectively. In this case, the depth (D) may be referred to as the number of channels.
[0087] The first layer (L1) can generate the second feature map (FM2) by convolving the first feature map (FM1) with a weight kernel (WK). The weight kernel (WK) can filter the first feature map (FM1) and may also be referred to as a filter or a map. The depth of the weight kernel (WK), i.e., the number of channels, is the same as the depth of the first feature map (FM1), i.e., the number of channels, and the same channels of the weight kernel (WK) and the first feature map (FM1) can be convolved. The weight kernel (WK) can be shifted in a manner that traverses the first feature map (FM1) using a sliding window. The amount of shifting can be referred to as a "stride length" or "stride."
[0088] During each shift, each of the weight values included in the weight kernel (WK) may be multiplied and added to all pixel data in the area overlapping with the first feature map (FM1). The data of the first feature map (FM1) in the area where each of the weight values included in the weight kernel (WK) overlaps with the first feature map (FM1) may be referred to as extracted data. As the first feature map (FM1) and the weight kernel (WK) are convolved, one channel of the second feature map (FM2) may be generated. Although one weight kernel (WK) is shown in FIG. 3, in reality, multiple weight maps may be convolved with the first feature map (FM1) to generate multiple channels of the second feature map (FM2). In other words, the number of channels of the second feature map (FM2) may correspond to the number of weight maps.
[0089] The second layer (L2) can generate a third feature map (FM3) by changing the spatial size of the second feature map (FM2) through pooling. Pooling may be referred to as sampling or down-sampling. A two-dimensional pooling window (PW) is shifted on the second feature map (FM2) in units of the size of the pooling window (PW), and the maximum value (or the average value of the pixel data) among the pixel data in the area overlapping with the pooling window (PW) may be selected. Accordingly, a third feature map (FM3) with a changed spatial size may be generated from the second feature map (FM2). The number of channels of the third feature map (FM3) is the same as the number of channels of the second feature map (FM2). In one embodiment, the third feature map (FM3) may correspond to the output feature map on which convolution is completed, as described above in FIG. 3.
[0090] The nth layer (Ln) can classify the class (CL) of the input data by combining the features of the nth feature map (FMn). In addition, the nth layer (Ln) can generate output data (QR) corresponding to the class. In an embodiment, the input data may correspond to data corresponding to the current state, and the nth layer (Ln) can generate output data (QR) for determining the optimal action by extracting classes corresponding to multiple actions from the nth feature map (FMn) provided from the previous layer.
[0091] FIG. 3 is a diagram illustrating a flowchart of a learning method of an artificial intelligence learning model (NN of FIG. 2) according to the present disclosure.
[0092] Referring to the flowchart (310) of FIG. 3, the input feature maps (201) include D channels, and the input feature map of each channel can have a size of H rows and W columns (D, H, W are natural numbers). Each of the kernels (202) has a size of R rows and S columns, and the kernels (202) can include a number of channels corresponding to the number of channels (or depth) (D) of the input feature maps (201) (R, S are natural numbers). The output feature maps (203) can be generated through a 3D convolution operation between the input feature maps (201) and the kernels (202), and can include Y channels depending on the convolution operation.
[0093] The process of generating an output feature map through a convolution operation between one input feature map and one kernel can be explained with reference to FIG. 2b, and the two-dimensional convolution operation explained in FIG. 5b can be performed between the input feature maps (201) of all channels and the kernels (202) of all channels, thereby generating output feature maps (203) of all channels.
[0094] Referring back to FIG. 3, for convenience of explanation, it is assumed that the input feature map (210) has a size of 6x6, the original kernel (220) has a size of 3x3, and the output feature map (230) has a size of 4x4, but this is not limited thereto, and the artificial intelligence learning model (NN of FIG. 2) can be implemented with feature maps and kernels of various sizes. In addition, the values defined in the input feature map (210), the original kernel (220), and the output feature map (230) are all exemplary values, and embodiments according to the present disclosure are not limited thereto.
[0095] The original kernel (220) can perform a convolution operation by sliding in a 3x3 window unit on the input feature map (210). The convolution operation can represent an operation for obtaining each feature data of the output feature map (230) by summing all values obtained by multiplying each feature data of a window of the input feature map (210) and each weight value of a corresponding position in the original kernel (220).
[0096] In other words, a convolution operation between one input feature map (210) and one original kernel (220) can be processed by repeatedly performing multiplication of extracted data of the input feature map (210) and corresponding weight values of the original kernel (220) and summing of the multiplication results, and an output feature map (230) can be generated as a result of the convolution operation.
[0097] FIG. 4 is a flowchart illustrating an artificial intelligence-based load and heavy equipment collision avoidance method according to the present disclosure.
[0098] According to one embodiment of the present disclosure, an artificial intelligence-based load and heavy equipment collision prevention method comprises a step (S100) in which a processor implements a learned artificial neural network, a step (S200) in which a camera senses an object and a background environment, a step (S300) in which a communication module transmits image data sensed from the camera to the processor, a step (S400) in which the processor extracts dangerous objects and people including the load and heavy equipment from the image data, and a step (S500) in which the processor prevents the risk of collision by estimating the distance between the dangerous objects and the people.
[0099] FIGS. 5 to 7 are drawings illustrating examples of the operation of artificial intelligence-based load and heavy equipment collision avoidance according to the present disclosure.
[0100] Referring to 510 to 710 of FIGS. 5 to 7, the load and the salvage are indicated as the load.
[0101] In industrial settings, heavy equipment has been frequently used to transport loads, resulting in numerous fatalities and collisions with surrounding objects. Most collisions involve the heavy equipment body and surrounding objects / people, or the load and surrounding objects / people.
[0102] Despite the advancement of these cameras, accidents resulting in human casualties and collisions with cargo are still commonplace. Despite the rapid advancement of artificial intelligence technology, due to the limitations of the computational capacity of artificial intelligence models, only human recognition systems are currently used to assist drivers with notifications. While a control system using a third-person camera exists, a collision avoidance system is absent.
[0103] The core ideas of the present invention are 1. a method for preventing collisions between loads and heavy equipment using artificial intelligence technology, 2. a system configuration and operating method for applying the collision method, and 3. a system for preventing collisions between loads and heavy equipment using the method and artificial intelligence technology.
[0104] 1. Method for preventing collisions between loads and heavy equipment using artificial intelligence and a system applying the method.
[0105] The core feature of the present disclosure is to detect and prevent collisions with surrounding objects and people when loads and heavy equipment are moving.
[0106] The implementation method utilizes AI technology to accurately extract the shape of highly variable loads using one or more wearable / portable wireless cameras, while simultaneously detecting surrounding objects and people. Camera-based distance detection AI technology is then used to measure the distance between the extracted load and the surrounding environment, thereby detecting and preventing collisions.
[0107] As another implementation method, artificial intelligence technology that detects real-time loads through the fusion of multiple sensors such as cameras and LiDAR, and artificial intelligence technology that detects surrounding objects and people, extracts specific objects, and then extracts distance data through calibration with distance detection sensors and cameras, and estimates the distance between surrounding objects, people, and loads to detect and prevent collisions in advance.
[0108] The part that wears / attaches / mounts the above sensors is a collision detection system that installs and mounts cameras or sensors on third parties other than heavy equipment operators, third-party mobile objects (motorcycles, bicycles, etc.), or fixed objects such as buildings, pipes, or exterior walls. The third parties here could be signalmen, guides, workers, or work supervisors at industrial sites.
[0109] In an exemplary embodiment, the load is extracted through an artificial intelligence in the field of segmentation based on camera-based unsupervised learning that finds variable loads.
[0110] According to an exemplary embodiment, the surrounding environment is recognized through artificial intelligence in the field of segmentation that detects surrounding objects and people based on camera-based supervised learning.
[0111] In an exemplary embodiment, the distance to all objects on the screen is estimated using an artificial intelligence that estimates distances based on camera-based supervised learning.
[0112] According to an exemplary embodiment, the pixel-by-pixel distance estimate values of the extracted object and the load are extracted and the distance between the two is estimated using these.
[0113] In an exemplary embodiment, the information is transmitted in real time to heavy equipment operators and third parties, and the risk level is notified after setting specific thresholds in multiple sections.
[0114] In an exemplary embodiment, sensors mounted on a third party communicate with the driver in real time, either wired or wirelessly. They transmit real-time video information along with hazard warnings.
[0115] 2. Camera-based unsupervised learning payload detection AI technology
[0116] The payload has a highly variable form, making it suitable for unsupervised learning rather than supervised learning. This disclosure proposes an AI technology that combines a self-designed cost function (loss function) with existing prompt engineering techniques based on surrounding objects.
[0117] 3. Distance data collection device
[0118] Data collection is possible by applying refined artificial intelligence technology along with existing camera and LiDAR calibration techniques.
[0119] According to the present disclosure, image data can be collected by capturing images with a camera. Furthermore, data can be collected by fusing image data captured by the camera with distance data measured using LiDAR.
[0120] We have developed our own tool for data collection, which can be used to implement distance estimation artificial intelligence model learning based on supervised learning.
[0121] 3. Collision detection between loads and heavy equipment vehicles and surrounding objects and people from a third-person perspective.
[0122] Portable / wearable sensors can be attached or worn / mounted on third parties or buildings to detect collisions between heavy equipment / loads and surrounding objects and people in real time.
[0123] The present disclosure can detect and prevent collisions in advance by estimating real-time relative speed and distance through tracking.
[0124] 4. How to use the sensor
[0125] Methods of utilizing the sensor of the present disclosure include a method in which a person wears it (wearable), a method in which a person carries it (portable), a method in which a person fixes it to a moving object such as a bicycle / motorcycle that a person rides, or a method in which the sensor is fixed to a surrounding wall or pipe.
[0126] According to the exemplary embodiment of the present disclosure as described above, the collision avoidance system for loads (salvageable objects) and heavy equipment using artificial intelligence can significantly reduce accidents caused by blind spots and lack of visibility by using third parties or fixed objects such as surrounding buildings / pipes.
[0127] In addition, according to the present disclosure, a great advantage can be achieved by providing video information through a real-time monitoring system for areas where the driver cannot see.
[0128] Furthermore, the present disclosure can significantly reduce safety accidents by enabling AI to proactively detect and alert drivers of accidents they may not otherwise be aware of. Of course, the scope of the present disclosure is not limited by these benefits.
[0129] Furthermore, the present disclosure can significantly reduce safety accidents by enabling AI to proactively detect and alert drivers of accidents they may not otherwise be aware of. Of course, the scope of the present disclosure is not limited by these benefits.
[0130] Although the present disclosure has been described with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will understand that various modifications and equivalent other embodiments are possible from the disclosure. Therefore, the true technical protection scope of the present disclosure should be determined by the technical spirit of the appended claims.
[0131] FIG. 8 is a diagram illustrating a configuration of an artificial intelligence-based load and heavy equipment collision avoidance device according to the present disclosure.
[0132] Referring to FIG. 8, the artificial intelligence-based load and heavy equipment collision avoidance device (100) includes an input module (110), a sensor module (120), a processor (130), a display module (140), a memory (150), a communication module (160), and a camera module (170).
[0133] The input module (110) acquires data.
[0134] The sensor module (120) senses data. The sensor module (120) includes a distance sensor. The distance sensor measures the distance between heavy equipment and a person, and the distance between a load and a person.
[0135] The processor (130) performs an artificial intelligence-based load-bearing equipment collision avoidance method according to the process.
[0136] The processor (130) extracts images of hazardous materials including the load and the heavy equipment and images of people from the image using an artificial intelligence model, measures the distance between the extracted images of hazardous materials and the images of people, determines the risk of collision based on the measured distance, and outputs a warning message if there is a risk of collision.
[0137] The processor (130) determines that there is a risk of collision if the distance between the human image and the hazardous material image is less than 1 m.
[0138] The processor (130) determines the separation distance and safety level based on the following criteria.
[0139] If the separation distance exceeds 2m, it is considered safe.
[0140] If the separation distance is 1 to 2 m, it is considered that there is a risk of collision.
[0141] If the separation distance is 0.5 to 1 m, it is considered a risk of collision.
[0142] If the separation distance is less than 0.5 m, it is considered very dangerous.
[0143] The display module (140) displays a graphic image according to a control command from the processor (130).
[0144] The memory (150) stores at least one process for performing an operation for preventing collisions between artificial intelligence-based loads and heavy equipment according to the present disclosure, stores user input and data, and stores processing code for implementing an artificial intelligence model and parameters of learned artificial intelligence.
[0145] The communication module (160) transmits and receives data with an external device (200).
[0146] Here, the external device (200) includes an external device such as a smartphone, PC, laptop, tablet PC, etc.
[0147] The camera module (170) captures images of the front.
[0148] The camera module (170) photographs a subject in front according to a control command from the processor (130).
[0149] The processor (130) extracts an image of a hazardous material including the cargo from the image using an artificial intelligence model learned based on unsupervised learning, and the size of the cargo has a variable characteristic. A detailed description thereof is provided in FIG. 11.
[0150] The processor (130) recognizes heavy equipment from the image using segmentation, provides at least one point (pixel point) based on the recognized heavy equipment to the artificial intelligence model, predicts the shape of the heavy equipment using the artificial intelligence model, and extracts an image of the heavy equipment from the prediction result. A detailed description thereof is provided in FIG. 20.
[0151] The processor (130) recognizes a section where a load is loaded for each piece of heavy equipment from image information using an artificial intelligence model, selects 1 to N pixel coordinates around the recognized section, trains the artificial intelligence model based on unsupervised learning using the selected pixel coordinates as prompt information, operates the training 1 to N times to select an artificial intelligence model with the highest reliability, predicts the shape of the load based on the selected artificial intelligence model, and extracts an image of the load from the prediction result. A detailed description thereof is provided in FIG. 21.
[0152] The camera (170) is installed in a position capable of capturing at least one of heavy equipment, loads, surrounding objects, and people from a third-person perspective. A detailed description thereof is provided in Fig. 18.
[0153] The collection device includes a camera (170).
[0154] The collection device can collect data by using a camera (170) or by fusing a camera (170) and LiDAR. A detailed description of this is provided in Fig. 18.
[0155] The processor (130) generates a blind spot image from the image using the artificial intelligence model, and transmits the generated blind spot image to the external device (200) via the communication unit (160). A detailed description thereof is provided in Fig. 14.
[0156] The processor (130) measures the relative velocity between the extracted image of the hazardous material and the image of the person, determines the risk of collision based on the measured relative velocity, and outputs a warning message if there is a risk of collision. A detailed description thereof is provided in Fig. 10.
[0157] The processor (130) extracts images of hazardous materials, including heavy equipment, and images of people from the image using an artificial intelligence model trained based on supervised learning. A detailed description of this is provided in Fig. 10.
[0158] If the load includes a protruding object image, the processor (130) extracts a hazardous material image by processing the load using a prompt technology. A detailed description thereof is provided in FIG. 11.
[0159] However, the components illustrated in FIG. 8 are not essential for implementing the present disclosure according to the present disclosure, and thus the present disclosure described in this specification may have more or fewer components than the components listed above.
[0160] The communication module (160) may include one or more components that enable communication with an external device, and may include, for example, at least one of a broadcast reception module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0161] The input module (110) is for inputting image information (or signal), audio information (or signal), data, or information input from a user, and may include at least one camera, at least one microphone, and at least one user input unit. Voice data or image data collected by the input module (110) may be analyzed and processed into a user control command.
[0162] The display module (140) displays (outputs) information processed in the present disclosure. For example, the present disclosure may display execution screen information of a running application program (e.g., an application), or UI (User Interface) or GUI (Graphical User Interface) information based on such execution screen information.
[0163] The memory (150) can store data supporting various functions of the present disclosure, programs for the operation of the control unit, input / output data (e.g., music files, still images, videos, etc.), and a plurality of application programs (application programs or applications) driven by the artificial intelligence-based user behavior pattern analysis device (100), data for the operation of the device, and commands. At least some of these application programs can be downloaded from an external server via wireless communication.
[0164] The memory (150) may include at least one type of storage medium among a flash memory type, a hard disk type, an SSD (Solid State Disk type), an SDD (Silicon Disk Drive type), a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. In addition, the memory (150) may be a database connected by wire or wirelessly, although separate from the present disclosure, and may be implemented as a database system.
[0165] The processor (130) may be implemented as at least one core, a memory storing data for an algorithm for controlling the operation of components within the present disclosure or a program reproducing the algorithm, and at least one processor (not shown) that performs the aforementioned operations using the data stored in the memory. In this case, the memory and the processor may be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.
[0166] In addition, the processor (130) can control any one or a combination of the components described above to implement various embodiments according to the present disclosure described in FIGS. 1 to 22.
[0167] At least one component may be added or deleted to correspond to the performance of the components illustrated in Fig. 8. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered to correspond to the performance or structure of the system.
[0168] Meanwhile, each component illustrated in FIG. 8 represents software and / or hardware components such as a Field Programmable Gate Array (FPGA) and an Application Specific Integrated Circuit (ASIC).
[0169] FIG. 9 is a flowchart illustrating an artificial intelligence-based load and heavy equipment collision avoidance method according to the present disclosure.
[0170] The present disclosure is performed by an artificial intelligence-based load and heavy equipment collision avoidance device (100) or a processor (130) of the artificial intelligence-based load and heavy equipment collision avoidance device (100).
[0171] Referring to FIG. 9, the processor (130) controls the camera (170) to capture a front image (S210).
[0172] The processor (130) extracts images of hazardous materials including loads and heavy equipment and images of people from the image using an artificial intelligence model (S220).
[0173] The processor (130) measures the distance between the extracted image of the hazardous material and the image of the person (S230).
[0174] The processor (130) determines the risk of collision based on the measured separation distance (S240).
[0175] If there is a risk of collision, the processor (130) outputs a warning message (S250).
[0176] If there is a risk of collision, the processor (130) controls the display module (140) to output a warning message.
[0177] The processor (130) extracts an image of a hazardous material including the load from the image using an artificial intelligence model learned based on unsupervised learning.
[0178] The processor (130) generates a blind spot image from the image using the artificial intelligence model, and transmits the generated blind spot image to the external device (200) through the communication unit (160).
[0179] FIG. 10 is a drawing illustrating the concept of the present disclosure according to the present disclosure.
[0180] As illustrated in 1010 of FIG. 10, the present disclosure is an artificial intelligence-based load and heavy equipment collision avoidance device (100) equipped with a third-person view camera (170) worn by a signalman.
[0181] The present disclosure (100) transmits blind spot images in real time to a tablet PC (200) of a heavy equipment operator, detects loads, the heavy equipment itself, the surrounding environment, and people, and transmits a warning message when it is determined that there is a risk of collision.
[0182] According to the present disclosure, the size of a load can be measured, a warning message notifying a collision with a surrounding object can be output, and a blind spot image can be transmitted to an external device (200) in real time.
[0183] An embodiment of outputting a warning message based on relative speed is described.
[0184] The processor (130) measures the relative speed between the extracted image of the hazardous material and the image of the person, determines the risk of collision based on the measured relative speed, and outputs a warning message if there is a risk of collision.
[0185] For example, when the speed of the image of a hazardous material is 10 m / s and the speed of the image of a person is 4 m / s, the processor (130) outputs a warning message if the distance between the image of a hazardous material and the image of a person is less than 1 m when considering the relative speed.
[0186] The processor (130) determines the distance and safety between the image of a hazardous material and the image of a person, taking into account the relative speed, based on the following criteria.
[0187] If the distance is more than 2m, it is considered safe.
[0188] If the distance is 1 to 2 m, it is considered that there is a risk of collision.
[0189] If the distance is 0.5 to 1 m, it is considered a risk of collision.
[0190] If the distance is less than 0.5 m, it is considered very dangerous.
[0191] The processor (130) extracts images of hazardous materials including heavy equipment and images of people from the image using an artificial intelligence model learned based on supervised learning.
[0192] The processor (130) includes a GPU board.
[0193] The GPU board is equipped with supervised heavy equipment data and processes the load with prompt technology.
[0194] Here, heavy equipment data includes forklift data and crane data.
[0195] The GPU board recognizes protruding loads such as pipes by processing them as points.
[0196] The processor (130) can measure the distance between objects by measuring the size of supervised-learned heavy equipment and unsupervised-learned, variable-sized loads.
[0197] The processor (130) can measure the distance to the load according to the rotation angle of the heavy equipment based on the learned size of the heavy equipment and the estimated size of the load.
[0198] For example, if heavy equipment rotates 90 degrees, the distance between a person and a load can be measured as 1.5 m. In this case, the processor (130) determines it to be safe.
[0199] When the heavy equipment rotates 180 degrees, the distance between the person and the load can be measured as 0.99 m. In this case, the processor (130) determines that there is a danger and outputs a warning message.
[0200] The main technical features of this disclosure are described below.
[0201] First, it is a technology that uses artificial intelligence to detect jack materials with highly variable shapes.
[0202] This disclosure is an artificial intelligence technology that recognizes and utilizes Cost Function (Loss function) and surrounding objects, and Prompt Engineering technology.
[0203] Here, since the payload has a variable shape, unsupervised learning is applied rather than supervised learning.
[0204] Second, it detects collisions between cargo and heavy equipment vehicles and surrounding objects and people from a third-person perspective.
[0205] In the conventional technology, a server (main body) is installed on a vehicle body (heavy equipment) to detect collisions between the heavy equipment and surrounding objects, whereas the present invention is in the form of a wearable device that a signalman wears to detect collisions between a specific object (load, heavy equipment) and surrounding objects from a third-person viewpoint and to notify of the risk of collision.
[0206] The present invention can detect and prevent collisions in advance by estimating relative speed and distance in real time through tracking.
[0207] Explains how to install and use it.
[0208] The present invention can be fixed to a wearable device, a portable device, a bicycle, a motorcycle, or a moving object, or installed on a surrounding wall or pipe, or a surrounding object.
[0209] Third, it is a system that uses only one camera.
[0210] The present invention can detect all objects using only one camera, estimate their relative speed and velocity, and predict the risk of collision in advance.
[0211] FIG. 11 is a diagram illustrating an embodiment of unsupervised learning according to the present disclosure.
[0212] Figure 11 includes Figures 11(a) and 11(b).
[0213] 1110 of Fig. 11(a) is a diagram illustrating an example of measuring distance using an artificial intelligence model learned based on unsupervised learning.
[0214] Figure 11(b) is a diagram illustrating an example of performing clustering based on unsupervised learning.
[0215] Unsupervised learning is a method in which a computer (algorithm) learns from data without including labels. Since labels do not exist, a model is created by specifying specific patterns or rules.
[0216] In supervised learning, training data and labels play the role of X and Y, respectively, while unsupervised learning infers results only from data.
[0217] That is, the goal is to find hidden patterns in X through a series of rules (f(x)).
[0218] Because it operates on data without training data, there is no target value, so unlike supervised learning, pretraining is not required. Therefore, unlike supervised learning, the lack of labels makes performance evaluation difficult.
[0219] Unsupervised learning mainly includes clustering, outlier detection (anomaly detection), and dimensionality reduction.
[0220] As shown in 1120 of Figure 11(b), clustering is an algorithm that divides similar data into several clusters based on input data.
[0221] Data is classified based on the characteristics of the input data. Data classified into the same group have similar properties (location, mean, variance, etc.). Conversely, data classified into different groups have different properties.
[0222] 1120 in Fig. 11(b) is the result of classification through a series of rules without prior learning.
[0223] Three types of clusters were classified based on similar data characteristics. Because clustering has no universal definition, different algorithms can classify different clusters.
[0224] According to the present disclosure, the size of the load is not fixed in shape and has variable characteristics.
[0225] The processor (130) extracts an image of a hazardous material including the load from the image using an artificial intelligence model learned based on unsupervised learning.
[0226] When detecting surrounding objects and people, the processor (130) uses a segmentation technique based on map learning using a camera.
[0227] When detecting a load, the processor (130) uses a segmentation technique based on unsupervised learning using a camera.
[0228] Describes distance estimation for objects.
[0229] The processor (1300) estimates the distance to all objects on the screen using map learning-based artificial intelligence technology using a camera.
[0230] Here, the size of the load has a variable characteristic.
[0231] Specifically, the size of the payload has a variable characteristic, and unsupervised learning is more suitable than supervised learning.
[0232] This disclosure uses a combination of Cost function (Loss Function) and Prompt Engineering techniques based on surrounding objects.
[0233] The processor (130) extracts a hazardous material image by processing the load with a prompt technology when the load includes a protruding object image.
[0234] Figure 12 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0235] As illustrated in 1210 of FIG. 12, according to the present disclosure, real-time relative distance and relative speed can be measured using only one camera.
[0236] According to the present disclosure, a lightweight artificial intelligence solution can be provided for utilizing low-cost on-device artificial intelligence hardware.
[0237] For example, performance can be improved by 3 to 5%, speed can be increased by 6 to 14 times, computational complexity can be reduced by 0.3 to 0.7 times, and hardware size can be reduced.
[0238] Figure 13 is a diagram showing a system configuration diagram in which a signal number according to the present disclosure is worn according to the present invention.
[0239] As shown in 1310 of Figure 13, it can be divided into a signalman wearing appearance and a forklift system configuration diagram.
[0240] The signalman wears a wearable belt equipped with a camera (170).
[0241] Here, the wearable belt may include an artificial intelligence-based load and heavy equipment collision avoidance device (100).
[0242] Install a tablet PC (200) in the driver's seat of the heavy equipment.
[0243] When heavy equipment moves, a signalman wearing a wearable belt accompanies it.
[0244] The signalman captures the blind spot caused by the load with a camera (170).
[0245] When a dangerous situation occurs, the processor (130) predicts a collision of heavy equipment or load in advance and transmits a warning message to the driver's tablet PC (200) in real time.
[0246] Figure 14 is a drawing showing a system configuration diagram in which the present invention according to the present disclosure is installed on heavy equipment.
[0247] As shown in 1410 of Figure 14, it can be divided into a signalman wearing appearance and a forklift system configuration diagram.
[0248] The signalman can wear the camera (170) in a wearable form.
[0249] Heavy equipment may include artificial intelligence-based load and heavy equipment collision avoidance devices (100).
[0250] The processor (130) can generate a blind spot image in real time through an artificial intelligence model and transmit it.
[0251] The processor (130) can measure the size of the load.
[0252] When a collision risk is predicted, the processor (130) outputs a visual notification message and an auditory notification message.
[0253] The processor (130) generates a blind spot image from the image using the artificial intelligence model and transmits the generated blind spot image to the external device (200) via the communication unit (160). Here, the external device (200) includes a driver's tablet.
[0254] FIG. 15 is a drawing showing an example of comparing the prior art and the present invention in a tower crane according to the present disclosure.
[0255] Figure 15 includes Figures 15(a) and 15(b).
[0256] 1510 of Fig. 15(a) illustrates a blind spot of a conventional tower crane.
[0257] 1520 of FIG. 15(b) illustrates a blind spot of the present invention in a tower crane.
[0258] Comparing 1510 of Fig. 15(a) and 1520 of Fig. 15(b), it can be seen that the blind spot of the present invention is significantly reduced compared to the prior art.
[0259] According to the present disclosure, the size of the blind spot area is significantly reduced compared to the prior art, thereby lowering the possibility of a safety accident, thereby improving user convenience.
[0260] FIG. 16 is a drawing showing an example of comparing the prior art and the present invention in a forklift according to the present disclosure.
[0261] Figure 16 includes Figures 16(a) and 16(b).
[0262] 1610 of Fig. 16(a) illustrates a blind spot of a conventional forklift.
[0263] 1620 of FIG. 16(b) illustrates a blind spot of the present invention in a forklift.
[0264] Comparing 1610 of Fig. 16(a) and 1620 of Fig. 16(b), it can be seen that the blind spot of the present invention is significantly reduced compared to the prior art.
[0265] According to the present disclosure, the size of the blind spot area is significantly reduced compared to the prior art, thereby lowering the possibility of a safety accident, thereby improving user convenience.
[0266] Figure 17 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0267] As illustrated in 1710 of FIG. 17, according to the present disclosure, there is an advantage in that the size of an object can be accurately measured in real time to prevent collision with surrounding objects, even in the case of heavy equipment work that transports variable loads and salvage objects each time.
[0268] FIG. 18 is a drawing illustrating an embodiment in which a third-party viewpoint camera is installed according to the present disclosure.
[0269] Figure 18 includes Figures 18(a), 18(b), 18(c), and 18(d).
[0270] Figure 18(a) is a drawing showing a camera installed in a body cam format.
[0271] Figure 18(b) is a drawing showing a camera being attached to and detached from heavy equipment.
[0272] Figure 18(c) is a drawing showing a camera attached to an electric bicycle.
[0273] Figure 18(d) is a drawing showing a camera attached to a robot.
[0274] As illustrated in 1810 of FIG. 18(a), the camera (170) can be installed in a hands-free wearable body camera fashion. In this case, there is an advantage in that the user can conveniently wear and move the camera in a clip-type manner.
[0275] As illustrated in 1820 of FIG. 18(b), the camera (170) can be attached or detached from heavy equipment. In this case, there is an advantage in that the user can easily attach or detach the camera from the main equipment and perform high-altitude work or delicate work.
[0276] As illustrated in 1830 of FIG. 18(c), a camera (170) can be attached to an electric bicycle. A user can drive an electric bicycle equipped with the camera (170). This provides the advantage of allowing the user to move flexibly within a wide workspace.
[0277] As illustrated in 1840 of Fig. 18(d), a camera (170) can be attached to the robot. The robot can move while the camera (170) is installed. In this case, customization is possible, such as for safety distances, collisions with hazardous objects and people, and oil leak detection, and the robot can move flexibly in a wide work space.
[0278] The camera (170) is installed in a position where it can capture at least one of heavy equipment, load, surrounding objects, and people from a third-person viewpoint.
[0279] The artificial intelligence-based load and heavy equipment collision avoidance device (100) of the present invention can be implemented as an integrated unit with a camera (170) or in a form including a camera (170).
[0280] Additional information on portable, wearable cameras and sensors is provided.
[0281] Cameras and sensors are installed in a third-person perspective (to detect heavy equipment, surrounding objects, and people) rather than a first-person perspective (the perspective of the heavy equipment operator).
[0282] The processor (130) can measure the relative speed and relative distance of at least one of heavy equipment, load, surrounding objects and people in real time to detect and warn in advance.
[0283] Explains the case where there are multiple cameras.
[0284] When the first camera, the second camera, and the third camera are installed in appropriate locations, each camera can capture the first image, the second image, and the third image.
[0285] The processor (130) synthesizes the first image, the second image, and the third image to generate a composite image.
[0286] The processor (130) generates a blind spot image from a composite image and measures the relative distance and relative distance of at least one of heavy equipment, load, surrounding objects, and people, and if a possibility of danger is recognized, a warning can be issued in advance.
[0287] An artificial intelligence-based load and heavy equipment collision avoidance device (100) includes a collection device. The collection device includes a camera (170) and a lidar.
[0288] Lidar includes a laser transmitter, laser receiver, GPS receiver, and IMU (Inertial Measurement Unit).
[0289] The laser transmitter fires a laser pulse.
[0290] The laser receiver senses the reflected laser beam.
[0291] GPS receivers provide location information.
[0292] The IMU measures the attitude and movement of the device.
[0293] The collection device can collect data by using a camera (170) or by fusing a camera (170) and LiDAR.
[0294] LiDAR (Light Detection and Ranging) is a technology that uses laser beams to measure the distance to objects and collect three-dimensional terrain information.
[0295] The principle of lidar works as follows:
[0296] First, a laser pulse is fired from the LiDAR device toward the target.
[0297] Second, the laser beam hits the target object and reflects.
[0298] Third, the reflected laser beam reaches the LiDAR receiver.
[0299] Fourth, measure the time it takes for the laser to fire and the reflected signal to return.
[0300] Fifth, calculate the distance to the object using time and the speed of light.
[0301] Through this process, numerous points are collected, forming a 3D data set called a point cloud. Because LiDAR offers high accuracy and precision, it can provide more efficient and reliable results than traditional surveying and data collection methods.
[0302] The processor (130) can measure the relative speed and relative distance of at least one of heavy equipment, load, surrounding objects, and people from the image captured by the camera (170).
[0303] The processor (130) can measure the relative speed and relative distance of at least one of heavy equipment, load, surrounding objects, and people by fusing the image captured by the camera (170) and the distance measured by the lidar.
[0304] The processor (130) can fuse the image captured by the camera (170) and the distance measured by the lidar by applying different weights.
[0305] For example, when the weather is clear, the weight of the image captured by the camera (170) may be w1 = 0.66, and the weight of the distance measured by the lidar may be w2 = 0.33. Here, w1 + w2 = 1. When the weather is clear, the weight of the image captured by the camera (170) is set high because the resolution of the image is high.
[0306] When the weather is cloudy, the weight of the image captured by the camera (170) may be w1 = 0.33, and the weight of the distance measured by the lidar may be w2 = 0.66. Here, w1 + w2 = 1. When the weather is cloudy, the weight of the image captured by the camera (170) is set low because the resolution of the image is low.
[0307] According to the present disclosure, the distance between a user and a hazard can be measured using a camera alone or a combination of a camera and a lidar, thereby providing highly reliable results.
[0308] Figure 19 is a drawing illustrating the effect of the present disclosure according to the present disclosure.
[0309] As shown in 1910 of FIG. 19, the prior art had the disadvantages of causing collision accidents, loss of life, blind spot accidents, and material damage.
[0310] According to the present disclosure, there are advantages in preventing material damage, preventing human casualties, and enabling efficient process management.
[0311] Figure 20 is a drawing illustrating an example of a load shape extraction embodiment in the present disclosure.
[0312] Figure 20 includes Figures 20(a), 20(b) and 20(c).
[0313] Figure 20(a) is a drawing illustrating recognition of heavy equipment.
[0314] Figure 20(b) is a diagram illustrating providing at least one point to an artificial intelligence model based on a recognized heavy equipment.
[0315] Figure 20(c) is a drawing illustrating extraction of the shape of a load.
[0316] As shown in Fig. 20(a), the processor (130) recognizes heavy equipment from the image using segmentation.
[0317] As shown in Fig. 20(b), the processor (130) provides at least one point (pixel point) based on the recognized heavy equipment to the artificial intelligence model.
[0318] As shown in Fig. 20(c), the processor (130) predicts the shape of the load using the artificial intelligence model and extracts an image of the load from the prediction result.
[0319] FIG. 21 is a diagram illustrating a load shape extraction algorithm according to the present disclosure.
[0320] As illustrated in Fig. 21, this algorithm is performed by a processor (130).
[0321] The processor (130) acquires an image from a camera (S310).
[0322] The processor (130) operates an artificial intelligence model (S320).
[0323] Here, the artificial intelligence model includes a heavy equipment recognition artificial intelligence model.
[0324] The processor (130) recognizes the section where the heavy equipment load is loaded from the image (S330).
[0325] The processor (130) selects 1 to N pixel coordinates around the recognized section (S340).
[0326] The processor (130) trains the artificial intelligence model based on unsupervised learning using the selected pixel coordinates as prompt information (S350).
[0327] The processor (130) selects the artificial intelligence model with the highest reliability by performing the above learning 1 to N times (S360).
[0328] The processor (130) predicts the shape of the load based on the selected artificial intelligence model and extracts an image of the load from the prediction result (S370).
[0329] FIG. 22 is a diagram illustrating a configuration of an artificial intelligence-based load and heavy equipment collision avoidance device according to the present disclosure.
[0330] Referring to FIG. 22, the present disclosure includes a device (1600). The device (1600) may include a memory (1602), a processor (1603), a transceiver (1604), and a peripheral device (1601). In addition, as an example, the device (1600) may further include other configurations and is not limited to the above-described embodiment.
[0331] More specifically, the device (1600) of FIG. 22 may be an exemplary hardware / software architecture, such as an NDN device, an NDN server, a content router, etc. In this case, as an example, the memory (1602) may be a non-removable memory or a removable memory. In addition, as an example, the peripheral device (1601) may include a display, GPS, or other peripheral devices, and is not limited to the above-described embodiment.
[0332] In addition, as an example, the above-described device (1600) may include a communication circuit such as the transceiver (1604), and may perform communication with an external device based thereon.
[0333] Additionally, as an example, the processor (1603) may be at least one of a general-purpose processor, a digital signal processor (DSP), a DSP core, a controller, a microcontroller, ASICs (Application Specific Integrated Circuits), FPGA (Field Programmable Gate Array) circuits, any other type of integrated circuit (IC), and one or more microprocessors associated with a state machine. In other words, it may be a hardware / software configuration that performs a control role for controlling the above-described device (1600).
[0334] At this time, the processor (1603) may execute computer-executable instructions stored in the memory (1602) to perform various essential functions of the present disclosure. For example, the processor (1603) may control at least one of signal coding, data processing, power control, input / output processing, and communication operations. In addition, the processor (1603) may control the physical layer, the MAC layer, and the application layers. In addition, for example, the processor (1603) may perform authentication and security procedures in the access layer and / or the application layer, and is not limited to the above-described embodiment.
[0335] For example, the processor (1603) can communicate with other devices via the transceiver (1604). For example, the processor (1603) can control a node to communicate with other nodes via a network through the execution of computer-executable instructions. That is, the communication performed in the present disclosure can be controlled. For example, the other nodes can be NDN servers, content routers, and other devices. For example, the transceiver (1604) can transmit RF signals via an antenna and transmit signals based on various communication networks.
[0336] In addition, as an example, MIMO technology, beamforming, etc. can be applied as antenna technology, and are not limited to the above-described embodiment. In addition, the signal transmitted and received through the transceiver (1604) can be modulated and demodulated and controlled by the processor (1603), and are not limited to the above-described embodiment.
[0337] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0338] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0339] Although the embodiments described above have been described with limited drawings, those skilled in the art will recognize that various modifications and variations are possible based on the above teachings. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents. Therefore, other implementations, other embodiments, and equivalents of the claims also fall within the scope of the following claims.
Claims
1. In a device for preventing collisions between loads and heavy equipment based on artificial intelligence, A camera that captures images of the front; A communication module that transmits and receives data with external devices; A memory storing at least one process for preventing collisions between the artificial intelligence-based load and heavy equipment; and Includes a processor that performs operations according to the above process, The above processor, Using an artificial intelligence model, images of hazardous materials including the load and heavy equipment and images of people are extracted from the images, Measure the distance between the extracted hazardous material image and the extracted human image, Based on the measured distance above, determine the risk of collision, If there is a risk of collision as identified above, a warning message is output, Using an artificial intelligence model, the section where the heavy equipment is loaded is recognized from the image above, Select 1 to N pixel coordinates around the above recognized section, The artificial intelligence model is trained based on unsupervised learning using the selected pixel coordinates as prompt information, The above learning is performed 1 to N times to select the artificial intelligence model with the highest reliability, Predict the shape of the load based on the selected artificial intelligence model, Extracting an image of the load from the predicted result, AI-based load and heavy equipment collision avoidance device.
2. In paragraph 1, The above processor extracts an image of a hazardous material including the load from the image using an artificial intelligence model learned based on the above unsupervised learning, The size of the above load has variable characteristics, AI-based load and heavy equipment collision avoidance device.
3. In the first paragraph, the processor Recognize the heavy equipment from the image using segmentation, Predicting the shape of the load by providing at least one pixel point based on the recognized heavy equipment to the artificial intelligence model. AI-based load and heavy equipment collision avoidance device.
4. In the first paragraph, the camera, Installed in a position where at least one of the above heavy equipment, the above load, surrounding objects and people can be captured from a third-person viewpoint. AI-based load and heavy equipment collision avoidance device.
5. In the first paragraph, the processor, Using the above artificial intelligence model, a blind spot image is generated from the above image, Transmitting the generated blind spot image to the external device through the communication module, AI-based load and heavy equipment collision avoidance device.
6. In the first paragraph, the processor, Measure the relative speed between the extracted hazardous material image and the extracted human image, Based on the measured relative speed, the collision risk is determined, If there is a risk of collision as identified above, output the warning message. AI-based load and heavy equipment collision avoidance device.
7. In the first paragraph, the processor, Using the artificial intelligence model, an image of a hazardous material including heavy equipment and an image of a person are extracted from the image. AI-based load and heavy equipment collision avoidance device.
8. In an artificial intelligence-based load and heavy equipment collision avoidance method performed by the processor of the device, Step of capturing the front image; A step of extracting an image of a hazardous material including the load and the heavy equipment and an image of a person from the image using an artificial intelligence model; A step of measuring the distance between the extracted hazardous material image and the extracted human image; A step of checking the risk of collision based on the measured separation distance; and If there is a risk of collision identified above, a step of outputting a warning message is included, The above method, A step of recognizing a section in which heavy equipment is loaded using the artificial intelligence model from the image; A step of selecting 1 to N pixel coordinates around the above recognized section; A step of training the artificial intelligence model based on unsupervised learning using the selected pixel coordinates as prompt information; A step of selecting an artificial intelligence model with the highest reliability by performing the above learning 1 to N times; A step of predicting the shape of the load based on the selected artificial intelligence model; and Further comprising a step of extracting an image of the load from the predicted result. AI-based collision avoidance method for loads and heavy equipment.
Citation Information
Patent Citations
Crane work monitoring system, crane work monitoring method, dangerous state determination device, and program
JP2020093890A
Crane monitoring device and crane monitoring method as well as overhead crane
JP2022061682A
Construction assistance system, construction assistance method, height calculation method, and construction assistance program
JP2022175364A
Throwing management system and throwing management method for sand carrier
JP2023121309A
Sealant film for cell pouch having barrier properties using hygroscopicity of inorganic materals, cell pouch including the same and method for preparing the same
KR102665239B1
Cited By
Unmanned aerial vehicle patrol intelligent early warning method, device and equipment based on construction safety distance and medium
CN121415291A