Electronic apparatus and method for detecting vision-based equipment anomaly and generating natural language report

KR103000031B1Active Publication Date: 2026-08-05D-TAP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
D-TAP CO LTD
Filing Date
2025-11-04
Publication Date
2026-08-05

Smart Images

  • Figure R1020250164573_ABST
    Figure R1020250164573_ABST
Patent Text Reader

Abstract

An electronic device for detecting equipment anomalies based on vision images and generating natural language reports according to the present disclosure comprises: a vision camera for capturing real-time vision images of industrial equipment; an anomaly detection model for receiving the real-time vision images and hierarchically detecting anomalies of the industrial equipment; a natural language report generation model for generating a report by processing the detected anomalies of the industrial equipment in natural language; a memory storing at least one process for performing operations of hierarchically detecting anomalies of the industrial equipment and generating a user-customized report for the detected anomalies; and at least one processor for performing the operations according to the process. The at least one processor may be configured to receive real-time vision image data of the industrial equipment, generate a spatial-temporal feature map of the vision image data through the anomaly detection model to first detect anomaly candidates, secondarily detect whether an anomaly exists in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area, and automatically generate a natural language-based report based on a preset template for the area determined to be an anomaly and the type of defect.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to an electronic device and method for performing monitoring based on a vision image, and more specifically, to an electronic device and method for generating a vision image-based facility anomaly detection and natural language report that improves the efficiency of safety management by hierarchically detecting anomaly areas and automatically generating a natural language report based on the anomaly detection results. Background Technology

[0002] In industrial sites, monitoring is performed using various sensor data such as temperature, vibration, and pressure for predictive maintenance technology, and the condition of the equipment is determined through this monitoring.

[0003] However, due to structural constraints of the equipment, it is difficult to attach sensors to all parts, and in particular, installation is often impossible inside rotating shafts or in enclosed structures, making it difficult to monitor the entire equipment structure using sensors.

[0004] Furthermore, attaching sensors to all possible locations leads to excessively high installation and maintenance costs, and there was a problem in that it was difficult to automatically reflect the overall behavioral patterns or external changes of the equipment over time, or changes due to the external environment, making early detection of abnormalities or inference of their causes difficult.

[0005] Meanwhile, when an anomaly is detected in equipment, managers are required to report the status; however, since the detection results output by monitoring devices are often provided in a format difficult for non-experts to understand, managers must prepare a separate report that decision-makers can comprehend. This results in additional time required for report preparation, making immediate reporting of the anomaly difficult and consequently raising concerns about missing the golden time for intervention. Prior art literature

[0006] Republic of Korea Published Patent Application 10-2024-0065436 A (May 14, 2024) The problem to be solved

[0007] To solve the aforementioned conventional problems, the embodiments disclosed in this disclosure aim to provide an electronic device and method for detecting equipment anomalies based on vision images and generating natural language reports, which hierarchically detect abnormal areas of equipment based on vision images and automatically generate a natural language-based report summarized for the detected anomalies tailored to the recipient.

[0008] The problems that this disclosure aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0009] An electronic device for detecting equipment anomalies based on vision images and generating natural language reports according to the present disclosure, for achieving the technical objectives described above, comprises: a vision camera for capturing real-time vision images of industrial equipment; an anomaly detection model for receiving the real-time vision images and hierarchically detecting anomalies of the industrial equipment; a natural language report generation module for generating a report by processing the detected anomalies of the industrial equipment in natural language; a memory storing at least one process for performing operations of hierarchically detecting anomalies of the industrial equipment and generating a user-customized report for the detected anomalies; and at least one processor for performing the operations according to the process. The at least one processor may be configured to receive real-time vision image data of the industrial equipment, generate a spatial-temporal feature map of the vision image data through the anomaly detection model to first detect anomaly candidates, secondarily detect whether an anomaly exists in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area, and automatically generate a natural language-based report based on a preset template for the area determined to be an anomaly and the type of defect.

[0010] A method for detecting equipment anomalies based on vision images and generating natural language reports, performed by a processor of a device according to the present disclosure to achieve the technical problem described above, wherein the method comprises: receiving real-time vision image data for industrial equipment; generating a spatial-temporal feature map for the vision image data through an anomaly detection model; detecting anomaly candidates in a first step and detecting anomalies in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area in a second step; and automatically generating a natural language-based report based on a preset template for the area determined to be anomaly and the type of defect.

[0011] In addition to this, a computer program stored on a computer-readable recording medium for implementing the present disclosure may be further provided.

[0012] In addition to this, a computer-readable recording medium for recording a computer program for implementing the present disclosure may be further provided. Effects of the invention

[0013] According to the aforementioned means for solving the problem of the present disclosure, the abnormal condition of the equipment is detected in real time using only a vision camera without the installation of a separate sensor, and information regarding the type of abnormality and predicted lifespan is provided, while continuously improving model accuracy through a feedback learning function, thereby providing the effect of providing accurate safety information in real time.

[0014] Furthermore, according to the aforementioned means for solving the problem of the present disclosure, by summarizing reports differently according to the information required by field managers and decision-makers and automatically attaching visual evidence, it provides the effect of enabling even non-experts to immediately understand the status of the equipment and manage the equipment safely and efficiently.

[0015] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0016] FIG. 1 is a simplified block diagram illustrating the configuration of an electronic device for generating a vision image-based equipment anomaly detection and natural language report according to the present disclosure. FIG. 2 is a block diagram briefly illustrating the configuration of a system including an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. FIG. 3 is a block diagram briefly illustrating the configuration of an image processing module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. FIG. 4 is a block diagram briefly illustrating the functions of an anomaly detection model of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. FIG. 5 is a block diagram briefly illustrating the configuration of a natural language report generation module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. FIG. 6 is a process diagram illustrating the process of a natural language report generation module according to the present disclosure generating a natural language-based report. FIG. 7 is a configuration diagram illustrating a visualization module and an alarm module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. FIG. 8 is a process diagram illustrating the feedback learning process of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure. Specific details for implementing the invention

[0017] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and general content in the art to which this disclosure pertains or content that overlaps between embodiments is omitted. The terms 'part, module, component, block' as used in the specification may be implemented in software or hardware, and depending on the embodiments, a plurality of 'parts, modules, components, blocks' may be implemented as a single component, or a single 'part, module, component, block' may include a plurality of components.

[0018] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are directly connected but also cases where they are indirectly connected, and indirect connections include connections made via a wireless communication network.

[0019] Furthermore, when it is stated that a part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0020] Throughout the specification, when it is stated that a component is located "on" another component, this includes not only cases where a component is in contact with another component, but also cases where another component exists between the two components.

[0021] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.

[0022] Singular expressions include plural expressions unless there is an obvious exception in the context.

[0023] In each step, identification codes are used for convenience of explanation and do not describe the order of the steps; the steps may be performed differently from the specified order unless a specific order is clearly indicated in the context.

[0024] The operating principles and embodiments of the present disclosure will be described below with reference to the attached drawings.

[0025] In this specification, the term "device according to the present disclosure" includes all various devices capable of performing computational processing and providing results to a user. For example, the device according to the present disclosure may include all of a computer, a server device, and a portable terminal, or may be in the form of any one of these.

[0026] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.

[0027] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.

[0028] The above portable terminal may include, for example, all types of handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminals, smartphones, etc., as well as wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMDs).

[0029] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0030] The predefined operating rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined operating rules or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0031] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.

[0032] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that enables a machine to learn by mimicking human biological neurons. Methodologies of artificial intelligence can be classified according to the learning method into supervised learning, where input and output data are provided together as training data and the solution (output data) to the problem (input data) is predetermined; unsupervised learning, where only input data is provided without output data and the solution (output data) to the problem (input data) is not predetermined; and reinforcement learning, where a reward is given from an external environment whenever an action is taken from the current state, and learning proceeds in a direction that maximizes such reward. In addition, artificial intelligence methodologies can be classified according to the architecture, which is the structure of the learning model. The architectures of widely used deep learning technologies can be classified into Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, and Generative Adversarial Networks (GAN).

[0033] The device and system may include an artificial intelligence model. The artificial intelligence model may be a single model or may be implemented as multiple models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms in machine learning and cognitive science that mimic biological neurons. A neural network may refer to a model that possesses problem-solving capabilities by having artificial neurons (nodes) that form a network through synaptic connections and change the strength of synaptic connections through learning. The neurons of a neural network may include combinations of weights or biases. A neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a result (output) to be predicted from an arbitrary input by changing the weights of the neurons through learning.

[0034] The processor can create a neural network, train or learn a neural network, perform computations based on received input data, generate an information signal based on the results of the computation, or retrain the neural network. Neural network models may include, but are not limited to, various types of models such as Convolutional Neural Networks (CNN), Region with Convolutional Neural Networks (R-CNN), Region Proposal Networks (RPN), Recurrent Neural Networks (RNN), Stacking-based Deep Neural Networks (S-DNN), State-Space Dynamic Neural Networks (S-SDNN), Deconvolutional Networks, Deep Belief Networks (DBN), Restricted Boltzmann Machines (RBM), Fully Convolutional Networks, Long Short-Term Memory Networks (LSTM), and Classification Networks, such as GoogleNet, AlexNet, and VGG Network. The processor may include one or more processors to perform computations according to the neural network models. For example, a neural network is a deep neural It may include a network (Deep Neural Network).

[0035] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Function), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), GAN (Generative Adversarial Network), LSM (Liquid State Machine), ELM (Extreme Learning Machine), ESN (Echo A person skilled in the art will understand that any neural network may be included, but is not limited to, State Network, Deep Residual Network, Differential Neural Computer, Neural Turing Machine, Capsule Network, Kohonen Network, and Attention Network.

[0036] According to an exemplary embodiment of the present disclosure, the processor comprises a Convolutional Neural Network (CNN) such as GoogleNet, AlexNet, VGG Network, Region with Convolutional Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based Deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolutional Network, Deep Belief Network (DBN), Restricted Boltzmann Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA, Text Analysis, Dialog System, GPT-3, GPT-4 for Natural Language Processing, Visual Analytics, Visual Understanding, Video Synthesis for Vision Processing, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization for ResNet Data Intelligence, Various artificial intelligence structures and algorithms, such as recommendation and data creation, may be used, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0037] FIG. 1 is a block diagram briefly illustrating the configuration of an electronic device (100) that generates a vision image-based equipment anomaly detection and natural language report according to the present disclosure.

[0038] Referring to FIG. 1, the electronic device (100) according to the present disclosure may include an input / output module (110), a communication module (120), a memory (130), and a processor (140). Hereinafter, the electronic device (100) is an electronic device that performs hierarchical anomaly detection based on vision images and then automatically generates a natural language report on the anomaly detection results, and the method thereof is assumed to be implemented through the electronic device (100).

[0039] The input / output module (110) may be various interfaces or connection ports that receive user input or output information to the user. The input / output module (110) may be divided into an input module and an output module.

[0040] The input module receives user input from the user. The input module is for inputting video information (or signals), audio information (or signals), data, or information input by the user, and may include at least one of at least one camera, at least one microphone, and a user input unit. Voice data or image data collected by the input unit may be analyzed and processed into a user control command.

[0041] User input can take various forms, including key input, touch input, and voice input. Examples of input modules capable of receiving such user input include traditional keypads, keyboards, and mice; as well as touch sensors that detect user touch; microphones that receive voice signals; cameras that recognize gestures through image recognition; proximity sensors consisting of light or infrared sensors that detect user approach; motion sensors that recognize user movements using accelerometers or gyroscopes; and all other diverse forms of input means that detect or receive various types of user input. This is a comprehensive concept.

[0042] Here, the touch sensor can be implemented as a piezoelectric or capacitive touch sensor that detects touch through a touch panel or touch film attached to the display panel, or as an optical touch sensor that detects touch by an optical method. In addition, the input module may be implemented in the form of an input interface (USB port, PS / 2 port, etc.) that connects an external input device to receive user input, instead of a device that detects user input itself.

[0043] The output module can output various types of information and provide it to the user. The output module is a comprehensive concept that includes a display for outputting video, a speaker for outputting sound (and / or an amplifier connected thereto), a haptic device for generating vibration, and various other forms of output means. In addition, the output module may be implemented in the form of a port-type output interface that connects the individual output means described above.

[0044] For example, an output module in the form of a display can display text, still images, and videos. The term "display" refers to a broad concept of an image display device that includes all types of devices capable of performing image output functions, such as Liquid Crystal Displays (LCDs), Light Emitting Diode (LED) displays, Organic Light Emitting Diode (OLED) displays, Flat Panel Displays (FPDs), transparent displays, Curved Displays, flexible displays, 1D displays, holographic displays, projectors, and others. Such a display may also take the form of a touch display integrated with the touch sensor of an input module.

[0045] In other words, the input / output module (110) can receive user input or provide output to the user based on a user interface.

[0046] The communication module (120) can communicate with an external device. Accordingly, the device can transmit and receive information with an external device through the communication module. For example, the device can communicate with an external device using the communication module so that information stored and generated within the electric vehicle charging management system is shared. The communication module (120) may include, for example, at least one of a wired communication module, a wireless communication module, a short-range communication module, and a location information module.

[0047] Here, communication, that is, the transmission and reception of data, can be performed via wired or wireless means. To this end, the communication module may be composed of a wired communication module that connects to the Internet, etc., via a Local Area Network (LAN); a mobile communication module that connects to a mobile communication network via a mobile communication base station to transmit and receive data; a short-range communication module that uses a Wireless Local Area Network (WLAN) family communication method such as Wi-Fi or a Wireless Personal Area Network (WPAN) family communication method such as Bluetooth or Zigbee; a satellite communication module that uses a Global Navigation Satellite System (GNSS) such as GPS; or a combination thereof. The wireless communication technology used for communication may include Narrowband Internet of Things (NB-IoT) for low-power communication. In this case, for example, NB-IoT technology may be an example of LPWAN (Low Power Wide Area Network) technology and may be implemented according to standards such as LTE Cat (category) NB1 and / or LTE Cat NB2, but is not limited to the names mentioned above. Additionally, or generally, wireless communication technology implemented in wireless devices according to various embodiments may perform communication based on LTE-M technology. In this case, for example, LTE-M technology may be an example of LPWAN technology and may be referred to by various names such as eMTC (enhanced Machine Type Communication).For example, LTE-M technology may be implemented in at least one of various standards such as 1) LTE CAT 0, 2) LTE Cat M1, 3) LTE Cat M2, 4) LTE non-BL (non-Bandwidth Limited), 5) LTE-MTC, 6) LTE Machine Type Communication, and / or 7) LTE M, and is not limited to the names mentioned above. Additionally or generally, wireless communication technology implemented in wireless devices according to various embodiments may include at least one of ZigBee, Bluetooth, and Low Power Wide Area Network (LPWAN) for low-power communication, and is not limited to the names mentioned above. As an example, ZigBee technology can create personal area networks (PANs) related to small / low-power digital communication based on various standards such as IEEE 802.15.4 and may be referred to by various names.

[0048] The memory (130) can store various types of information. The memory can store data temporarily or semi-permanently. For example, the memory may store an operating system (OS) for operating the first device and / or the second device, data for hosting a website, or data regarding a program or application (e.g., a web application) for generating Braille. In addition, the memory may store modules in the form of computer code as described above.

[0049] Examples of memory (130) may include a hard disk drive (HDD), a solid state drive (SSD), flash memory, ROM (Read-Only Memory), and RAM (Random Access Memory). These memories may be provided as built-in or removable types.

[0050] The processor (140) controls the overall operation of the electronic device (100). To this end, the processor (140) performs computation and processing of various information and can control the operation of the components of the first device and / or the second device.

[0051] The processor (140) may be implemented as a computer or a similar device according to hardware, software, or a combination thereof. Hardware-wise, the processor (140) may be provided in the form of an electronic circuit that processes electrical signals to perform control functions, and software-wise, it may be provided in the form of a program that drives the hardware processor. Meanwhile, unless otherwise specifically mentioned in the following description, the operation of the first device and / or the second device may be interpreted as being performed by the control of the processor (140). That is, the modules may be interpreted as the processor (140) controlling the first device and / or the second device to perform the following operations.

[0052] The processor (140) may be implemented with a memory that stores data for an algorithm or a program that reproduces the algorithm for controlling the operation of components within the device, and at least one sub-processor (not shown) that performs the aforementioned operation using the data stored in the memory. In this case, the memory and the processor may each be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.

[0053] Additionally, the processor (140) may control one or a combination of the components described above in order to implement various embodiments according to the present disclosure, which will be described in the drawings below, on the device.

[0054] FIG. 2 is a block diagram briefly illustrating the configuration of a system including an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0055] An electronic device (100) for detecting equipment anomalies based on vision images and generating natural language reports according to the present disclosure may include: a sensor module (200) including a vision camera that captures real-time vision images of industrial equipment (10); an anomaly detection model (400) that receives the real-time vision images and hierarchically detects anomalies of the industrial equipment (10); a natural language report generation module (500) that generates a report by processing the detected anomalies of the industrial equipment (10) in natural language; a memory (130) that stores at least one process for performing an operation of hierarchically detecting anomalies of the industrial equipment (10) and generating a user-customized report for the detected anomalies; and at least one processor (140) that performs the operation according to the process.

[0056] The above at least one processor (140) may be configured to receive real-time vision image data for the industrial equipment (10), generate a spatial-temporal feature map for the vision image data through the anomaly detection model (400) to detect anomaly candidates first, detect anomalies secondarily in a secondary extended area including the area of ​​the first detected anomaly candidate and the surrounding area of ​​the anomaly candidate, and automatically generate a natural language-based report based on a preset template for the area determined to be anomaly and the type of defect.

[0057] Specifically, as illustrated in FIG. 2, a system including an electronic device (100) according to one embodiment of the present disclosure may include a sensor module (200) including a vision camera for photographing industrial equipment (10), an image processing module (300) for preprocessing a vision image, an anomaly detection model (400) for detecting anomalies in a preprocessed vision image, a natural language report generation module (500) for generating a natural language-based report on the detected anomalies, and a feedback learning unit (600) for relearning an anomaly detection method and feedback generation rules based on feedback on the anomaly detection and report.

[0058] The industrial equipment (10) is equipment that operates at an industrial site, such as a motor, pump, press, conveyor, etc., and can be shown in a normal operating state or an abnormal operating state in an image.

[0059] The sensor module (200) may include a vision camera for capturing the exterior of the industrial equipment (10). The vision camera may capture the surface condition of the equipment in real time and provide image data to the image processing module (300).

[0060] For example, video streams can be received in real time from multiple vision cameras installed in key parts of the equipment, such as the drive unit, heating unit, and vibration unit. The received video is temporarily stored in a buffer, and metadata including the equipment ID, camera ID, and shooting time can be transmitted along with video segments at regular time intervals for time-series analysis.

[0061] The image processing module (300) can preprocess image data received from the sensor module (200) and perform resolution normalization, noise removal, contrast correction, etc. of the image. In addition, the image processing module (300) can generate a spatio-temporal feature map that combines spatial and temporal changes of the equipment from consecutive frames in a time series.

[0062] The anomaly detection model (400) can analyze the spatial-temporal feature map input from the image processing module (300) and detect an area where the deviation exceeds a threshold value based on the statistical distance or similarity difference from a reference normal state as a first-order anomaly candidate area. The anomaly detection model (400) can set a second-order extended area including a surrounding area centered on the first-order anomaly candidate area and re-detect whether an anomaly exists by considering the surrounding extended area together with spatial continuity and temporal change patterns.

[0063] At this time, the first detection step is performed as a first screening step to quickly extract areas where anomalies are likely to occur, and the second extended detection step can be performed as a step to finely determine whether to confirm an actual anomaly by considering the condition of the surrounding area together. Through this, false positives that may occur due to environmental factors such as reflected light, changes in lighting, and temporary vibrations can be reduced.

[0064] The natural language report generation module (500) receives an anomaly determination result output from the anomaly detection model (400), structures the location of the anomaly area, the type of anomaly, the severity, and the prediction information, and maps it to a preset sentence template to automatically generate a natural language-based report.

[0065] The report may include video frames showing the abnormal areas, and the cause of the abnormality and recommended corrective actions may be included in sentence form.

[0066] The feedback learning unit (600) can collect report feedback information provided by the manager. The feedback information may include anomaly detection feedback and report writing feedback, such as the accuracy of the report, the suitability of anomaly judgment, and whether there is a false positive. If the number of accumulated feedbacks exceeds a certain standard, an automatic retraining procedure for the anomaly detection model (400) and the natural language report generation module (500) can be performed.

[0067] This allows for the continuous improvement of the detection model's judgment accuracy and the quality of the report.

[0068] FIG. 3 is a block diagram briefly illustrating the configuration of an image processing module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0069] As illustrated in FIG. 3, the image processing module (300) of the present disclosure may include a preprocessing unit (310), a feature extraction unit (320), and a feature map generation unit (330).

[0070] The preprocessing unit (310) can perform a preprocessing process to convert a raw image received from a vision camera into a form suitable for anomaly detection.

[0071] For example, the preprocessing unit (310) can clearly express changes in the appearance of the equipment by performing preprocessing such as resolution normalization, contrast adjustment, and noise removal on a frame-by-frame basis on the received raw image segment. Alternatively, to reduce the amount of computation, the area of ​​interest of the equipment can be set automatically or by user.

[0072] Additionally, the preprocessing unit (310) can perform frame alignment to correct distortion caused by changes in camera position or lighting conditions in the time-series image.

[0073] The feature extraction unit (320) receives a preprocessed image and can extract spatial features and temporal features together.

[0074] Spatial features may include spatial information such as color distribution, texture, and shape of specific areas within the equipment in the image, and feature vectors can be extracted at the equipment part level by subdividing the area by equipment part.

[0075] Temporal features may include the amount of change between consecutive frames, movement vectors containing dynamic information such as vibration or rotation of equipment parts, or repeating patterns reflecting the trend of change of feature vectors along the time axis.

[0076] The feature extraction unit (320) can generate a multilayer feature vector that simultaneously reflects information in the time axis and the spatial axis using a convolutional neural network (CNN) based filter, a 3D convolution (3D-CNN), or a recurrent neural network (RNN) structure.

[0077] The feature map generation unit (330) can generate a spatial-temporal feature map by integrating the multidimensional feature vector obtained from the feature extraction unit (320).

[0078] The feature map generation unit (330) can comprehensively represent where, when, and what changes occurred by combining extracted spatial features and temporal features to form a spatial-temporal feature map in the form of a three-dimensional tensor. The deviation can be quantified by comparing the real-time spatial-temporal feature map with the normal operation spatial-temporal feature map.

[0079] Additionally, the feature map generation unit (330) can calculate the correlation between multiple frames of the same time period and fuse macroscopic and microscopic features of the equipment based on feature intensity to generate a comprehensive spatial-temporal feature map that fuses the two scales.

[0080] For example, macroscopic features including overall positional movement, tilting, and shape changes due to large-scale vibration of the equipment, and microscopic features including local heat distribution on the surface of the equipment, fine deformation, wear marks, loosened bolts, and leakage marks can be generated into a spatial-temporal feature map in the form of a 3D tensor and transmitted to an anomaly detection model (400).

[0081] The anomaly detection model (400) can detect an anomaly region by receiving the spatial-temporal feature map as input and comparing it with the spatial-temporal feature map of a normal state.

[0082] FIG. 4 is a block diagram briefly illustrating the functions of an anomaly detection model of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0083] An anomaly detection model (400) according to one embodiment of the present disclosure can detect a spatial-temporal feature map that exceeds a normal spatial-temporal feature map and a preset statistical distance as an anomaly region, and after the first anomaly candidate detection (411), can perform a second expanded region analysis (412) on an expanded region including the detected first anomaly candidate region and the surrounding region.

[0084] Specifically, the anomaly detection model (400) can detect primary anomaly candidates (411) that are anomaly regions exceeding a normal spatial-temporal feature map and a preset statistical distance.

[0085] Then, a second extended region analysis (412) including the detected first abnormal candidate region and the surrounding region of the first abnormal candidate region can be performed. The continuity can be examined to see if the first abnormal candidate region shows an abnormal pattern consecutively in the first detection and the second detection, and the extended region can be examined to see if an abnormal pattern appears in the surrounding region as well.

[0086] In one embodiment, the anomaly detection model (400) may be configured to eliminate false detections by comparing the detected anomaly area with the surrounding area of ​​the anomaly area, and if the first anomaly candidate area detected as an anomaly area in the first detection is continuously determined to be an anomaly in the second detection or the anomaly is extended to the surrounding area in the second detection, the first anomaly candidate area may be determined to be the final anomaly area.

[0087] For example, if a vision camera is installed on the upper part of a conveyor belt in a food factory, and an anomaly with an increasing positional deviation is detected in the left-side area of ​​the belt while no anomaly is detected in the surrounding area of ​​the left-side area, it can be determined that the anomaly is isolated only in the left-side area of ​​the belt and is not caused by temporary lighting changes or interference from external objects, and thus it can be judged as a normal detection.

[0088] Subsequently, when performing a second detection, if an anomaly occurs in which the position deviation increases in the left area of ​​the belt (continuity) and the position deviation increases in the surrounding area (scalability), the anomaly can be determined as the final anomaly area.

[0089] Alternatively, if an area where brightness changes temporarily due to light reflected on the metal surface of industrial equipment (10) is classified as an anomaly in the first detection, the anomaly detection model (400) may re-determine the area as normal if the anomaly is not observed within the second extended area (discontinuity) or if it shows the same trend as the surrounding area. Conversely, if a similar anomaly pattern persists in both the first detection area and the surrounding extended area, the area may be confirmed as the final anomaly area.

[0090] The anomaly detection model (400) can perform anomaly cause classification (413) for anomalies determined as final anomaly regions. Based on the morphological characteristics of the anomaly pattern in the final anomaly region, it can analyze the type of anomaly and output the cause of the anomaly. The results of the anomaly pattern analysis can be classified into predefined cause categories.

[0091] For example, cause categories can be classified into broad categories such as sudden change, gradual change, periodic fluctuation, localized heat generation, and foreign matter ingress, and further classified into detailed categories such as impact, damage, and sudden stop within sudden change, and wear, corrosion, and leakage within gradual change.

[0092] The accuracy of cause estimation can be improved by matching each cause category with a database of past failure cases.

[0093] Additionally, the anomaly detection model (400) can output a failure probability or a predicted remaining lifespan based on the amount of change over time of the anomaly pattern and the history of past anomaly occurrences.

[0094] It is possible to predict how much time remains until failure occurs by considering the amount of change over time of abnormal patterns, and it is also possible to predict how much time remains until failure occurs based on the deterioration rate over time by referring to the progression history of similar past abnormal patterns.

[0095] The anomaly detection model (400) can output a risk level according to the cause of the anomaly, determine and output a severity score according to the probability of failure occurrence, and an urgency grade according to the predicted remaining lifespan.

[0096] FIG. 5 is a block diagram briefly illustrating the configuration of a natural language report generation module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0097] As illustrated in FIG. 5, the natural language report generation module (500) may include a vision language model (511), a context analysis model (512), a thought chain reasoning engine (513), an insight generation unit (514), and an expert knowledge DB (515).

[0098] The vision language model (511) receives data of the vision image and the anomaly detection result received from the anomaly detection model (400) and can convert the main features of the image and the location of the anomaly into text information.

[0099] For example, by analyzing vision images, descriptive sentences such as “traces of leakage around the valve” or “discoloration of the pipe surface” can be automatically generated.

[0100] The context analysis model (512) can combine metadata containing a history of past anomalies with the descriptive sentences generated by the vision language model (511) to contextually analyze the current status of the equipment and add it to the report.

[0101] For example, by considering the past anomaly history of the equipment where the anomaly occurred, contextual information including the number of anomalies within a specific recent period, the interval between occurrences, changes in risk levels compared to the past history, and the history of past actions can be reflected.

[0102] For example, when considering the predicted causes of the anomaly and the history of past anomalies, it is possible to distinguish whether the current anomaly is a normal deviation associated with a rise in ambient temperature or a sign of a structural defect such as a leak.

[0103] The thought chain reasoning engine (513) can analyze the judgment basis of the anomaly detection model (400) in the form of multi-stage reasoning and add it to the report.

[0104] For example, by analyzing the logical flow of “increase in traces of pipe leakage within the vision image → same location as past anomalies → maintenance required due to repetitive anomalies,” one can derive the conclusion that “maintenance is required due to repeated leakage because the increase in traces of pipe leakage is similar to past leakage patterns.”

[0105] The insight generation unit (514) can generate a report in natural language form that is easy for a manager or maintenance person to understand based on the inferred results.

[0106] For example, an automatically summarized report can be generated in the form of, “Anomaly in heat distribution has been detected in the valve area, and the rate of temperature rise is about 2.5 times faster than past similar patterns. Inspection within 48 hours is recommended.”

[0107] The expert knowledge database (515) includes structural information of industrial facilities, maintenance guidelines, history of past abnormal response, and safety regulations, and can be used to accurately describe technical terms, defect classification systems, and action recommendations during the report generation process.

[0108] For example, regarding a detection result of “cooling system valve leak,” it can automatically supplement with a professional explanation such as “possibility of O-ring damage in the cooling line valve.”

[0109] Accordingly, the natural language report generation module (500) can automatically organize and provide anomaly detection results in a form that is easy for humans to understand, thereby reducing the workload of an administrator who previously had to manually write anomaly reports and improving the consistency and accuracy of the report quality.

[0110] FIG. 6 is a process diagram illustrating the process of a natural language report generation module according to the present disclosure generating a natural language-based report.

[0111] As illustrated in FIG. 6, the natural language report generation module (500) can perform a data structuring step (S510), template-based basic sentence generation (S520), priority-based sentence alignment (S530), context information and corresponding image addition (S540), and receiver-level summary information generation (S550).

[0112] In the data structuring step (S510), the natural language report generation module (500) can organize the result data output from the anomaly detection model (400) into a structured data format.

[0113] Specifically, the equipment name, equipment ID, the vision camera ID that captured the equipment, the anomaly type, cause category, risk level, time of occurrence, location of occurrence, predicted remaining lifespan, urgency grade, and the number of past anomaly occurrences of the same equipment and the interval between recent occurrences can be organized by dividing them into intervals.

[0114] In the template-based basic sentence generation (S520), the natural language report generation module (500) can generate a natural language sentence-based report by mapping the structured data to a preset sentence template. The templates are stored in various ways according to anomaly types and can be automatically selected based on the anomaly type and the cause of the anomaly.

[0115] For example, data organized according to a structured data format can be mapped to the sentence, “[Anomaly Type] was detected at [Time of Occurrence] in [Part Name] of [Equipment Name].”

[0116] In the priority-based sentence alignment (S530), when multiple anomalies occur, the natural language report generation module (500) determines the priority based on the severity according to the risk level of the anomaly and the urgency according to the predicted remaining lifespan, and can adjust the order so that the natural language sentence corresponding to the anomaly with the highest priority among the multiple anomalies is described first.

[0117] For example, if an anomaly occurs simultaneously in multiple facilities, the anomalies can be classified according to severity and urgency, and priority scores can be calculated so that items requiring faster verification by the manager are listed first.

[0118] In addition, if multiple abnormalities within the same facility are related, natural language sentences can be generated by inferring causal relationships, adjusting the narrative order, and integrating them into a single narrative paragraph.

[0119] In the addition of context information and corresponding images (S540), the natural language report generation module (500) generates a natural language sentence that reflects context information, including the number of anomalies within a specific period, the interval between occurrences, the change in risk level compared to the past occurrence history, and the past action history, by considering the past anomaly history of the equipment where the anomaly occurred, and can add the natural language sentence reflecting the context information to the report.

[0120] Additionally, the natural language report generation module (500) can attach a video frame containing the above-mentioned final abnormal region to the generated report.

[0121] For example, contextual information such as “anomalies occurred 3 times in the same part over the past 7 days” or “severity increased by 20% compared to the previous” can be automatically inserted into sentences by referring to the history of past abnormalities of the same equipment stored in the database, or sentences such as “recurred after recent maintenance and parts replacement” can be added by referring to the history of past actions.

[0122] After natural language sentences are inserted, vision image frames related to the aforementioned anomalies can be attached as evidence for each anomaly item within the report. Visual highlighting can be performed by marking the anomaly occurrence area on the vision image frames with bounding boxes or heatmaps, or the process of change in the anomaly state can be reflected by rapidly converting time-series images.

[0123] In the recipient-level based summary information generation (S550), the natural language report generation module (5000) can automatically adjust the summary level of the report according to the recipient's role and skill level.

[0124] Specifically, if the recipient is an engineer, the report can be generated based on a detailed technical report that includes all numerical values, coordinates, and analysis processes related to the anomaly pattern, using technical terminology as is.

[0125] Alternatively, if the recipient is a field manager, the report can be generated by basing it on an interim summary report that includes key details, causes, actions, and recommendations related to the abnormal pattern, and by appropriately replacing technical terms.

[0126] Alternatively, if the recipient is management, the report can be generated based on a key summary report that includes the overall facility status, major anomalies where risk levels exceed a threshold, and expected impacts, by replacing technical terms with basic terminology. In particular, if facility replacement is required, the report can be generated to include the costs incurred due to the replacement and the potential cost savings when considering future benefits.

[0127] In other words, the natural language report generation module (500) can generate a report by automatically performing an appropriate level of summary according to the recipient settings. The generated report can be provided in a fixed format, such as email, messenger, or a dashboard within the system, and an alarm can be sent immediately if the risk level is very high.

[0128] Accordingly, the natural language report generation module (500) can generate different reports by adjusting the summary level of the natural language sentences generated according to the recipient information.

[0129] FIG. 7 is a configuration diagram illustrating a visualization module and an alarm module of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0130] As illustrated in FIG. 7, the visualization module (700) can be additionally connected to the facility drawing DB (910) and the past anomaly history DB (920).

[0131] The equipment drawing DB (910) contains information such as the structure, piping, and sensor locations of actual industrial equipment, and the visualization module (700) can map the area where an abnormality occurred and the abnormality information to the equipment drawing and output it to the display screen.

[0132] Additionally, the visualization module (700) can mark defects in areas where abnormalities are currently detected on a 2D or 3D drawing of the facility with preset colors according to risk level.

[0133] For example, if an abnormality caused by repeated leakage occurred at the same valve point in the past, the visualization module (700) can indicate the leakage progression status by gradually expressing the area in a darker color.

[0134] The past anomaly occurrence history DB (920) contains data on the area where the anomaly occurred, and the visualization module (700) can visualize the change process of the anomaly over time by mapping the time of occurrence of the anomaly, the classified type of the anomaly, and the result of the action to the past anomaly detection history of the area where the anomaly was detected.

[0135] The visualization module (700) can display a suggested action or warning message for the currently detected anomaly type by referring to the record of response measures included in the history of past anomalies.

[0136] As illustrated in FIG. 7, the visualization module (700) may be linked with the alarm module (800) to output an anomaly occurrence and a risk level, and at the same time, the alarm module (800) may provide a multi-layer alarm corresponding to the output risk level.

[0137] Reports regarding abnormal situations are provided through a designated channel depending on the recipient, and evaluation feedback on the abnormal detection and report can be received through the provided channel. The feedback data can be stored together with the abnormally detected vision image and the abnormal detection analysis results to be used for retraining the abnormal detection model (400).

[0138] FIG. 8 is a process diagram illustrating the feedback learning process of an electronic device that generates vision image-based equipment anomaly detection and natural language reports according to the present disclosure.

[0139] When an anomaly occurs and measures regarding it are completed, retraining of the anomaly detection model (400) or natural language report generation module (500) can be performed periodically to further learn about normal data and anomaly data or to further learn rules for report generation.

[0140] As illustrated in FIG. 8, feedback on the anomaly detection and report is received through the feedback interface (111), and when the received feedback exceeds a preset number, retraining of the anomaly detection model and the natural language report generation module can be automatically performed.

[0141] The feedback interface (111) is a user interface for evaluating the quality of anomaly detection results and natural language reports, and can be provided in the form of a GUI in which an administrator or inspector can directly input feedback.

[0142] In one embodiment, the feedback interface (111) may be displayed on the report screen in the form of a checkbox or selection item and may provide a plurality of items for false positives of the anomaly detection model (400), such as accuracy, false positive, overestimation of severity, underestimation of severity, and cause error.

[0143] In addition, detailed errors or improvements to the report can be freely entered through the comment input field of the feedback interface (111).

[0144] The database (900) can receive and store evaluation information input from the feedback interface (111). The feedback data may store results output by an anomaly detection model (400) or a natural language report generation module (500) mapped to the user's evaluation results. For example, it may be stored in a form such as “Anomaly area detection accuracy: 95%”, “Severity overestimation: present”, “Comment: Actual defect is minor”.

[0145] When the number of feedback data in the database (900) exceeds a preset number, a retraining trigger signal is automatically transmitted to the feedback learning unit (600) to perform retraining on the anomaly detection model (400) or the natural language report generation module (500).

[0146] When the feedback learning unit (600) receives a trigger signal, it can perform automatic relearning to update the parameters of the anomaly detection model (400) and the natural language report generation module (500) based on accumulated feedback data.

[0147] For example, if the anomaly detection model (400) repeatedly outputs an overestimated severity in a specific equipment part, the feedback learning unit (600) can reduce the frequency of overestimation by re-evaluating the learning data portion for the same part and adjusting the severity prediction weights within the loss function.

[0148] In addition, if the natural language report generation module (500) uses a sentence template that is different from the actual result, or if there is a discrepancy between the result and the sentence template, the sentence generation can be corrected based on the expression included in the feedback comment.

[0149] Accordingly, an electronic device (100) according to one embodiment of the present disclosure can detect abnormal conditions of equipment in real time using only a vision camera without installing a separate sensor, provide information on the type of abnormality and predicted lifespan, and continuously improve model accuracy through a feedback learning function.

[0150] In addition, by summarizing reports differently based on the information required by on-site managers and decision-makers and automatically attaching visual evidence, even non-experts can immediately understand the status of the equipment and manage it safely and efficiently.

[0151] Meanwhile, a method for detecting equipment anomalies based on vision images and generating natural language reports, performed by a processor of a device according to the present disclosure, the method may include: receiving real-time vision image data for industrial equipment; generating a spatial-temporal feature map for the vision image data through an anomaly detection model; detecting anomaly candidates in a first step, and detecting anomalies in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area in a second step; and automatically generating a natural language-based report based on a preset template for the area determined to be anomaly and the type of defect.

[0152] Content that overlaps with the above is omitted for the sake of brevity in the specification.

[0153] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium that stores instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operation of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0154] Computer-readable recording media include all types of recording media that store instructions that can be decoded by a computer. Examples include ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, optical data storage devices, etc.

[0155] As described above, the disclosed embodiments have been explained with reference to the attached drawings. Those skilled in the art will understand that the present disclosure may be practiced in forms different from the disclosed embodiments without changing the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be interpreted restrictively. Explanation of the symbols

[0156] 100: Electronic device 110: I / O module 111: Feedback Interface 120: Communication module 130: Memory 140: Processor 200: Sensor module 300: Image processing module 310: Preprocessing Unit 320: Feature Extraction Unit 330: Feature Map Generation Unit 400: Anomaly detection model 500: Natural Language Report Generation Module 511: Vision Language Model 512: Context Analysis Model 513: Thought Chain Reasoning Engine 514: Insight Generation Unit 515: Expertise DB 600: Feedback Learning Department 700: Visualization Module 800: Alarm Module 900: DB 910: Facility Drawing DB 920: Past Anomaly History DB

Claims

Claim 1 A vision camera for capturing real-time vision images of industrial equipment; an anomaly detection model that receives the real-time vision images and hierarchically detects anomalies in the industrial equipment; a natural language report generation module that generates a report by processing the detected anomalies in the industrial equipment in natural language; and a memory storing at least one process for performing the operation of hierarchically detecting anomalies in the industrial equipment and generating a user-customized report for the detected anomalies. and at least one processor that performs the operation according to the above process; wherein the at least one processor receives real-time vision image data for the industrial equipment, generates a spatial-temporal feature map for the vision image data through the anomaly detection model to first detect anomaly candidates, secondarily detects whether an anomaly exists in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area, and automatically generates a natural language-based report based on a preset template regarding the area determined to be anomaly and the type of defect, wherein the anomaly detection model detects a spatial-temporal feature map that exceeds a preset statistical distance from a normal spatial-temporal feature map as an anomaly area, eliminates false detections by comparing the detected anomaly area with the surrounding area of ​​the anomaly area, and if the first anomaly candidate area detected as an anomaly area in the first detection is continuously determined to be an anomaly in the second detection or if the anomaly is extended to the surrounding area in the second detection, the first anomaly candidate area is determined to be a final anomaly area, and the anomaly detection model outputs the type of anomaly and the cause of the anomaly based on the morphological characteristics of the anomaly pattern in the final anomaly area, and An electronic device for generating vision image-based equipment anomaly detection and natural language reports configured to output failure probability or predicted remaining life based on the amount of change over time of anomaly patterns and past anomaly history. Claim 2 delete Claim 3 delete Claim 4 An electronic device for generating vision image-based equipment anomaly detection and natural language reports, wherein, in claim 1, at least one processor structures data output from the anomaly detection model, the natural language report generation module maps the structured data to a preset sentence template to generate a natural language sentence-based report, and attaches an image frame containing the final anomaly region to the generated report. Claim 5 An electronic device for generating vision image-based equipment anomaly detection and natural language reports, wherein, in the event of a plurality of anomalies, the at least one processor is configured such that, the natural language report generation module determines a priority based on the severity according to the risk level of the anomaly and the urgency according to the predicted remaining lifespan, and adjusts the order so that the natural language sentence corresponding to the anomaly with the highest priority among the plurality of anomalies is described first. Claim 6 In claim 4, the electronic device for generating vision image-based equipment anomaly detection and natural language reports is configured such that at least one processor generates a natural language sentence reflecting contextual information, including the number of anomalies within a recent specific period, the interval between occurrences, the change in risk level compared to the past occurrence history, and the past action history, by considering the past anomaly history of the equipment where the anomaly occurred, and adds the natural language sentence reflecting the contextual information to the report. Claim 7 In claim 1, the at least one processor is an electronic device for generating vision image-based equipment anomaly detection and natural language reports, configured such that the natural language report generation module generates different reports by differently adjusting the summary level of natural language sentences generated according to recipient information. Claim 8 An electronic device for generating vision image-based equipment anomaly detection and natural language reports, wherein the at least one processor receives feedback on the anomaly detection and report through a feedback interface, and is configured to automatically perform retraining on the anomaly detection model and the natural language report generation module when the received feedback exceeds a preset number. Claim 9 An electronic device according to claim 1, wherein the at least one processor is configured to generate a vision image-based facility anomaly detection and natural language report, wherein the anomaly detection model extracts macroscopic characteristics including changes in the overall shape of the facility and microscopic features including changes in the shape of specific parts of the facility, and fuses the macroscopic characteristics and the microscopic features to determine a final anomaly. Claim 10 A method for generating a natural language report based on the result of equipment anomaly detection based on vision images performed by a processor of a device, the method comprising: receiving real-time vision image data for industrial equipment; generating a spatial-temporal feature map for the vision image data through an anomaly detection model; detecting anomaly candidates in a first step and detecting anomalies in a second extended area including the first detected anomaly candidate area and the surrounding area of ​​the anomaly candidate area in a second step; and automatically generating a natural language-based report based on a preset template for the area determined to be anomaly and the type of defect. A method for generating a natural language report based on a vision image-based equipment anomaly detection result, comprising: a second detection step in which the anomaly detection model detects a spatial-temporal feature map exceeding a normal spatial-temporal feature map and a preset statistical distance as an anomaly region, eliminates false detections by comparing the detected anomaly region with the surrounding region of the anomaly region, and if the first anomaly candidate region detected as an anomaly region in the first detection is continuously determined as an anomaly in the second detection or if the anomaly is extended to the surrounding region in the second detection, the first anomaly candidate region is determined as the final anomaly region; and an automatic generation step in which the anomaly detection model outputs the type of anomaly and the cause of the anomaly based on the morphological characteristics of the anomaly pattern in the final anomaly region, and outputs the probability of failure or the predicted remaining life based on the amount of change over time of the anomaly pattern and the history of past anomaly occurrences.

Citation Information

Patent Citations

  • For users of data processing systems to loosely group the sliders on the interface.

    KR1019950001504A

  • Sensor

    KR1020100129989A

  • Apparatus and method for safety condition prediction of industrial site based on image

    KR1020210050711A

  • Method, apparatus and system for detecting manufactured articles defects based on deep learning feature extraction

    KR1020230156512A

  • System and method for anomaly detection using images

    US20220207691A1