Method for predicting dangerous situation of object by analyzing plurality of image frames and utilizing physics-based model
By analyzing multiple image frames and incorporating static and dynamic physical information through physics-based models into a time series prediction model, the method effectively predicts dangerous situations of objects with improved accuracy and robustness.
Patent Information
- Application Number
- PCT/KR2024/017604
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-12
AI Technical Summary
Existing methods for predicting dangerous situations of objects rely solely on image frames without considering physical information, leading to inaccurate predictions and lack of robustness to environmental changes.
A method that analyzes multiple image frames and extracts both static and dynamic physical information of objects using physics-based models, which are then input into a time series prediction model to generate encoded features for predicting risk situations.
This approach enables accurate prediction of dangerous situations by integrating physical information into the prediction model, enhancing robustness to changes in data distribution due to environmental changes.
Smart Images

Figure KR2024017604_12062025_PF_FP_ABST
Abstract
Description
A method for predicting the risk of an object by analyzing multiple image frames and utilizing physics-based models.
[0001] The present disclosure relates to a method for predicting a dangerous situation of an object, and more particularly, to a method for predicting a dangerous situation of an object by analyzing a plurality of image frames.
[0002] Data-driven machine learning refers to systems that automatically build models and improve their performance using empirical data. Advances in data-driven machine learning have revolutionized computer vision, reinforcement learning, and various other scientific and engineering fields.
[0003] With recent advances in machine learning, deep neural networks possess powerful abstraction capabilities, and research is actively underway to apply deep neural networks to scientific problems that model physical systems.
[0004] Physics-informed machine learning (PIML) is a paradigm that builds models by leveraging prior knowledge of physics and empirical data representing high-level abstractions of natural phenomena or human behaviors to improve the performance of tasks involving physical mechanisms.
[0005] Unlike typical neural networks that rely entirely on data without considering physical laws, a physics-informed neural network (PINN) uses the physical equations themselves in its loss function to learn how to ensure that the sampled data satisfies the physical laws.
[0006] Korean Patent No. 10-2437392 (registration date: August 24, 2022) discloses a method for predicting molecular characteristics using a physics-based deep neural network model.
[0007] The present disclosure relates to a method for predicting a dangerous situation of an object by utilizing a plurality of image frames and physical information of the object included in the plurality of image frames.
[0008] Meanwhile, the technical task to be achieved by the present disclosure is not limited to the technical task mentioned above, and may include various technical tasks within a scope obvious to a person skilled in the art from the contents described below.
[0009] To solve the above-described problem, a method for predicting a dangerous situation of an object by analyzing a plurality of image frames, performed by a computing device, is disclosed. The method may include the steps of: acquiring a plurality of image frames; acquiring physical information of an object included in the plurality of image frames; inputting the plurality of image frames and the physical information of the object into a time-series prediction model; and utilizing the time-series prediction model, acquiring encoded time-series features based on the plurality of image frames and the physical information of the object.
[0010] In one embodiment, the step of obtaining physical information of an object included in the plurality of image frames may include a step of obtaining static information of an object included in each of the plurality of image frames by utilizing a first physics-based model.
[0011] In one embodiment, the step of inputting the plurality of image frames and the physical information of the object into a time series prediction model may include the step of inputting the image frame of the corresponding time point and the static information of the object included in the image frame of the previous time point into each of the plurality of encoder cells included in the encoder of the time series prediction model.
[0012] In one embodiment, the step of obtaining physical information of an object included in the plurality of image frames may further include a step of obtaining dynamic information of an object included in each of the plurality of image frames by utilizing a second physics-based model.
[0013] In one embodiment, the step of inputting the plurality of image frames and the physical information of the object into a time series prediction model may include a step of integrating the dynamic information of the object included in each of the image frames up to a previous point in time to obtain the dynamic information of the integrated object, and a step of inputting the image frame of the corresponding point in time, the static information of the object included in the image frame of the previous point in time, and the dynamic information of the integrated object into each of the plurality of encoder cells included in the encoder of the time series prediction model.
[0014] In one embodiment, the method may further include a step of predicting a risk situation of the object based on the encoded time series features by utilizing the time series prediction model.
[0015] In one embodiment, the step of inputting the plurality of image frames and the physical information of the object into a time series prediction model may include the step of inputting the plurality of image frames and the physical information of the object into an encoder of the time series prediction model.
[0016] In one embodiment, the step of predicting a dangerous situation of the object based on the encoded time series feature by utilizing the time series prediction model may include the steps of inputting the encoded time series feature into a decoder of the time series prediction model, the step of predicting an image frame at a predicted time based on the encoded time series feature by utilizing the decoder of the time series prediction model, and the step of predicting a dangerous situation of the object at the predicted time based on the predicted image frame.
[0017] In one embodiment, the step of predicting the risk situation of the object based on the encoded time series features by utilizing the time series prediction model may include the step of inputting the encoded time series features into a classifier of the time series prediction model, and the step of predicting the risk situation of the object based on the encoded time series features by utilizing the classifier of the time series prediction model.
[0018] In one embodiment, the step of predicting a risk situation of the object based on the encoded time series features using the time series prediction model may include at least one of a step of predicting an event occurrence related to the risk situation of the object, or a step of determining a risk level for the risk situation of the object.
[0019] A computer program stored in a computer-readable storage medium is disclosed for solving the above-described problem. When the computer program is executed by one or more processors, the computer program causes the one or more processors to perform operations for analyzing a plurality of image frames to predict a dangerous situation of an object, the operations including: an operation for acquiring a plurality of image frames; an operation for acquiring physical information of an object included in the plurality of image frames; an operation for inputting the plurality of image frames and the physical information of the object into a time-series prediction model; and an operation for acquiring encoded time-series features based on the plurality of image frames and the physical information of the object by utilizing the time-series prediction model.
[0020] In addition, a computing device is disclosed for solving the above-described problem. The computing device includes at least one processor and a memory, and the at least one processor is configured to acquire a plurality of image frames, acquire physical information of an object included in the plurality of image frames, input the plurality of image frames and the physical information of the object into a time-series prediction model, and, by utilizing the time-series prediction model, acquire encoded time-series features based on the plurality of image frames and the physical information of the object.
[0021]
[0022] The present disclosure can accurately predict a dangerous situation of an object by predicting a dangerous situation of an object by utilizing a plurality of image frames and physical information of the object included in the plurality of image frames.
[0023] Furthermore, by incorporating crowd dynamics into time-series prediction models, we can predict the risk of an object by reflecting the characteristics of crowd dynamics that change over time. Therefore, we can build a predictive model that is robust to changes in data distribution due to environmental changes.
[0024] Meanwhile, the effects of the present disclosure are not limited to the effects mentioned above, and various effects may be included within a range apparent to those skilled in the art from the contents described below.
[0025] FIG. 1 is a block diagram of a computing device performing operations according to one embodiment of the present disclosure.
[0026] FIG. 2 is a schematic diagram illustrating a neural network according to one embodiment of the present disclosure.
[0027] FIG. 3 is a flowchart illustrating a method for predicting a risk situation of an object by analyzing a plurality of image frames according to an embodiment of the present disclosure.
[0028] FIG. 4 is a block diagram schematically illustrating a risk situation prediction system according to one embodiment of the present disclosure.
[0029] FIG. 5 is a diagram specifically illustrating a time series prediction model according to an embodiment of the present disclosure.
[0030] FIG. 6 is a block diagram schematically illustrating a first physics-based model according to an embodiment of the present disclosure.
[0031] FIG. 7 is a block diagram schematically illustrating a second physics-based model according to an embodiment of the present disclosure.
[0032] FIGS. 8 and 9 are drawings for explaining how a predictor according to one embodiment of the present disclosure predicts a dangerous situation of an object.
[0033] FIG. 10 is a diagram illustrating a method for a classifier according to an embodiment of the present disclosure to predict a risk situation of an object.
[0034] FIG. 11 is a simplified, general schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented.
[0035] Various embodiments are now described with reference to the drawings. In this disclosure, various descriptions are provided to facilitate understanding of the present disclosure. However, it will be apparent that these embodiments may be practiced without these specific descriptions.
[0036] The terms "component," "module," "system," and the like, as used herein, refer to computer-related entities, hardware, firmware, software, a combination of software and hardware, or an execution of software. For example, a component may be, but is not limited to, a procedure running on a processor, a processor, an object, a thread of execution, a program, and / or a computer. For example, both an application running on a computing device and the computing device may be a component. One or more components may reside within a processor and / or a thread of execution. A component may be localized within a single computer. A component may be distributed between two or more computers. Furthermore, these components may execute from various computer-readable media having various data structures stored therein. Components may communicate via local and / or remote processes, for example, by signals comprising one or more data packets (e.g., data from one component interacting with another component in a local system, a distributed system, and / or data transmitted to another system via a network such as the Internet via signals).
[0037] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from context, "X utilizes A or B" is intended to mean one of its natural inclusive permutations. That is, if X utilizes A; X utilizes B; or X utilizes both A and B, "X utilizes A or B" can apply to any of these cases.
[0038] Additionally, the terms "comprises" and / or "comprising" should be understood to imply the presence of the features and / or components. However, it should be understood that the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other features, components, and / or groups thereof. Furthermore, unless otherwise specified or clear from context to refer to the singular form, the singular in the present disclosure and claims should generally be construed to mean "one or more."
[0039] And, the term "at least one of A or B" should be interpreted to mean "if it includes only A", "if it includes only B", or "if it is combined in the composition of A and B".
[0040] Those skilled in the art should further appreciate that the various illustrative logical blocks, configurations, modules, circuits, means, logics, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate the interchangeability of hardware and software, various illustrative components, blocks, configurations, means, logics, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0041] The description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art. The general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Thus, the present disclosure is not limited to the embodiments set forth herein. The present disclosure is to be construed in the widest scope consistent with the principles and novel features disclosed herein.
[0042]
[0043] FIG. 1 is a block diagram of a computing device performing operations according to one embodiment of the present disclosure.
[0044] The configuration of the computing device (100) illustrated in FIG. 1 is merely a simplified example. In one embodiment of the present disclosure, the computing device (100) may include other configurations for performing the computing environment of the computing device (100), and only some of the disclosed configurations may constitute the computing device (100).
[0045] A computing device (100) may include a processor (110), memory (130), and network unit (150).
[0046] The processor (110) may be configured with one or more cores, and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU) of a computing device. The processor (110) may read a computer program stored in the memory (130) and perform data processing for machine learning according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the processor (110) may perform operations for learning a neural network model. The processor (110) may perform calculations for learning a neural network model, such as processing input data for learning in deep learning (DL), extracting features from input data, calculating errors, and updating weights of a neural network model using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor (110) may process learning of the neural network model. For example, a CPU and a GPGPU can work together to train a neural network model and classify data using the neural network model. Furthermore, in one embodiment of the present disclosure, processors of multiple computing devices can be used together to train a neural network model and classify data using the neural network model. Furthermore, a computer program executed on a computing device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.
[0047]
[0048] A computing device (100) according to one embodiment of the present disclosure may be a risk situation prediction system. The risk situation prediction system may acquire a plurality of image frames collected from a camera, such as a CCTV, and analyze the acquired plurality of image frames to predict a risk situation of an object.
[0049] The risk situation prediction system can acquire physical information about an object contained in the plurality of image frames. The physical information may include static information about the object and dynamic information about the object. Static information about the object may include the object's volume, angle, appearance, shape, etc., while dynamic information about the object may include the object's moving speed, direction of movement, and behavioral pattern.
[0050] The risk situation prediction system can obtain static information about objects included in image frames from previous time points at each point in time. Furthermore, the computing device (100) can integrate dynamic information about objects included in image frames from previous time points at each point in time using cumulative operations, and obtain dynamic information about the integrated objects.
[0051] A risk situation prediction system can predict the risk situation of an object by utilizing multiple image frames and physical information about the object contained in the multiple image frames. Specifically, the risk situation prediction system can predict the risk situation of an object by utilizing multiple image frames, static information about the object contained in image frames from previous points in time, and dynamic information about the object integrated. Therefore, the risk situation of an object can be predicted more accurately than when the risk situation of an object is predicted using only multiple image frames.
[0052] Furthermore, by incorporating crowd dynamics into time-series prediction models, we can predict the risk of an object by reflecting the characteristics of crowd dynamics that change over time. Therefore, we can build a predictive model that is robust to changes in data distribution due to environmental changes.
[0053]
[0054] According to one embodiment of the present disclosure, the memory (130) can store any form of information generated or determined by the processor (110) and any form of information received by the network unit (150).
[0055] According to one embodiment of the present disclosure, the memory (130) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. The computing device (100) may also operate in relation to web storage that performs the storage function of the memory (130) on the internet. The description of the above-described memory is merely an example, and the present disclosure is not limited thereto.
[0056] The network unit (150) according to one embodiment of the present disclosure can use various wired communication systems such as a public switched telephone network (PSTN), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and a local area network (LAN).
[0057] In addition, the proposed network unit (150) according to one embodiment of the present disclosure can use various wireless communication systems such as CDMA (Code Division Multi Access), TDMA (Time Division Multi Access), FDMA (Frequency Division Multi Access), OFDMA (Orthogonal Frequency Division Multi Access), SC-FDMA (Single Carrier-FDMA) and other systems.
[0058] In one embodiment, the network unit (150) may be configured regardless of the communication mode, such as wired or wireless, and may be configured as various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the network may be the well-known World Wide Web (WWW), and may also utilize a wireless transmission technology used for short-distance communication, such as infrared (IrDA) or Bluetooth. The technologies described in the present disclosure may be used not only in the networks mentioned above but also in other networks.
[0059]
[0060] FIG. 2 is a schematic diagram illustrating a neural network according to one embodiment of the present disclosure.
[0061] Throughout this disclosure, the terms "neural network," "neural network model," and "neural network" may be used interchangeably. A neural network may be comprised of a set of interconnected computational units, generally referred to as nodes. These nodes may also be referred to as neurons. A neural network comprises at least one node. The nodes (or neurons) comprising the neural network may be interconnected by one or more links.
[0062] At this time, within the neural network model, one or more nodes connected through links can relatively form a relationship between input nodes and output nodes. The concept of input nodes and output nodes is relative, and any node in an output node relationship with respect to one node can also be in an input node relationship with respect to another node, and vice versa. As described above, the relationship between input nodes and output nodes can be created based on links. One or more output nodes can be connected to one input node through links, and vice versa.
[0063] In a relationship between input nodes and output nodes connected through a single link, the data of the output node can have its value determined based on the data input to the input node. Here, the link interconnecting the input nodes and the output nodes can have a weight (in this case, parameters and weights can be used with the same meaning throughout the present disclosure). The weight can be variable and can be varied by a user or an algorithm so that the neural network model can perform a desired function. For example, when one or more input nodes are interconnected to one output node through respective links, the output node can determine the output node value based on the values input to the input nodes connected to the output node and the weights set for the links corresponding to the respective input nodes.
[0064] As described above, a neural network is a network in which one or more nodes are interconnected through one or more links, forming input and output node relationships within the network. The characteristics of a neural network can be determined based on the number of nodes and links within the network, the relationships between the nodes and links, and the weights assigned to each link. For example, if two neural networks have the same number of nodes and links but different weight values for the links, the two neural networks can be perceived as different from each other.
[0065] A neural network can be composed of a set of one or more nodes. A subset of the nodes comprising the neural network can form a layer. Some of the nodes comprising the neural network can form a layer based on their distances from the initial input node. For example, a set of nodes that are n distances from the initial input node can form layer n. The distance from the initial input node can be defined by the minimum number of links that must be traversed to reach the node from the initial input node. However, this definition of a layer is arbitrary for illustrative purposes, and the layer order within a neural network can be defined in a different way than described above. For example, a layer of nodes can be defined by its distance from the final output node.
[0066] An initial input node may refer to one or more nodes within a neural network into which data is directly input without going through links with other nodes. Alternatively, within a neural network, it may refer to nodes that do not have other input nodes connected by links in the relationship between nodes based on links. Similarly, a final output node may refer to one or more nodes within a neural network that do not have output nodes in their relationship with other nodes. Furthermore, a hidden node may refer to nodes that constitute a neural network other than the initial input node and the final output node.
[0067] A neural network according to one embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be the same as the number of nodes in an output layer, and the number of nodes decreases and then increases as it progresses from the input layer to the hidden layer. In addition, a neural network according to another embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be less than the number of nodes in an output layer, and the number of nodes decreases as it progresses from the input layer to the hidden layer. In addition, a neural network according to another embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be greater than the number of nodes in an output layer, and the number of nodes increases as it progresses from the input layer to the hidden layer. A neural network according to another embodiment of the present disclosure may be a neural network in a combined form of the neural networks described above.
[0068] A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to input and output layers. Using a deep neural network, one can identify latent structures in data. That is, one can identify latent structures in photos, text, videos, voices, and music (e.g., what objects are in a photo, what the content and emotion of a text are, what the content and emotion of a voice are, etc.). A deep neural network can include a convolutional neural network (CNN), a recurrent neural network (RNN), an autoencoder, a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, a generative adversarial network (GAN), and the like. The description of the above-described deep neural network is merely an example, and the present disclosure is not limited thereto.
[0069] In one embodiment of the present disclosure, the network function may include an autoencoder. An autoencoder may be a type of artificial neural network that outputs output data similar to input data. The autoencoder may include at least one hidden layer, and an odd number of hidden layers may be arranged between input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding), and then expanded symmetrically from the bottleneck layer to the output layer (symmetrical to the input layer). The autoencoder may perform nonlinear dimensionality reduction. The number of input layers and output layers may correspond to the dimensionality after preprocessing of the input data. In the autoencoder structure, the number of nodes in the hidden layer included in the encoder may have a structure in which the number of nodes decreases as it moves away from the input layer. The number of nodes in the bottleneck layer (the layer with the fewest nodes between the encoder and decoder) may be kept above a certain number (e.g., more than half of the input layer), as too few nodes may not transmit enough information.
[0070] A neural network model including a neural network can be trained using at least one of supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Training a neural network model can be a process of applying knowledge to the neural network model to perform a specific action.
[0071] Neural network models can be trained to minimize output errors. Training involves repeatedly inputting training data into the neural network model, calculating the neural network model output and target error for the training data, and backpropagating the neural network model error from the output layer to the input layer to update the weights of each node in the neural network model to reduce the error. In supervised learning, training data with the correct answer labeled for each training data is used (i.e., labeled training data). In unsupervised learning, the correct answer may not be labeled for each training data. For example, in the case of supervised learning for data classification, the training data may be data in which each training data category is labeled. Labeled training data is input to the neural network model, and the error can be calculated by comparing the output (category) of the neural network model with the labels of the training data. As another example, in unsupervised learning for data classification, the error can be calculated by comparing the input training data with the neural network model output. The calculated error is backpropagated in the neural network model in the backward direction (i.e., from the output layer to the input layer), and the connection weights of each node in each layer of the neural network model can be updated according to the backpropagation. The amount of change in the connection weights of each node to be updated can be determined by the learning rate. The calculation of the neural network model for the input data and the backpropagation of the error can constitute an epoch. The learning rate can be applied differently depending on the number of iterations of the epoch of the neural network model. For example, a high learning rate can be used in the early stage of training a neural network model so that the neural network model quickly achieves a certain level of performance, thereby increasing efficiency, and a low learning rate can be used in the later stage of training to increase accuracy.
[0072] In neural network model training, the training data can typically be a subset of the actual data (i.e., the data to be processed using the trained neural network model). Therefore, there may be epochs where the error on the training data decreases but the error on the actual data increases. Overfitting is a phenomenon where the error on the actual data increases due to excessive training on the training data. For example, a neural network model trained on yellow cats may fail to recognize cats of any color other than yellow, which could be a form of overfitting. Overfitting can increase the error in machine learning algorithms. Various optimization methods can be used to prevent overfitting. These methods include increasing the training data, regularization, dropout (inactivating some nodes in the network during the learning process), and batch normalization.
[0073]
[0074] FIG. 3 is a flowchart illustrating a method for predicting a risk situation of an object by analyzing a plurality of image frames according to an embodiment of the present disclosure.
[0075] Referring to FIG. 3, the computing device (100) can acquire a plurality of image frames (S110). The computing device (100) can acquire a plurality of image frames collected from a camera such as a CCTV. When a video captured with a first number of frames per second is collected, the computing device (100) can acquire all of the first number of frames every second, or can sample and acquire a second number of frames smaller than the first number of frames among the first number of frames.
[0076] The computing device (100) can obtain physical information of an object included in the plurality of image frames (S120). The computing device (100) can obtain static information of an object included in each of the plurality of image frames by utilizing a first physics-based model. The static information of the object can include a volume of the object, an angle of the object, an appearance of the object, a shape of the object, etc. The computing device (100) can obtain static information of an object included in an image frame of a previous time point at each time point.
[0077] In addition, the computing device (100) can obtain dynamic information of an object included in each of the plurality of image frames by utilizing the second physics-based model. The dynamic information of the object can include the object's moving speed, the object's moving direction, the object's behavior pattern, etc. At each point in time, the computing device (100) can obtain the integrated object's dynamic information by integrating the object's dynamic information included in the image frames up to the previous point in time using a cumulative operation.
[0078] That is, the computing device (100) can model crowd dynamics such as crowd movement, interaction, and density based on the physical information of the object.
[0079] A computing device (100) can input a plurality of image frames and physical information of objects contained in the plurality of image frames into a time series prediction model (S130). The process of inputting the physical information of objects into the time series prediction model may mean a process of integrating modeled crowd dynamics with the time series prediction model.
[0080] A time series prediction model according to one embodiment of the present disclosure may refer to a model that acquires time series data prior to a prediction point in time, acquires an internal representation based on the time series data, and predicts data after the prediction point in time based on the internal representation. The internal representation may refer to a representation inherent in the data itself or the totality of information contained in the data. For example, the internal representation may be a memory state, a latent vector, a context vector, etc.
[0081] In particular, video prediction, which predicts future frames using previous frames, can be performed using deep learning models such as RNN, CNN, and ViT (Vision Transformer). For example, the time series prediction model may have a multi-RNN (stacked RNN, RNN-RNN-RNN) structure with multiple RNN layers stacked. According to embodiments, the time series prediction model may have a CNN-RNN-CNN structure that projects video frames into a latent space and predicts future latent states using RNN. According to embodiments, the time series prediction model may have a CNN-ViT-CNN structure that introduces ViT to model latent video dynamics. LSTM, GRU, and Spatiotemporal LSTM (ST-LSTM), which are modified RNN models, can be applied to the above models. The time series prediction model is not limited to the above-described structure.
[0082] The computing device (100) can obtain encoded time series features based on a plurality of image frames and physical information of the object by utilizing a time series prediction model (S140). The computing device (100) can predict the risk situation of the object based on the encoded time series features by utilizing the time series prediction model.
[0083] For example, the computing device (100) can utilize the predictor of the time series prediction model to predict an image frame at a predicted time based on the encoded time series features and the final predicted output. The computing device (100) can predict the occurrence of an event related to a dangerous situation of an object based on the predicted image frame. The computing device (100) can generate a safety-related event, such as transmitting a safety-related message, based on the predicted occurrence of the event.
[0084] According to embodiments, the computing device (100) may utilize a classifier of a time series prediction model to determine the risk level of an object's dangerous situation based on the encoded time series features. For example, the computing device (100) may receive encoded time series features and classify the risk level of an object's dangerous situation into one of several classes.
[0085] According to embodiments of the present disclosure, the computing device (100) can predict a dangerous situation of an object by utilizing physical information of the object included in a plurality of image frames, and thus can predict a dangerous situation of the object more accurately than when predicting a dangerous situation of the object by utilizing only a plurality of image frames.
[0086] Additionally, the computing device (100) models crowd dynamics and integrates the modeled crowd dynamics into a time-series prediction model, thereby predicting the risk of an object by reflecting the characteristics of crowd dynamics that change over time. Therefore, this has the effect of building a predictive model that is robust to changes in data distribution due to environmental changes.
[0087]
[0088] FIG. 4 is a block diagram schematically illustrating a risk situation prediction system according to one embodiment of the present disclosure.
[0089] Referring to FIG. 4, the risk situation prediction system (100) may include a time series prediction model (110), a first physics-based model (140), and a second physics-based model (150).
[0090] The first physics-based model (140) can receive a plurality of image frames (FRM) as input and output static information (SI) of an object contained in each of the plurality of image frames (FRM). As described above, the static information (SI) of an object can include the volume, angle, appearance, shape, etc. of the object.
[0091] For example, the first physics-based model (140) can analyze the optical properties of an object in an image to determine the shape or appearance of the object. The optical properties of the object may include the object's color, reflection, refraction, etc. However, the present invention is not limited thereto, and the first physics-based model (140) may be a model utilizing image processing and computational vision technologies such as geometric modeling techniques, outline extraction algorithms, and contour and moment calculation algorithms.
[0092] The second physics-based model (150) can receive a plurality of image frames (FRM) as input and acquire dynamic information about an object contained in each of the plurality of image frames (FRM). As described above, the dynamic information about an object can include the object's movement speed, movement direction, behavioral pattern, etc.
[0093] For example, the second physics-based model (150) can utilize optical flow techniques to track pixel movement between image frames and extract the movement distance and direction of a moving object. Optical flow techniques may include algorithms such as Lucas-Kanade, Farneback, and Horn-Schunck.
[0094] According to embodiments, the second physics-based model (150) can model pedestrian movement patterns in crowded environments by utilizing a social force model (SFM) to model the movement patterns of moving objects based on their interactions with each other. Additionally, the second physics-based model (150) may be a model utilizing the Helmholtz decomposition theorem, etc.
[0095] The second physics-based model (150) can integrate dynamic information of an object using cumulative operations and output dynamic information (DI) of the integrated object.
[0096] A time series prediction model (110) may include an encoder (120) and a decoder (130). The time series prediction model (110) may obtain k sequential image frames (St-1, St-2, St-3, ..., St-k; FRM) and physical information (SI, DI) of an object. The time series prediction model (110) may be a model that predicts the future by analyzing the image frames (FRM) and physical information (SI, DI) of an object.
[0097] More specifically, the computing device (100) can model crowd dynamics, such as crowd movement, interaction, and density, based on the physical information (SI, DI) of the object, and integrate the modeled crowd dynamics with a time-series prediction model. That is, when predicting the future based on image frames (FRM), the computing device (100) can utilize the physical information (SI, DI) of the object, which is the output of the first physics-based model (140) and the second physics-based model (150), as additional characteristics. The computing device (100) can integrate the modeled crowd dynamics into the time-series prediction model by taking temporal and spatial interactions into consideration.
[0098] The encoder (120) of the time series prediction model (110) can output encoded time series features (160) based on image frames (FRM) and physical information (SI, DI) of the object. The decoder (130) of the time series prediction model (110) can include at least one of a predictor and a classifier.
[0099] The predictor can predict an image frame (170) at a prediction time point based on the encoded time series features (160) and the final prediction output. The time series prediction model may be a single-step time series prediction model that predicts only one time step at a time, or a multi-step time series prediction model that predicts multiple time steps at a time. When the time series prediction model is a multi-step time series prediction model, the predictor can predict an image frame (170) at each of a plurality of prediction time points based on the encoded time series features (160) and the final prediction output.
[0100] The image frame at the prediction point can be used to predict the occurrence of an event related to a dangerous situation of an object through a data processing process. For example, the computing device (100) can perform a data processing process on the image frame at the prediction point and calculate a crowd density as a result of the processing. The computing device (100) can predict the occurrence of an event indicating that a dangerous situation will occur based on the calculated crowd density. The computing device (100) can generate a safety-related event, such as transmitting a safety-related message, based on the predicted event occurrence.
[0101] The classifier can determine the risk level (170) of an object's dangerous situation based on encoded time series features (160). For example, the classifier can receive encoded time series features as input and classify the risk level of an object's dangerous situation into one of several classes.
[0102]
[0103] FIG. 5 is a diagram specifically illustrating a time series prediction model according to an embodiment of the present disclosure, FIG. 6 is a block diagram briefly illustrating a first physics-based model according to an embodiment of the present disclosure, and FIG. 7 is a block diagram briefly illustrating a second physics-based model according to an embodiment of the present disclosure.
[0104] First, referring to FIG. 5, the encoder (120) of the time series prediction model (110) may include a plurality of encoder cells (121-126). Image frames (St-k, ..., St-4, St-3, St-2, St-1, St) prior to the prediction time may be input to the encoder (120). In addition, the image frames (St-k, ..., St-4, St-3, St-2, St-1) may be input to each of the first physics-based model (MODEL1) and the second physics-based model (MODEL2).
[0105] The first physics-based model (MODEL1) can input image frames (St-k, ..., St-4, St-3, St-2, St-1) and output static information (..., SIt-5, SIt-4, SIt-3, SIt-2, SIt-1) of objects contained in each of the image frames (St-k, ..., St-4, St-3, St-2, St-1).
[0106] Referring to the example illustrated in FIG. 6, the first physics-based model (MODEL1) can receive an image frame (St-4) at time t-4, extract static information (SIt-4) of an object included in the image frame (St-4) at time t-4, and output the extracted static information (SIt-4) of the object to an encoder cell (123) corresponding to time t-3.
[0107] In addition, the first physics-based model (MODEL1) can receive an image frame (St-3) at time t-3, extract static information (SIt-3) of an object included in the image frame (St-3) at time t-3, and output the extracted static information (SIt-3) of the object to an encoder cell (124) corresponding to time t-2.
[0108] In addition, the first physics-based model (MODEL1) can receive an image frame (St-2) at time t-2 as input, extract static information (SIt-2) of an object included in the image frame (St-2) at time t-2, and output the extracted static information (SIt-2) of the object to an encoder cell (125) corresponding to time t-1.
[0109] In addition, the first physics-based model (MODEL1) can receive an image frame (St-1) at time t-1 as input, extract static information (SIt-1) of an object included in the image frame (St-1) at time t-1, and output the extracted static information (SIt-1) of the object to an encoder cell (126) corresponding to time t.
[0110] In the embodiment illustrated in FIG. 6, each of the plurality of modules included in the first physics-based model (MODEL1) is illustrated as processing each of the plurality of image frames in parallel, but an embodiment in which one module sequentially receives and serially processes the plurality of image frames is also possible.
[0111] Referring back to FIG. 5, the second physics-based model (MODEL2) can receive image frames (St-k, ..., St-4, St-3, St-2, St-1) as input and obtain dynamic information of objects included in each of the image frames (St-k, ..., St-4, St-3, St-2, St-1). The second physics-based model (MODEL2) can integrate the dynamic information of objects using a cumulative operation and output the integrated dynamic information of objects (..., DIt-5, DIt-4, DIt-3, DIt-2, DIt-1).
[0112] Referring to the example illustrated in FIG. 7, the second physics-based model (MODEL2) can receive image frames (St-k, ..., St-4) up to time point t-4 and extract dynamic information of objects included in each of the image frames (St-k, ..., St-4). The second physics-based model (MODEL2) can integrate the extracted dynamic information of objects to obtain dynamic information (DIt-4) of the integrated object, and output the dynamic information (DIt-4) of the integrated object to the encoder cell (123) corresponding to time point t-3.
[0113] In addition, the second physics-based model (MODEL2) can receive image frames (St-k, ..., St-4, St-3) up to time t-3 and extract dynamic information of objects included in each of the image frames (St-k, ..., St-4, St-3). The second physics-based model (MODEL2) can integrate the extracted dynamic information of objects to obtain integrated dynamic information (DIt-3) of the object, and output the integrated dynamic information (DIt-3) of the object to the encoder cell (124) corresponding to time t-2.
[0114] In addition, the second physics-based model (MODEL2) can receive image frames (St-k, ..., St-4, St-3, St-2) up to time t-2 and extract dynamic information of objects included in each of the image frames (St-k, ..., St-4, St-3, St-2). The second physics-based model (MODEL2) can integrate the extracted dynamic information of objects to obtain integrated dynamic information (DIt-2) of the object, and output the integrated dynamic information (DIt-2) of the object to the encoder cell (125) corresponding to time t-1.
[0115] In addition, the second physics-based model (MODEL2) can receive image frames (St-k, ..., St-4, St-3, St-2, St-1) up to time point t-1 and extract dynamic information of objects included in each of the image frames (St-k, ..., St-4, St-3, St-2, St-1). The second physics-based model (MODEL2) can integrate the extracted dynamic information of objects to obtain integrated dynamic information (DIt-1) of the object, and output the integrated dynamic information (DIt-1) of the object to the encoder cell (125) corresponding to time point t.
[0116] In the embodiment illustrated in FIG. 7, each of the plurality of modules included in the second physics-based model (MODEL2) is illustrated as receiving image frames up to a previous point in time and processing them in parallel. However, an embodiment in which one module receives image frames up to a previous point in time and processes them serially is also possible.
[0117] Referring back to FIG. 5, static information of an object included in an image frame at a given time point and an image frame at a previous time point can be inputted into each of the plurality of encoder cells (121-126). For example, static information (SIt-2) of an object included in an image frame (St-1) at a given time point and an image frame (St-2) at a previous time point (t-2) can be inputted into an encoder cell (125) corresponding to time point t-1.
[0118] Similarly, static information (SIt-1) of an object included in the image frame (St) of the corresponding time point and the image frame (St-1) of the previous time point (t-1) can be input to the encoder cell (126) corresponding to the time point t.
[0119] Each of the plurality of encoder cells (121-126) can receive static information about an object included in an image frame of the current time point and an image frame of a previous time point and predict an image frame of the next time point.
[0120] According to embodiments, each of the plurality of encoder cells (121-126) may be input with the image frame of the corresponding time, the static information of the object included in the image frame of the previous time, and the dynamic information of the integrated object. For example, the encoder cell (125) corresponding to the time t-1 may be input with the image frame (St-1) of the corresponding time, the static information (SIt-12) of the object included in the image frame (St-2) of the previous time, and the dynamic information (DIt-2) of the integrated object. Here, the dynamic information (DIt-2) of the integrated object may be information that integrates the dynamic information of the object included in each of the image frames (St-k, ..., St-4, St-3, St-2) up to the previous time (t-2) using a cumulative operation.
[0121] Likewise, the encoder cell (126) corresponding to time point t may be input with the image frame (St) of the time point, the static information (SIt-1) of the object included in the image frame (St-1) of the previous time point, and the dynamic information (DIt-1) of the integrated object. Here, the dynamic information (DIt-1) of the integrated object may be information in which the dynamic information of the object included in each of the image frames (St-k, ..., St-4, St-3, St-2, St-1) up to the previous time point (t-1) is integrated using a cumulative operation.
[0122] Each of the plurality of encoder cells (121-126) can receive an image frame of the current time, static information of an object included in an image frame of a previous time, and dynamic information of an integrated object to predict an image frame of the next time.
[0123] The encoder (120) of the time series prediction model can output encoded time series features (160) and a final prediction output (St+1) based on image frames (St-k, ..., St-4, St-3, St-2, St-1, St) and physical information of an object (..., SIt-5, SIt-4, SIt-3, SIt-2, SIt-1, and ..., DIt-5, DIt-4, DIt-3, DIt-2, DIt-1).
[0124] The decoder (130) of the time series prediction model may include at least one of a predictor and a classifier. The predictor may predict an image frame at a prediction time based on the encoded time series features (160) and the final prediction output (St+1), and may predict the occurrence of an event related to the object's risk situation based on the predicted image frame. The classifier may determine the risk level of the object's risk situation based on the encoded time series features (160).
[0125]
[0126] FIG. 8 and FIG. 9 are diagrams for explaining how a predictor according to an embodiment of the present disclosure predicts a dangerous situation of an object, and FIG. 10 is a diagram for explaining how a classifier according to an embodiment of the present disclosure predicts a dangerous situation of an object.
[0127] First, referring to FIG. 8, the encoder (120) of the time series prediction model (110) can input image frames (St-k, ..., St-4, St-3, St-2, St-1, St) prior to the prediction time and output encoded time series features (160) and the final prediction output (St+1). That is, each of the plurality of encoder cells of the time series prediction model (110) can input actual (true) image frames (St-k, ..., St-4, St-3, St-2, St-1, St) of the corresponding time and output encoded time series features (160) and the final prediction output (St+1). Additionally, each of the plurality of encoder cells of the time series prediction model (110) can utilize the physical information (SI, DI) of the object as an additional feature. In this case, the physical information (SI, DI) of the object can be used to model crowd dynamics. The computing device (100) can integrate the modeled crowd dynamics into a time series prediction model, taking temporal and spatial interactions into account.
[0128] The decoder (130) of the time series prediction model (110) may include a plurality of decoder cells. The decoder (130) of the time series prediction model (110) may receive encoded time series features (160) and a final prediction output (St+1) as input and predict an image frame (St+2) of a prediction time point (t+1). When the time series prediction model (110) is a multi-step time series prediction model that predicts multiple time steps at once, the decoder (130) of the time series prediction model (110) may receive encoded time series features (160) and a final prediction output (St+1) as input and predict an image frame (St+2, ..., St+n+1; FRM_P) for each of a plurality of prediction time points (t+1, ..., t+n).
[0129] That is, each of the plurality of decoder cells of the time series prediction model (110) can receive encoded time series features (160) and prediction frames (St+1, St+2, ..., St+n) of previous time points as input and output image frames (St+2, ..., St+n+1) of prediction time points (t+1, ..., t+n).
[0130] The encoder (120) operates based on the actual (true) image frames (St-k, ..., St-4, St-3, St-2, St-1, St) of the corresponding time point, and the decoder (130) can perform prediction based on the prediction frames (St+1, St+2, ..., St+n) of the previous time point.
[0131]
[0132] Referring to FIG. 9, the computing device (100) performs a data processing process on image frames (St+2, St+3, ... St+n+1; FRM_P) for prediction time points (t+1, ..., t+n), and can predict the occurrence of an event related to a dangerous situation of an object based on the result of the processing.
[0133] For example, the computing device (100) can calculate crowd density through a data processing process for image frames (St+2, St+3, ... St+n+1; FRM_P) for predicted time points (t+1, ..., t+n). Here, crowd density refers to the density of a population gathered in a specific area or place, and the higher the crowd density, the higher the risk of an accident occurring.
[0134] The computing device (100) can detect objects in an image using an object detection algorithm and estimate crowd density using the location of the objects. According to embodiments, the computing device (100) can estimate crowd density by analyzing the color or pixel value distribution of the image, or can calculate crowd density using a statistical method.
[0135] The computing device (100) can determine whether the calculated crowd density reaches a preset level. If the calculated crowd density is determined to be above the preset level, the computing device (100) can predict the occurrence of an event indicating that a dangerous situation will occur.
[0136] The computing device (100) may generate a safety-related event, such as transmitting a safety-related message based on a prediction of an event occurrence (OUT1). For example, the computing device (100) may transmit a safety-related message such as "There is an expectation of a large concentration of people in the OO area, so please use a different route" based on a prediction of an event occurrence (OUT1) indicating that a dangerous situation will occur.
[0137] Referring to Figure 10, the classifier can determine the risk level of an object's risk situation based on encoded time-series features. For example, the classifier can receive encoded time-series features as input and classify the risk level of an object's risk situation into one of several classes.
[0138] The classifier may include a fully connected layer and an output layer. The fully connected layer may consist of a number of neurons equal to the number of classes to be classified. The output layer may be implemented using a softmax function that estimates the probability for each class.
[0139] For example, a classifier can classify the risk level of an object's dangerous situation into three classes: non-risky (CLASS1), low-risk (CLASS2), and high-risk (CLASS3). In the example described with reference to FIG. 9, in a situation where the crowd density is above a preset level, the classifier can classify the risk level of an object's dangerous situation as high-risk (CLASS3).
[0140]
[0141] According to one embodiment of the present disclosure, a computer-readable medium storing a data structure is disclosed.
[0142] A data structure can refer to the organization, management, and storage of data to enable efficient access and modification. A data structure can refer to the organization of data to solve a specific problem (e.g., data retrieval, data storage, or data modification in the shortest possible time). A data structure can also be defined as the physical or logical relationships between data elements designed to support specific data processing functions. The logical relationships between data elements can include connections between user-defined data elements. The physical relationships between data elements can include actual relationships between data elements physically stored on a computer-readable storage medium (e.g., persistent storage). Specifically, a data structure can include a collection of data, relationships between data, and functions or commands applicable to the data. An effectively designed data structure allows a computing device to perform operations while minimizing the use of its resources. Specifically, a computing device can improve the efficiency of operations, reading, inserting, deleting, comparing, exchanging, and searching through an effectively designed data structure.
[0143] Data structures can be categorized as linear or nonlinear, depending on their form. A linear data structure can be a structure in which only one data item is linked to the next. Linear data structures can include lists, stacks, queues, and deques. A list can refer to a series of data sets with an internal order. Lists can also include linked lists. A linked list is a data structure in which data is linked in a single line, each item having a pointer. In a linked list, a pointer can contain information about the next or previous item. Linked lists can be expressed as singly linked lists, doubly linked lists, or circular linked lists, depending on their form. A stack can be a data listing structure with limited data access. A stack can be a linear data structure in which data operations (e.g., insertion or deletion) can only be performed at one end of the data structure. Data stored in a stack can be a Last-in-First-out (LIFO) data structure. A queue is a data structure with limited access to data. Unlike a stack, it can be a first-in, first-out (FIFO) data structure, with later data being retrieved later. A deck can be a data structure that can process data at both ends.
[0144] A nonlinear data structure can be a structure in which multiple data are connected behind a single data. Nonlinear data structures can include graph data structures. A graph data structure can be defined by vertices and edges, and an edge can include a line connecting two different vertices. A graph data structure can include a tree data structure. A tree data structure can be a data structure in which there is only one path connecting two different vertices among multiple vertices included in the tree. In other words, it can be a data structure that does not form a loop in a graph data structure.
[0145] A data structure may include a neural network model. The data structure including the neural network model may be stored on a computer-readable medium. The data structure including the neural network model may include preprocessed data for processing by the neural network model, data input to the neural network model, weights of the neural network model, hyperparameters of the neural network model, data obtained from the neural network model, activation functions associated with each node or layer of the neural network model, loss functions for learning the neural network model, etc. The data structure including the neural network model may include any of the components among the components disclosed above. That is, the data structure including the neural network model may be configured to include all or any combination of the following: preprocessed data for processing by the neural network model, data input to the neural network model, weights of the neural network model, hyperparameters of the neural network model, data obtained from the neural network model, activation functions associated with each node or layer of the neural network model, loss functions for learning the neural network model, etc. In addition to the above-described components, the data structure including the neural network model may include any other information that determines the characteristics of the neural network model. Additionally, the data structure may include any form of data used or generated in the computational process of the neural network model, and is not limited to the aforementioned. The computer-readable medium may include a computer-readable recording medium and / or a computer-readable transmission medium. The neural network model may be composed of a set of interconnected computational units, which may generally be referred to as nodes. These nodes may also be referred to as neurons. The neural network model is composed of at least one node.
[0146] The data structure may include data input to a neural network model. The data structure including the data input to the neural network model may be stored on a computer-readable medium. The data input to the neural network model may include training data input during the training process of the neural network model and / or input data input to the neural network model after training has been completed. The data input to the neural network model may include data that has undergone preprocessing and / or data that is the target of preprocessing. Preprocessing may include a data processing process for inputting data to the neural network model. Accordingly, the data structure may include data that is the target of preprocessing and data generated by the preprocessing. The above-described data structure is merely an example, and the present disclosure is not limited thereto.
[0147] The data structure may include weights of a neural network model. (In the present disclosure, the terms "weight" and "parameter" may be used interchangeably.) The data structure including the weights of the neural network model may be stored in a computer-readable medium. The neural network model may include a plurality of weights. The weights may be variable and may be varied by a user or an algorithm so that the neural network model can perform a desired function. For example, when one or more input nodes are interconnected to one output node by respective links, the output node may determine a data value output from the output node based on values input to the input nodes connected to the output node and weights set for links corresponding to each input node. The above-described data structure is merely an example, and the present disclosure is not limited thereto.
[0148] By way of example and not limitation, the weights may include weights that vary during the training process of the neural network model and / or weights that have completed training of the neural network model. The weights that vary during the training process of the neural network model may include weights at the start of an epoch and / or weights that vary during the epoch. The weights that have completed training of the neural network model may include weights that have completed an epoch. Accordingly, a data structure including the weights of the neural network model may include a data structure including weights that vary during the training process of the neural network model and / or weights that have completed training of the neural network model. Therefore, the above-described weights and / or combinations of each weight are included in the data structure including the weights of the neural network model. The above-described data structures are merely examples and the present disclosure is not limited thereto.
[0149] A data structure including the weights of a neural network model can be stored in a computer-readable storage medium (e.g., memory, hard disk) after going through a serialization process. Serialization can be a process of converting a data structure into a form that can be stored on the same or a different computing device and later reconstructed and used. A computing device can serialize the data structure to transmit and receive data over a network. The data structure including the weights of a serialized neural network model can be reconstructed on the same or a different computing device through deserialization. The data structure including the weights of a neural network model is not limited to serialization. Furthermore, the data structure including the weights of a neural network model can include a data structure that increases computational efficiency while minimizing the use of computing device resources (e.g., a B-Tree, a Trie, an m-way search tree, an AVL tree, a Red-Black Tree in nonlinear data structures). The foregoing is merely an example, and the present disclosure is not limited thereto.
[0150] The data structure may include hyperparameters of a neural network model. Furthermore, the data structure including the hyperparameters of the neural network model may be stored on a computer-readable medium. The hyperparameters may be variables that can be varied by the user. The hyperparameters may include, for example, a learning rate, a cost function, a loss function, the number of epoch iterations, weight initialization (e.g., setting a range of weight values to be subject to weight initialization), and the number of hidden units (e.g., the number of hidden layers, the number of nodes in the hidden layer). The above-described data structure is merely an example, and the present disclosure is not limited thereto.
[0151]
[0152] FIG. 11 is a simplified, general schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented.
[0153] Although the present disclosure has been described above as being generally implemented by a computing device, those skilled in the art will appreciate that the present disclosure may also be implemented in combination with computer-executable instructions and / or other program modules that may be executed on one or more computers and / or as a combination of hardware and software.
[0154] Generally, program modules include routines, programs, components, data structures, and the like that perform particular tasks or implement particular abstract data types. Furthermore, those skilled in the art will appreciate that the methods of the present disclosure can be implemented with other computer system configurations, including single-processor or multiprocessor computer systems, minicomputers, mainframe computers, as well as personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, and the like, each of which may be operatively connected to one or more associated devices.
[0155] The described embodiments of the present disclosure can be practiced in a distributed computing environment, where certain tasks are performed by remote processing devices that are connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0156] Computers typically include a variety of computer-readable media. Computer-readable media can be any media that can be accessed by a computer, and includes both volatile and nonvolatile media, transitory and non-transitory media, removable and non-removable media. By way of example, and not limitation, computer-readable media can include computer-readable storage media and computer-readable transmission media. Computer-readable storage media includes both volatile and nonvolatile media, transitory and non-transitory media, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital video disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be accessed by a computer and used to store the desired information.
[0157] Computer-readable transmission media typically includes any information delivery media that embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism. The term modulated data signal means a signal that has one or more of its characteristics set or changed so as to encode information in the signal. By way of example, and not limitation, computer-readable transmission media includes wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared, or other wireless media. Combinations of any of the above are also intended to be included within the scope of computer-readable transmission media.
[0158] An exemplary environment for implementing various aspects of the present disclosure is illustrated, including a computer (1102), which includes a processing unit (1104), system memory (1106), and a system bus (1108). The system bus (1108) connects system components, including but not limited to the system memory (1106), to the processing unit (1104). The processing unit (1104) may be any of a variety of commercially available processors. Dual processors and other multiprocessor architectures may also be utilized as the processing unit (1104).
[0159] The system bus (1108) may be any of several types of bus structures that may be additionally interconnected to a memory bus, a peripheral bus, and a local bus using any of a variety of commercial bus architectures. The system memory (1106) includes read-only memory (ROM) (1110) and random access memory (RAM) (1112). A basic input / output system (BIOS) is stored in non-volatile memory (1110), such as ROM, EPROM, or EEPROM, and includes basic routines that help transfer information between components within the computer (1102), such as during start-up. The RAM (1112) may include high-speed RAM, such as static RAM, for caching data.
[0160] The computer (1102) includes an internal hard disk drive (HDD) (1114) (e.g., EIDE, SATA) - which may be configured for external use within a suitable chassis (not shown), a magnetic floppy disk drive (FDD) (1116) (e.g., for reading from or writing to a removable diskette (1118)), and an optical disk drive (1120) (e.g., for reading from or writing to a CD-ROM disk (1122) or other high-capacity optical media such as a DVD). The hard disk drive (1114), the magnetic disk drive (1116), and the optical disk drive (1120) may be connected to the system bus (1108) by a hard disk drive interface (1124), a magnetic disk drive interface (1126), and an optical drive interface (1128), respectively. The interface (1124) for implementing an external drive includes at least one or both of Universal Serial Bus (USB) and IEEE 1394 interface technologies.
[0161] These drives and their associated computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, and the like. In the case of the computer (1102), the drives and media correspond to storing any data in a suitable digital format. While the description of computer-readable media above refers to HDDs, removable magnetic disks, and removable optical media such as CDs or DVDs, those of ordinary skill in the art will appreciate that other types of computer-readable media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, and the like, may also be used in the exemplary operating environment, and that any such media may contain computer-executable instructions for performing the methods of the present disclosure.
[0162] A number of program modules, including an operating system (1130), one or more application programs (1132), other program modules (1134), and program data (1136), may be stored in the drive and RAM (1112). All or portions of the operating system, applications, modules, and / or data may be cached in RAM (1112). It will be appreciated that the present disclosure may be implemented in various commercially available operating systems or combinations of operating systems.
[0163] A user may enter commands and information into the computer (1102) via one or more wired / wireless input devices, such as a keyboard (1138) and a pointing device such as a mouse (1140). Other input devices (not shown) may include a microphone, an IR remote control, a joystick, a game pad, a stylus pen, a touch screen, and the like. These and other input devices are often connected to the processing unit (1104) via an input device interface (1142) that is connected to the system bus (1108), but may be connected by other interfaces such as a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR interface, and the like.
[0164] A monitor (1144) or other type of display device is also connected to the system bus (1108) via an interface, such as a video adapter (1146). In addition to the monitor (1144), the computer typically includes other peripheral output devices (not shown), such as speakers, a printer, and so on.
[0165] The computer (1102) may operate in a networked environment using logical connections to one or more remote computers, such as remote computer(s) (1148), via wired and / or wireless communications. The remote computer(s) (1148) may be a workstation, a computing device computer, a router, a personal computer, a portable computer, a microprocessor-based entertainment device, a peer device, or other conventional network node, and generally include many or all of the components described for the computer (1102), although for simplicity, only the memory storage device (1150) is shown. The logical connections shown include wired / wireless connections to a local area network (LAN) (1152) and / or a larger network, such as a wide area network (WAN) (1154). Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks, such as intranets, all of which may be connected to a worldwide computer network, such as the Internet.
[0166] When used in a LAN networking environment, the computer (1102) is connected to a local network (1152) via a wired and / or wireless communication network interface or adapter (1156). The adapter (1156) may facilitate wired or wireless communications to the LAN (1152), which may include a wireless access point installed therein for communicating with the wireless adapter (1156). When used in a WAN networking environment, the computer (1102) may include a modem (1158), be connected to a communications computing device on the WAN (1154), or have other means of establishing communications over the WAN (1154), such as via the Internet. The modem (1158), which may be internal or external and wired or wireless, is connected to the system bus (1108) via a serial port interface (1142). In a networked environment, program modules or portions thereof described for the computer (1102) may be stored in a remote memory / storage device (1150). It will be appreciated that the network connections depicted are exemplary and other means of establishing a communications link between the computers may be used.
[0167] The computer (1102) operates to communicate with any wireless device or object that is arranged and operates via wireless communication, such as a printer, a scanner, a desktop and / or portable computer, a portable data assistant (PDA), a communication satellite, any equipment or location associated with a radio-detectable tag, and a telephone. This includes at least Wi-Fi and Bluetooth wireless technologies. Accordingly, the communication may be a predefined structure as in a conventional network, or may simply be an ad hoc communication between at least two devices.
[0168] Wi-Fi (Wireless Fidelity) enables connections to the Internet and other devices without wires. Wi-Fi is a wireless technology that allows devices, such as computers, to send and receive data anywhere within the coverage area of a base station, both indoors and outdoors, similar to cell phones. Wi-Fi networks use wireless technologies called IEEE 802.11 (a, b, g, etc.) to provide secure, reliable, and high-speed wireless connections. Wi-Fi can be used to connect computers to each other, to the Internet, and to wired networks (using IEEE 802.3 or Ethernet). Wi-Fi networks can operate in the unlicensed 2.4 and 5 GHz radio bands, at data rates of, for example, 11 Mbps (802.11a) or 54 Mbps (802.11b), or in products that include both bands (dual-band).
[0169] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips referenced in the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0170] Those skilled in the art will appreciate that the various illustrative logical blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, various forms of programs or design code (referred to herein, for convenience, as software), or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0171] The various embodiments presented herein can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term article of manufacture includes a computer program, carrier, or media accessible from any computer-readable storage device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Furthermore, various storage media presented herein include one or more devices and / or other machine-readable media for storing information.
[0172] It should be understood that the specific order or hierarchy of steps in the presented processes is merely an example of exemplary approaches. It should be understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of the present disclosure based on design priorities. The appended method claims provide elements of various steps in a sample order, but are not intended to be limited to the specific order or hierarchy presented.
[0173] The description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments disclosed herein, but is to be construed in the broadest scope consistent with the principles and novel features disclosed herein.
[0174]
[0175] As described above, the relevant contents have been described in the best form for carrying out the invention.
Claims
1. A method for predicting a risk situation of an object by analyzing multiple image frames performed by a computing device, A step of obtaining multiple image frames; A step of obtaining physical information of an object included in the plurality of image frames; A step of inputting the plurality of image frames and the physical information of the object into a time series prediction model; and A step of obtaining encoded time series features based on the physical information of the object and the plurality of image frames by utilizing the above time series prediction model. Including, method.
2. In paragraph 1, The step of obtaining physical information of an object included in the above multiple image frames is: A step of obtaining static information of an object included in each of the plurality of image frames by utilizing the first physics-based model. Including, method.
3. In paragraph 2, The step of inputting the above plurality of image frames and the physical information of the object into the time series prediction model is: A step of inputting static information of an image frame at a corresponding time point and an object included in an image frame at a previous time point into each of a plurality of encoder cells included in the encoder of the above time series prediction model. Including, method.
4. In paragraph 2, The step of obtaining physical information of an object included in the above multiple image frames is: A step of obtaining dynamic information of an object included in each of the plurality of image frames by utilizing a second physics-based model. Including more, method.
5. In paragraph 4, The step of inputting the above plurality of image frames and the physical information of the object into the time series prediction model is: A step of inputting, into each of a plurality of encoder cells included in the encoder of the above time series prediction model, an image frame of the corresponding time point, static information of an object included in an image frame of a previous time point, and dynamic information of an integrated object, The dynamic information of the above integrated object is integrated by using cumulative operation, which is the dynamic information of the object included in each image frame up to the previous point in time. method.
6. In paragraph 1, The above method, A step of predicting the risk situation of the object based on the encoded time series features by utilizing the above time series prediction model. Including more, method.
7. In paragraph 6, The step of inputting the above plurality of image frames and the physical information of the object into the time series prediction model is: A step of inputting the plurality of image frames and the physical information of the object into the encoder of the above time series prediction model. Including, method.
8. In paragraph 7, The step of predicting the risk situation of the object based on the encoded time series features by utilizing the above time series prediction model is as follows. A step of inputting the encoded time series features and the final prediction output into the predictor of the time series prediction model; A step of predicting an image frame at a prediction time based on the encoded time series features and the final prediction output by utilizing a predictor of the above time series prediction model; and A step of predicting the risk situation of the object at the predicted time based on the predicted image frame. Including, method.
9. In paragraph 7, The step of predicting the risk situation of the object based on the encoded time series features by utilizing the above time series prediction model is as follows. a step of inputting the encoded time series features into a classifier of the time series prediction model; and A step of predicting the risk situation of the object based on the encoded time series features by utilizing the classifier of the above time series prediction model. Including, method.
10. In paragraph 6, The step of predicting the risk situation of the object based on the encoded time series features by utilizing the above time series prediction model is as follows. A step of predicting the occurrence of an event related to a risk situation of the above object; or Step for determining the risk level of the above object's risk situation Containing at least one of: method.
11. A computer program stored in a computer-readable storage medium, wherein the computer program, when executed by one or more processors, causes the one or more processors to perform operations for analyzing a plurality of image frames to predict a dangerous situation of an object, the operations comprising: An action to acquire multiple image frames; An operation of obtaining physical information of an object included in the plurality of image frames; An operation of inputting the plurality of image frames and the physical information of the object into a time series prediction model; and An operation of obtaining encoded time series features based on the physical information of the object and the plurality of image frames by utilizing the above time series prediction model. Including, A computer program stored on a computer-readable storage medium.
12. In paragraph 11, The operation of obtaining physical information of an object included in the above multiple image frames is: An operation of obtaining static information of an object contained in each of the plurality of image frames by utilizing the first physics-based model. Including, A computer program stored on a computer-readable storage medium.
13. In paragraph 12, The operation of inputting the above plurality of image frames and the physical information of the object into a time series prediction model is as follows. An operation of inputting static information of an object included in an image frame of a corresponding time point and an image frame of a previous time point to each of a plurality of encoder cells included in the encoder of the above time series prediction model. Including, A computer program stored on a computer-readable storage medium.
14. As a computing device, at least one processor; and Memory Including, At least one processor of the above, Obtain multiple image frames, Obtaining physical information of an object contained in the above multiple image frames, Inputting the above plurality of image frames and the physical information of the object into a time series prediction model, and, configured to obtain encoded time series features based on the physical information of the object and the plurality of image frames by utilizing the above time series prediction model. Computing device.
15. In paragraph 14, It is further configured to obtain static information of an object included in each of the plurality of image frames by utilizing the first physics-based model. Computing device.
Citation Information
Patent Citations
Semiconductor Die Formation and Packaging Thereof
KR102227858B1
Automatic Transportation System
KR102297574B1
Method and apparatus for detecting change
KR102345892B1
Canvas fixing frame
KR102763909B1
Training a machine learning system to detect an excursion of a CMP component using time-based sequence of images
US20230316502A1