Method and apparatus for detecting driver distraction in a DMS (Driver Monitoring System).
The method optimizes driver monitoring systems by using AI models to analyze gaze and posture data, reducing false alarms and enhancing reliability through clustering and voting mechanisms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NOTA INC
- Filing Date
- 2024-09-23
- Publication Date
- 2026-05-19
AI Technical Summary
Existing driver monitoring systems (DMS) face issues with frequent false alarms due to sensitive detection algorithms and variable camera positioning, leading to reduced accuracy and reliability, especially under changing lighting and occlusion conditions.
A method using artificial intelligence models to determine gaze classes and yaw/pitch values, combined with voting mechanisms to optimize alarm generation, considering driver attention distraction based on clustering and threshold comparisons.
Enhances the accuracy of driver distraction detection by minimizing false alarms and improving system reliability under varying driving conditions.
Smart Images

Figure 2026515619000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a driver monitoring system, and more specifically, to a method and apparatus for detecting distraction based on the results of driver monitoring.
Background Art
[0002] A driver monitoring system (DMS) is a technology that senses the driver's state to assist in safe driving. The DMS can execute functions such as detecting the driver's drowsiness, distraction, seat belt wearing status, and / or drunk driving, and warning the driver or controlling the vehicle. Such a DMS is evaluated as one of the core technologies of autonomous driving vehicles to ensure the safe operation of the vehicle, and can also contribute to protecting the driver's safety and preventing traffic accidents.
[0003] To sense the driver's state, the DMS can utilize various sensors such as cameras, infrared sensors, acceleration sensors, and gyroscopes. The camera can be used to determine the driver's drowsiness or distraction by sensing the expression on the driver's face and the movement of the eyelids. The infrared sensor can be used to track the driver's line of sight by sensing the movement of the driver's pupils. The acceleration sensor and gyroscope can be used to determine drowsiness or drunk driving by sensing the driver's posture and movement.
[0004] As a method for improving the accuracy and reliability of the DMS, the sensors have been enhanced. When the resolution and image quality of the camera are improved, the driver's expression and eyelid movement can be detected more accurately. When the performance of the infrared sensor is improved, the movement of the driver's pupils can be tracked more accurately. When the performance of the acceleration sensor and gyroscope is improved, the driver's posture and movement can be measured more accurately.
[0005] One factor that improves the performance of DMS is the advancement of artificial intelligence technology. When artificial intelligence technology is applied to DMS, it becomes possible to more accurately sense the driver's condition along with the sophistication of the sensors. [Overview of the project] [Problems that the invention aims to solve]
[0006] However, even if detection technology advances and the driver's state becomes more sophisticated, if the model used in the DMS is not equipped with a separate alarm-related algorithm, alarms may be triggered for every small action the driver takes, potentially interfering with driving. Furthermore, a sensitive driver detection algorithm may actually increase the frequency of false alarms. A high frequency of false alarms can lead to problems with the performance and reliability of the DMS product.
[0007] Furthermore, since aftermarket DMS products cannot specify their installation location, a problem may arise where the position of the camera must be customized for each vehicle based on the driver's position. Also, the size and angle of the face on the camera can change depending on the driver's driving habits and actions, which can reduce the accuracy of the DMS's sensing results. In addition, depending on the various occlusion and lighting conditions that occur during driving, there may be situations where objects that should be detected (e.g., seat belts) are not correctly detected in the image. [Means for solving the problem]
[0008] The inventors of this disclosure provide various embodiments for addressing various technical problems, including various technical problems in the related art and the problems identified by the inventors.
[0009] The various embodiments of this disclosure are the result of efforts to optimize the alarms provided by the DMS.
[0010] One embodiment of this disclosure is made by an effort to accurately detect the driver's state or condition by taking into account various circumstances that occur during the driving process.
[0011] The technical challenges described herein are not limited to those mentioned above, and other technical challenges not explicitly mentioned will be clearly understood by those skilled in the art from the following description.
[0012] One embodiment of the present disclosure discloses a method for detecting driver distraction in a Driver Monitoring System (DMS). The method includes receiving a first image including the driver inside a vehicle; in response to acquiring the first image, using a first model to determine a gaze class from a plurality of gaze classes corresponding to the first image; and determining whether or not driver distraction is present in the first image based on the gaze class corresponding to the first image, wherein the plurality of gaze classes include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking away from the front. Determining whether or not driver distraction is present in the first image includes extracting yaw and pitch values from the first image, and determining whether or not driver distraction is present in the first image using the distance between the gaze cluster corresponding to the first image and the extracted yaw and pitch values from a plurality of gaze clusters generated by clustering selected reference images.
[0013] In one embodiment, determining whether or not the driver's attention is distracted in the first image involves setting a higher attention distraction score indicating the possibility of the driver's attention distraction the greater the distance between at least one cluster belonging to the first gaze class and the yaw value and pitch value extracted from the first image, or the greater the distance between at least one cluster belonging to the second gaze class and the yaw value and pitch value extracted from the first image.
[0014] In one embodiment, the yaw value and pitch value corresponding to the driver's face in the first image are generated by a second model different from the first model, and the second model is an artificial intelligence model that has been pre-trained to output the yaw value and pitch value corresponding to the driver's face in the first image from the first image.
[0015] In one embodiment, the first model corresponds to an artificial intelligence-based model pre-trained to output a gaze class corresponding to the yaw and pitch values, and the distance between the yaw and pitch values and the gaze class, in response to yaw and pitch values extracted from an image. The pre-trained first model is updated by further acquiring the reference images. If the acquired reference images exceed the critical size of the queue of the first model, older reference images are removed from the queue with respect to the time of acquisition of the reference images.
[0016] In one embodiment, the first model corresponds to an artificial intelligence model that has been pre-trained using a training dataset generated based on clustering of a dataset consisting of reference images from a plurality of images that satisfy the condition that the vehicle's driving speed is above a predetermined critical speed.
[0017] In one embodiment, the training dataset is generated by labeling each of the plurality of gaze clusters, which are generated by clustering the reference images, with a first gaze class indicating that the driver is looking forward or a second gaze class indicating that the driver is looking somewhere other than forward, based on quantitative information of the images contained in each of the plurality of gaze clusters.
[0018] In one embodiment, determining whether or not driver distraction exists in the first image includes determining a first attention distraction score corresponding to the first image based on a gaze cluster corresponding to the first image from among a plurality of gaze clusters generated by clustering the selected reference images, and yaw and pitch values extracted from the first image; generating a first primary attention distraction prediction result indicating whether or not attention distraction exists in the first image by comparing the first attention distraction score with a first threshold; and determining an attention distraction alarm corresponding to the first image by performing first voting using the first primary attention distraction prediction result. The first voting generates a group attention distraction prediction result representing the group consisting of the first image and the first number of consecutive images, using the first primary attention distraction prediction results of the first image and a first number of consecutive images selected before the first image, and further generates a first secondary attention distraction prediction result used as a parameter for determining an attention distraction alarm corresponding to the first image. The group attention distraction prediction result includes a result value representing the group, among result values indicating the presence or absence of attention distraction and result values indicating the absence of attention distraction.
[0019] In one embodiment, determining whether or not the driver's attention is distracted in the first image includes determining a first attention distraction score corresponding to the first image based on a gaze cluster corresponding to the first image from among a plurality of gaze clusters generated by clustering the selected reference images, and yaw and pitch values extracted from the first image; generating a first primary attention distraction prediction result indicating whether or not attention is distracted in the first image by comparing the first attention distraction score with a first threshold; generating a first secondary attention distraction prediction result by performing a first voting using the first primary attention distraction prediction result; and determining an attention distraction alarm corresponding to the first image by performing a second voting using the first secondary attention distraction prediction result. The first voting generates a secondary attention distraction prediction result used in the second voting using the primary attention distraction prediction results of each of the plurality of images including the first image, and the second voting is used to determine an attention distraction alarm using the secondary attention distraction prediction results of each of the plurality of images including the first image.
[0020] In one embodiment, the second voting includes determining whether there is continuity in the secondary prediction results of distraction within an image set consisting of the first image and a predetermined second number of sequential images previously acquired for the first image, and determining a distraction alarm corresponding to the first image, which means determining to generate the distraction alarm if the continuity in the secondary prediction results of distraction exists.
[0021] In one embodiment, generating the secondary prediction result for the first distraction includes generating one or more current counter values corresponding to the first image in a manner that increases or decreases one or more previous counter values corresponding to a second image previously acquired based on the results of the first voting, and generating the secondary prediction result for the first distraction that includes the one or more current counter values. Determining the distraction alarm corresponding to the first image includes determining whether one or more distraction alarms corresponding to the first image are ON or OFF by comparing the one or more current counter values with one or more predetermined counter thresholds.
[0022] In one embodiment, determining the first attention-distraction secondary prediction result includes determining at least one current counter value corresponding to the first image by increasing or decreasing at least one previous counter value corresponding to a second image received before the first image based on the result of the first voting, wherein the unit of increase or decrease of the at least one previous counter value is determined based on the time difference between the time of reception of the second image and the time of reception of the first image.
[0023] In one embodiment, the first threshold is determined based on at least one prior primary prediction result of attentional distraction corresponding to at least one prior image previously acquired of the first image.
[0024] In one embodiment, determining the gaze class corresponding to the first image includes extracting yaw and pitch values from the first image, determining the gaze cluster corresponding to the first image from among a plurality of gaze clusters generated by clustering the reference images extracted based on selected conditions, based on the distance between the extracted yaw and pitch values and the clustering results of the reference images, and determining the gaze class among the plurality of gaze classes to which the gaze cluster corresponding to the first image belongs as the gaze class corresponding to the first image.
[0025] In one embodiment, a computer program stored on a computer-readable storage medium is disclosed. The computer program, when executed by at least one processor, allows the at least one processor to perform operations for sensing driver distraction in a Driver Monitoring System (DMS), the operations including receiving a first image including the driver in a vehicle, determining a gaze class from a plurality of gaze classes corresponding to the first image using a first model in response to the acquisition of the first image, and determining whether or not driver distraction is present in the first image based on the gaze class corresponding to the first image, wherein the plurality of gaze classes include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking away from the front. Determining whether or not driver distraction is present in the first image includes extracting yaw and pitch values from the first image, and determining whether or not driver distraction is present in the first image using the distance between the gaze cluster corresponding to the first image and the extracted yaw and pitch values from a plurality of gaze clusters generated by clustering selected reference images.
[0026] In one embodiment, a computing device including at least one processor and a memory is disclosed. The at least one processor receives a first image including a driver in a vehicle, and in response to the acquisition of the first image, uses a first model to determine a line-of-sight class corresponding to the first image among a plurality of line-of-sight classes, and based on the line-of-sight class corresponding to the first image, determines whether there is driver distraction in the first image. Here, the plurality of line-of-sight classes includes a first line-of-sight class in which the driver is looking straight ahead and a second line-of-sight class in which the driver is looking non-straight ahead. Determining whether there is driver distraction in the first image includes extracting a yaw value and a pitch value from the first image, and using a distance between the line-of-sight cluster corresponding to the first image and the extracted yaw value and pitch value among a plurality of line-of-sight clusters generated by clustering a selected reference image to determine whether there is driver distraction in the first image.
Advantages of the Invention
[0027] The technology according to an embodiment of the present disclosure can optimize the alarm provided by the DMS.
[0028] The technology according to an embodiment of the present disclosure can more accurately sense the state of the driver in consideration of various situations occurring during driving.
Brief Description of the Drawings
[0029] [Figure 1] A block configuration diagram of a computing device according to an embodiment of the present disclosure is schematically shown. [Figure 2] An exemplary structure of an artificial intelligence-based model according to an embodiment of the present disclosure is shown. [Figure 3] A method for detecting an abnormality and determining an abnormality alarm in a DMS according to an embodiment of the present disclosure is exemplarily shown. [Figure 4]A method for determining abnormal alarms in a DMS according to one embodiment of the disclosed information is illustrated as an example. [Figure 5] An exemplary method for determining driver distraction according to one embodiment of the content of this disclosure is provided. [Figure 6] This section illustrates the line-of-sight angles based on the camera's placement in the DMS (Digital Measure System). [Figure 7] An exemplary method for collecting training data to train a model according to one embodiment of the disclosure is provided. [Figure 8] This disclosure illustrates a model training and inference method according to one embodiment of the disclosed material. [Figure 9] An exemplary method for determining a driver's distraction alarm according to one embodiment of the content of this disclosure is provided. [Figure 10] An exemplary method for determining a driver's abnormal alarm according to one embodiment of the content of this disclosure is provided. [Figure 11] An exemplary method for determining a driver's abnormal alarm according to one embodiment of the content of this disclosure is provided. [Figure 12] An exemplary method for determining a driver's abnormal alarm according to one embodiment of the content of this disclosure is provided. [Figure 13] This disclosure illustrates a methodology for determining a driver's distraction alarm according to one embodiment of the disclosed information. [Figure 14] This is a schematic diagram of a computing environment according to one embodiment of the information disclosed herein. [Modes for carrying out the invention]
[0030] This application claims priority and enjoys the benefits thereof based on Korean Patent Application No. 10-2023-0182553, filed with the Korean Intellectual Property Office on December 15, 2023, the entire contents of which are incorporated herein by reference.
[0031] Various embodiments are described with reference to the drawings. Various explanations are provided herein to provide an understanding of the disclosure. Before describing specific details for implementing the disclosure, note that configurations not directly related to the technical essence of the disclosure have been omitted to the extent that they do not obscure the technical essence of the invention. Furthermore, terms or words used herein and in the claims should be interpreted as having meanings and concepts consistent with the technical idea of the invention, based on the principle that inventors may define appropriate terms to best describe their inventions.
[0032] As used herein, terms such as “model,” “system,” and / or “module” refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or software executions, and can be used interchangeably with each other. For example, a module may be, but is not limited to, a process executed on a processor, a processor, an object, an execution thread, a program, an application, and / or a computing device. One or more modules may reside within a processor and / or an execution thread. A module may be localized to one computer. A module may be distributed between two or more computers. Such a module may also be executed from various computer-readable media having various data structures stored within it. A module may communicate via local and / or teleprocessing according to signals having one or more data packets (e.g., data from one component interacting with other components in a local system, a distributed system, and / or data transmitted to other systems via signals over a network such as the Internet).
[0033] Furthermore, the term "or" is intended to mean inclusive "or" rather than exclusive "or". That is, unless otherwise specified or contextually clear, "X uses A or B" is intended to mean one of the natural inclusive substitutions. That is, if X uses A; X uses B; or X uses both A and B, "X uses A or B" can apply to any of these cases. Also, the terms "and / or" and "at least one" as used herein should be understood to refer to and include all possible combinations of one or more of the listed related items. For example, the terms "at least one of A or B" or "at least one of A and B" should be interpreted as meaning "including only A," "including only B," and "a combination of A and B."
[0034] Furthermore, the terms “contains” and / or “includes” should be understood to mean the presence of the feature and / or component in question. However, it should be understood that the terms “contains” and / or “includes” do not exclude the presence or addition of one or more other features, components, and / or groups thereof. Furthermore, where not otherwise specified or where it is not contextually clear that it refers to the singular, in this specification and claims, the singular should generally be interpreted as meaning “one or more.”
[0035] Furthermore, those skilled in the art should recognize that various exemplary logical components described in relation to the embodiments disclosed herein can be implemented in hardware, computer software, or a combination of both.
[0036] The description of the presented embodiments is provided so that a person with ordinary skill in the art of this disclosure may utilize or practice the invention. Various modifications to such embodiments will be obvious to a person with ordinary skill in the art of this disclosure. The general principles defined herein can be applied to other embodiments without departing from the scope of this disclosure. Thus, the invention is not limited by the embodiments presented herein. The invention should be analyzed in the broadest sense, consistent with the principles and novel features presented herein.
[0037] In this disclosure, terms such as the first, second, or third, or any other N, are used to distinguish at least one entity. For example, the entities represented by the first and second may be identical or different from each other.
[0038] The terms expressed in this disclosure as primary, secondary, or tertiary are used to distinguish at least one entity. For example, the terms N, such as primary, secondary, and tertiary, may be used to distinguish temporal order. In such examples, a larger value of N may indicate a later entity in time, and a smaller value of N may indicate an earlier entity in time.
[0039] As used in this disclosure, the term “model” can be used to encompass artificial intelligence-based models, AI models, computational models, neural networks, network functions, and neural networks. In one embodiment, a model may mean a model file, model identification information, the model's execution environment, the model's running time, and / or the model's framework.
[0040] In this disclosure, a Driver Monitoring System (DMS) can represent a software entity or hardware entity, or a combination thereof, to which vehicle technology for monitoring the driver's condition is applied. Such a DMS may be executed by a computing device according to one embodiment of this disclosure.
[0041] As used in this disclosure, the term "image" can be used to encompass one or more frames. For example, one image may correspond to one frame. In one embodiment, the image may be a still image and correspond to a frame obtained from a video.
[0042] In one embodiment, the image may be captured data in which the driver is included as an object, and may be acquired via a camera installed in the vehicle. In one embodiment, the image may be analyzed and / or processed in the DMS to detect an anomaly corresponding to the driver and / or determine whether, type, and / or intensity of an alarm corresponding to the anomaly is generated.
[0043] In this specification, the expression "determine an alarm" may be used to encompass determining whether or not to generate an alarm, determining the type of alarm, and / or determining the intensity of the alarm.
[0044] For the sake of clarity, the technology described herein will be referred to as "images" below. It will be apparent to those skilled in the art that the technology described herein can be implemented using the term "frames."
[0045] As used in this disclosure, the term “anomaly” may be used to describe any action, situation, or element within an image or frame that impedes driver safety. For example, an anomaly may include driver distraction, drowsiness, failure to wear a seatbelt, and / or the presence of flames (or fire).
[0046] As used in this disclosure, the term “voting” may mean an algorithm for correcting predicted anomaly results in order to improve the accuracy of anomaly detection and / or the accuracy of anomaly alarm generation. In one embodiment, voting may mean a rule-based algorithm that combines results from multiple images to determine, correct, and / or adjust anomaly results corresponding to the current image. In one embodiment, voting may mean a method for determining a predicted result from the current image based on predicted results from previous images. In this disclosure, predicted anomaly results and predicted results can be used interchangeably.
[0047] In one embodiment, each of the multiple factors used in the voting may correspond to each of the multiple images acquired over time. In one embodiment, each of the multiple factors used in the voting may correspond to each of the predicted results of the multiple images acquired over time (e.g., predicted results obtained from a model, and / or predicted results that reflect previous voting results). For example, a first predicted result for the first image, a second predicted result for the second image, and a third predicted result for the third image can be considered as factors used in the voting.
[0048] The technology according to one embodiment of the present disclosure allows for the sequential use of multiple voting processes. For example, the results of a first voting process may be used in a second voting process, and the results of a second voting process may be used in a third voting process. By sequentially using multiple voting processes that use predicted results corresponding to previous and current images as factors, more accurate anomaly alarms can be provided.
[0049] Figure 1 schematically shows a block diagram of a computing device 100 according to one embodiment of the present disclosure.
[0050] According to embodiments of this disclosure, the computing device 100 may include a processor 110 and memory 130.
[0051] The configuration of the computing device 100 shown in Figure 1 is merely a simplified example. In one embodiment of this disclosure, the computing device 100 may include other configurations for executing the computing environment of the computing device 100, and only a portion of the disclosed configuration may constitute the computing device 100. For example, if the computing device 100 described above includes a user terminal, an output unit (not shown) and an input unit (not shown) may be included within the scope of the computing device 100.
[0052] In this disclosure, computing device 100 can be used to encompass any form of server and any form of terminal. Computing device 100 may be used interchangeably with computing equipment.
[0053] In this disclosure, computing device 100 may mean any form of component that constitutes a system for realizing embodiments of this disclosure.
[0054] In one embodiment, the computing device 100 may mean a device on which the DMS is driven.
[0055] In one embodiment, the computing device 100 may mean a device for detecting anomalies from the driver's image and / or determining whether or not to generate an alarm corresponding to the anomaly.
[0056] In one embodiment, the computing device 100 may mean a device used to train a model for detecting anomalies from images of the driver.
[0057] In one embodiment, the computing device 100 may mean a device used to train a model for detecting anomalies from images of a driver. In one embodiment, the computing device 100 may mean a device used to infer a model for detecting anomalies from images of a driver.
[0058] In one embodiment, the computing device 100 may mean a server located remotely from the in-vehicle device that acquires the driver's image.
[0059] In one embodiment, the computing device 100 may acquire an image of the driver from a device in the vehicle, detect anomalies in the image, and / or determine an alarm corresponding to the anomaly in the image. For example, it may determine whether or not to generate an alarm corresponding to the anomaly, or the intensity of the alarm.
[0060] In one embodiment, the computing device 100 may obtain object detection results from a target image that includes the driver. For example, the computing device 100 may obtain detection results from the target image that relate to seat belts, eye closure, drowsiness, and / or inattention.
[0061] In one embodiment, the processor 110 may perform the overall operation of the computing device 100. The processor 110 may consist of at least one core. The processor 110 may include devices for data analysis and / or processing, such as a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU) of the computing device 100.
[0062] The processor 110 may read a computer program stored in the memory 130 and, in accordance with one embodiment of the present disclosure, detect an anomaly and / or determine an anomaly alarm.
[0063] In one embodiment of this disclosure, the processor 110 may perform calculations for learning a neural network. The processor 110 may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), feature extraction from the input data, error calculation, and weighting updates of the neural network using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor 110 may process the learning of network functions. For example, the CPU and GPGPU may both process the learning of network functions and data classification using network functions. Furthermore, in one embodiment of this disclosure, the processors of multiple computing devices can be used together to process the learning of network functions and data classification using network functions. Furthermore, the computer program executed on the computing device 100 in one embodiment of this disclosure may be a CPU, GPGPU, or TPU executable program.
[0064] Furthermore, the processor 110 may typically handle the overall operation of the computing device 100. For example, the processor 110 may provide the user with appropriate information or functions by processing data, information, or signals that are input or output through components included in the computing device 100, or by driving application programs stored in memory.
[0065] According to one embodiment of the present disclosure, the memory 130 may store any form of information generated or determined by the processor 110 and any form of information received by the computing device 100. According to one embodiment of the present disclosure, the memory 130 may also be a storage medium for storing computer software that causes the processor 110 to perform the operations according to the embodiments of the present disclosure. Thus, the memory 130 may mean a computer-readable medium for storing software code necessary to perform the embodiments of the present disclosure, data on which the code is executed, and the results of the code execution.
[0066] According to one embodiment of this disclosure, memory 130 may mean any type of storage medium. For example, memory 130 may include at least one type of storage medium from among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, magnetic disk, and optical disk. The computing device 100 may also operate in conjunction with web storage that performs the storage functions of memory 130 over the internet. The above descriptions of memory are illustrative, and memory 130 used in this disclosure is not limited to the above examples.
[0067] The communication unit (not shown) of this disclosure can be configured in any manner, including wired and wireless, and may be composed of various communication networks such as Personal Area Networks (PANs) and Wide Area Networks (WANs). Furthermore, the network unit 150 can operate on a known World Wide Web (WWW) basis and may utilize wireless transmission technologies used for short-range communication, such as infrared (IrDA) or Bluetooth®.
[0068] The computing device 100 in this disclosure may include any form of user terminal and / or any form of server. Therefore, embodiments of this disclosure may be performed by a server and / or user terminal.
[0069] In one embodiment, the user terminal may include any form of terminal capable of interacting with a server or other computing device. The user terminal may include, for example, a mobile phone, a smartphone, a laptop computer, a PDA (personal digital assistant), a slate PC, a tablet PC, and an ultrabook. In one embodiment, the user terminal may mean a device including a camera installed in a vehicle.
[0070] In one embodiment, the server may include any type of computing system or computing device, such as a microprocessor, a mainframe computer, a digital processor, a portable device, and a device controller.
[0071] In one embodiment, the server may include a storage unit (not shown) for storing data and / or information used in the present disclosure. Such a storage unit may be contained within the server or reside under the server's control. In another example, the storage unit may reside outside the server and be implemented in a manner that allows it to communicate with the server. In this case, the storage unit may be managed and controlled by another external server different from the server.
[0072] Figure 2 shows an exemplary structure of an artificial intelligence-based model according to one embodiment of the present disclosure.
[0073] In this disclosure, the terms model, artificial intelligence model, artificial intelligence-based model, computational model, neural network, network function, and neural network may be used interchangeably with each other.
[0074] The artificial intelligence-based models described herein may include models applicable to various domains, such as models for image processing including object segmentation, object detection, anomaly detection, and / or object classification, and models for text processing including data prediction, text semantic inference, and / or data classification.
[0075] A neural network can generally be composed of a set of interconnected computational units called nodes. Such nodes are sometimes called neurons. A neural network consists of at least one node. The nodes (or neurons) that make up a neural network may be interconnected by one or more links.
[0076] Nodes within an artificial intelligence model may be used to represent components that make up a neural network; for example, nodes in a neural network may correspond to neurons.
[0077] Within a neural network, one or more nodes connected by links can form a relative input-output node relationship. The concepts of input and output nodes are relative; any node that is an output node to another node is an input node to another node, and vice versa. As mentioned above, the relationship between input and output nodes may be generated around links. One input node may be connected to one or more output nodes via links, and vice versa.
[0078] In a relationship between input and output nodes connected via a single link, the data of the output node may be determined based on the data input to the input node. Here, the link connecting the input and output nodes may have weights. The weights may be variable and can be varied by the user or algorithm in order for the neural network to perform a desired function. For example, if one or more input nodes are interconnected to one output node by their respective links, the output node may determine its output node value based on the values input to the input nodes connected to the output node and the weights set for the links corresponding to each input node.
[0079] As mentioned above, a neural network consists of one or more nodes interconnected via one or more links, forming input-output node relationships within the network. The characteristics of a neural network may be determined by the number of nodes and links within the network, the relationships between nodes and links, and the weighting values assigned to each link. For example, if there are two neural networks with the same number of nodes and links but different link weighting values, the two neural networks can be recognized as distinct from each other.
[0080] A neural network may consist of a set of one or more nodes. A subset of nodes constituting a neural network can constitute a layer. A portion of the nodes constituting a neural network can constitute a single layer based on their distance from the initial input node. For example, a set of nodes that are n in distance from the initial input node can constitute an n-layer. The distance from the initial input node can be defined by the minimum number of links that must be traversed to reach the node in question from the initial input node. However, such a definition of a layer is arbitrary for illustrative purposes, and the order of layers within a neural network may be defined in ways different from those described above. For example, the layer of nodes may be defined by their distance from the final output node.
[0081] In embodiments of this disclosure, a collection of neurons or nodes may be defined as a “layer.”
[0082] The initial input node may mean one or more nodes in the neural network that receive data directly without links in relation to other nodes. Alternatively, it may mean a node in the neural network that does not have other input nodes connected by links in relation to other nodes based on links. Similarly, the final output node may mean one or more nodes in the neural network that do not have an output node in relation to other nodes. Furthermore, a hidden node may mean a node in the neural network that is neither the initial input node nor the final output node.
[0083] A neural network according to one embodiment of the present disclosure may have the same number of nodes in the input layer as the number of nodes in the output layer, and the number of nodes may decrease as you move from the input layer to the hidden layer, and then increase again. Another neural network according to another embodiment of the present disclosure may have fewer nodes in the input layer than the number of nodes in the output layer, and the number of nodes may decrease as you move from the input layer to the hidden layer. Yet another neural network according to yet another embodiment of the present disclosure may have more nodes in the input layer than the number of nodes in the output layer, and the number of nodes may increase as you move from the input layer to the hidden layer. Another neural network according to another embodiment of the present disclosure may be a neural network that is a combination of the neural networks described above.
[0084] A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to input and output layers. Deep neural networks can be used to understand the latent structures of data. For example, the latent structures of photographs, text, videos, audio, protein sequence structures, gene sequence structures, peptide sequence structures, and / or music may be understood via a deep neural network. Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, Generative Adversarial Networks (GANs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siam networks, and Generative Adversarial Networks (GANs). The above descriptions of deep neural networks are illustrative and this disclosure is not limited thereto.
[0085] The artificial intelligence models described herein can be represented by a network structure of any of the aforementioned structures, including an input layer, a hidden layer, and an output layer.
[0086] The neural networks that can be used in the clustering models of this disclosure may be trained in at least one of the following ways: supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Training of a neural network may be the process of applying knowledge to the neural network so that it can perform a particular action.
[0087] Neural networks can be trained to minimize the error in their output. Training a neural network involves repeatedly inputting training data, calculating the neural network's output and target error for the training data, and updating the weights of each node in the neural network by backpropagating the error from the output layer to the input layer in a way that reduces the error. In guided learning, training data with the correct answer labeled is used (i.e., labeled training data), while in unguided learning, the training data may not be labeled. For example, in guided learning for data classification, the training data may be data with a category labeled for each data point. The labeled training data may be input to the neural network, and the error may be calculated by comparing the neural network's output (category) with the labels on the training data. As another example, in unguided learning for data classification, the error may be calculated by comparing the input training data with the output of the neural network. The calculated errors are backpropagated in the reverse direction of the neural network (i.e., from the output layer to the input layer), and this backpropagation can update the connection weights of each node in each layer of the neural network. The amount of change in the connection weights of each node being updated may be determined by the learning rate. The computation of the neural network on the input data and the backpropagation of errors can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of iterations of the neural network's learning cycle. For example, a high learning rate can be used in the early stages of learning to increase efficiency by allowing the neural network to quickly achieve a certain level of performance, while a lower learning rate can be used in the later stages of learning to improve accuracy.
[0088] In neural network training, training data is generally a subset of real-world data (i.e., data that the trained neural network intends to process). Therefore, there can be training cycles where errors on training data decrease, but errors on real-world data increase. Overfitting is this phenomenon where the network over-trains on training data, leading to increased errors on real-world data. For example, a neural network trained to recognize cats by being shown yellow cats may fail to recognize cats that are not yellow; this is a type of overfitting. Overfitting can act as a cause of increased errors in machine learning algorithms. Various optimization methods can be used to prevent such overfitting. To prevent overfitting, methods such as increasing the amount of training data, regularization, dropout (deactivating some of the network nodes during the training process), and the use of a batch normalization layer can be applied.
[0089] One embodiment of the present disclosure discloses a computer-readable medium storing a data structure including an artificial intelligence-based model. The aforementioned data structure may be stored in a storage unit (not shown) of the present disclosure, executed by a processor 110, and transmitted and received by a communication unit (not shown).
[0090] A data structure may mean the organization, management, and storage of data that enables efficient access to and modification of data. A data structure may also mean the organization of data to solve a specific problem (e.g., data retrieval, data storage, data modification in the shortest time). A data structure can also be defined as physical or logical relationships between data elements designed to support specific data processing functions. Logical relationships between data elements may include user-defined linking relationships between data elements. Physical relationships between data elements may include actual relationships between data elements physically stored in a computer-readable storage medium (e.g., persistent storage). Specifically, a data structure may include a collection of data, relationships between data, and functions or instructions that can be applied to data. A well-designed data structure may enable a computing device to perform operations with minimal use of its resources. Specifically, a well-designed data structure can improve the efficiency of operations such as arithmetic, reading, insertion, deletion, comparison, exchange, and retrieval.
[0091] Data structures can be divided into linear and non-linear data structures depending on their form. A linear data structure is one in which only one piece of data is linked after another. Linear data structures may include lists, stacks, queues, and decks. A list may refer to a set of data that has an internal order. A list may also include linked lists. A linked list may be a data structure in which data is linked in a linear fashion, with each piece of data having a pointer. In a linked list, the pointer may contain linking information to the next or previous piece of data. Linked lists can be expressed as single linked lists, double linked lists, or circular linked lists depending on their form. A stack is a data arrangement structure in which data can be accessed in a restricted manner. A stack may be a linear data structure in which data can only be processed (e.g., inserted or deleted) at one end of the data structure. Data stored in a stack may be a LIFO (Last In First Out) data structure, where data entered later comes out earlier. A queue is a data arrangement structure that restricts access to data, and unlike a stack, it may be a data structure where data stored later is retrieved later (FIFO - First in First Out). A deck may be a data structure that allows data to be processed from both ends of the data structure.
[0092] A nonlinear data structure is a structure in which multiple data are linked after a single data item. Nonlinear data structures may include graph data structures. Graph data structures can be defined by vertices and edges, and edges may include lines connecting two different vertices. Graph data structures may also include tree data structures. A tree data structure may be a data structure in which a path connecting two different vertices among the multiple vertices included in the tree is a single data structure. In other words, a graph data structure may not form a loop.
[0093] The data structure may include a neural network. Furthermore, the data structure including a neural network may be stored on a computer-readable medium. The data structure including a neural network may also include pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. The data structure including a neural network may include any of the components of the disclosed configuration. That is, the data structure including a neural network may consist of all or any combination thereof of pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. In addition to the configurations described above, the data structure including a neural network may include any other information that determines the properties of the neural network. Furthermore, the data structure may include, but is not limited to, any form of data used or generated during the computational process of the neural network. Computer-readable media may include computer-readable recording media and / or computer-readable transmission media. A neural network can generally consist of a collection of interconnected computational units called nodes. Such nodes are sometimes called neurons. A neural network consists of at least one node.
[0094] The data structure may include data to be input to a neural network. The data structure including data to be input to a neural network may be stored on a computer-readable medium. The data to be input to a neural network may include training data input during the neural network's learning process and / or input data to be input to a neural network after training is complete. The data to be input to a neural network may include pre-processed data and / or data subject to pre-processing. Pre-processing may include data processing processes for inputting data to a neural network. Therefore, the data structure may include data subject to pre-processing and data generated during pre-processing. The data structures described above are illustrative and the disclosure is not limited thereto.
[0095] The data structure may include weights for the neural network (in this specification, weights and parameters may be used interchangeably). The data structure including the weights for the neural network can be stored on a computer-readable medium. The neural network may include multiple weights. The weights are variable and may be varied by the user or algorithm in order for the neural network to perform a desired function. For example, if one or more input nodes are interconnected to an output node by their respective links, the output node may determine the data values output from the output node based on the values input to the input nodes connected to the output node and the weights set for the links corresponding to each input node. The data structures described above are illustrative and the disclosure is not limited thereto.
[0096] As a non-limiting example, the weights may include weights that change during the neural network learning process and / or weights after the neural network has finished learning. The weights that change during the neural network learning process may include weights at the start of a learning cycle and / or weights that change during a learning cycle. The weights after the neural network has finished learning may include weights after the learning cycle has finished. Therefore, a data structure containing neural network weights may include a data structure containing weights that change during the neural network learning process and / or weights after the neural network has finished learning. Accordingly, the weights and / or each combination of weights described above shall be included in the data structure containing neural network weights. The data structures described above are illustrative and the disclosure is not limited thereto.
[0097] A data structure containing neural network weights can be stored on a computer-readable storage medium (e.g., memory, hard disk) after undergoing a serialization process. Serialization may be a process of converting the data structure into a format that can be stored on the same or different computing devices and later reconfigured for use. Computing devices can serialize the data structure to send and receive data over a network. A serialized data structure containing neural network weights may be reconfigured on the same or other computing devices through deserialization. A data structure containing neural network weights is not limited to serialization. Furthermore, a data structure containing neural network weights may include data structures that enhance computational efficiency while minimizing the use of computing device resources (e.g., in nonlinear data structures, B-trees, R-trees, tries, m-way search trees, AVL trees, Red-Black Trees). The foregoing is illustrative, and this disclosure is not limited thereto.
[0098] The data structure may include the hyperparameters of the neural network. The data structure containing the neural network hyperparameters can be stored on a computer-readable medium. The hyperparameters may be variable variables controlled by the user. Examples of hyperparameters may include the learning rate, cost function, number of iterations in the learning cycle, weight initialization (e.g., setting the range of weights to be initialized), and number of Hidden Units (e.g., number of hidden layers, number of nodes in hidden layers). The aforementioned data structure is illustrative and the disclosure is not limited thereto.
[0099] Figure 3 illustrates an example of a method for detecting an anomaly and determining an anomaly alarm in a DMS according to one embodiment of the present disclosure.
[0100] As shown in Figure 3, the computing device 100 may detect the driver's face (310).
[0101] In one embodiment, the computing device 100 may acquire images from a camera installed inside the vehicle. For example, the images may include the driver inside the vehicle.
[0102] In one embodiment, the computing device 100 may use a face detection model to detect the driver's face from the acquired image. For example, the model may include an object detection model and / or an object segmentation model. For example, the model may correspond to a pre-trained artificial intelligence-based model to detect and / or segment human faces in the image.
[0103] In one embodiment, the model for face detection may correspond to the detection model.
[0104] In one embodiment, the face detection model may output the result of segmenting contours that define faces in an image. In another embodiment, the face detection model may output bounding boxes containing faces from an image. In yet another embodiment, the face detection model may be configured to output together a region in an image corresponding to a face and a plurality of feature points that can identify a face within the region. In yet another embodiment, the face detection model may be configured to output together a region in an image corresponding to a face, a plurality of feature points that can identify a face within the region, and a plurality of feature points that can identify eyes within the face.
[0105] In one embodiment, the computing device 100 may detect a face landmark from the detected face (320).
[0106] In one embodiment, the computing device 100 may use a model for facial landmark detection to acquire feature points contained in the driver's face on the driver's face. In this disclosure, feature points and landmarks can be used interchangeably.
[0107] In one embodiment, the model for facial landmark detection may correspond to a pre-trained artificial intelligence-based model that determines a plurality of feature points for identifying facial features (e.g., eye features, nasal features, and / or mouth features) on the driver's face. In another embodiment, the model for facial landmark detection may be configured to output together a plurality of feature points that can identify a face within a facial region, and a plurality of feature points that can identify eyes within the face.
[0108] In one embodiment, the model for detecting facial landmarks may correspond to a detection model.
[0109] In one embodiment, the driver's identity can be identified based on the detection of the driver's facial landmark. Identification of the driver's identity may be performed based on a comparison between a pre-stored driver's facial landmark and the detected driver's facial landmark.
[0110] In one embodiment, the computing device 100 may detect eye landmarks from the image (330).
[0111] For example, eye landmark detection may be performed more efficiently using the results of face detection. For example, eye landmark detection may be included in the results of face detection. Based on eye landmark detection, the computing device 100 can determine whether the driver is drowsy and / or distracted. For example, eye landmark detection may detect the position of the eyes within the face, whether the eyes are closed, and / or the direction the eyes are looking.
[0112] For example, a model for detecting eye landmarks may correspond to a detection model.
[0113] In one embodiment, the computing device 100 may detect anomalies from the image (340).
[0114] In one embodiment, the computing device 100 may use the results of eye landmark detection to detect anomalies from the image. For example, the computing device 100 may use the results of eye landmark detection to determine whether the driver is distracted while driving (e.g., not paying attention to the road ahead). For example, the computing device 100 may use the results of eye landmark detection to determine whether the driver is drowsy while driving.
[0115] In this disclosure, "abnormality" may be used to describe abnormal behavior while driving. In this disclosure, "abnormality" may be used to describe situations and / or behaviors that impede safety while driving. For example, an abnormality may include a first abnormality corresponding to not wearing a seat belt, a second abnormality corresponding to distraction in the driving situation, a third abnormality corresponding to drowsiness in the driving situation, a fourth abnormality corresponding to the driver smoking, a fifth abnormality corresponding to a fire in the vehicle, and / or a sixth abnormality corresponding to closing one's eyes while driving.
[0116] In other embodiments, the computing device 100 may use one model to detect multiple anomalies. In other embodiments, the computing device 100 may operate to detect multiple anomalies using multiple models (for example, models dedicated to specific anomaly detection).
[0117] In other embodiments, the computing device 100 may detect the driver's body from the image. For example, after detecting the driver's face, the driver's body in the image may be detected based on the driver's face. In another example, driver body detection may be performed independently of driver face detection. Based on such body detection, it may be determined whether the driver is wearing a seat belt, whether the driver is smoking, and / or where the driver's hands are. Based on the driver's body detection, abnormalities related to the driver's body (e.g., not wearing a seat belt, smoking, and / or a fire inside the vehicle) may be detected.
[0118] For example, a model for detecting a body may correspond to a detection model.
[0119] In other embodiments, the computing device 100 may use one model to detect multiple anomalies. In other embodiments, the computing device 100 may operate to detect multiple anomalies using multiple models (for example, models dedicated to specific anomaly detection).
[0120] In one embodiment, the computing device 100 may use a classification model that utilizes eye detection results and / or face detection results to determine whether there are any anomalies (e.g., attention distractions) corresponding to the input image. As an unrestricted example, such a classification model may operate to divide the input image into forward or non-forward sections. As an unrestricted example, such a classification model may operate to output quantitative values indicating whether the input image is forward or non-forward.
[0121] In one embodiment, the computing device 100 may use the anomaly detection result to determine an alarm corresponding to the anomaly (350).
[0122] In one embodiment, an alarm or abnormal alarm corresponding to an anomaly may include various forms of output to cause the user to recognize the anomaly, such as sound, images, vibrations and / or light.
[0123] For example, if computing device 100 determines that an anomaly exists in an image using one or more models, it may decide whether or not to generate an alarm corresponding to the anomaly. If anomaly detection directly leads to an anomaly alarm, there is a problem that the driver may receive unnecessary or inaccurate alarms while driving. Therefore, the technology according to one embodiment of the present disclosure can provide the user with more optimal and accurate alarms by deciding whether or not to generate an anomaly alarm using the anomaly detection result. For example, the presence of an anomaly may be determined to be a situation in which driver distraction is detected, a situation in which driver drowsiness is detected, and / or a situation in which the driver is not wearing a seat belt.
[0124] For example, if computing device 100 determines that an anomaly exists in the image, it may determine the type of alarm corresponding to the anomaly and / or the intensity of the alarm. Computing device 100 may also use one or more models to determine the intensity of an anomaly alarm in the image. For example, computing device 100 may determine multiple anomaly alarms or the intensity of an anomaly alarm by comparing the expected result related to the anomaly with each of several thresholds. For example, computing device 100 may determine multiple anomaly alarms or the intensity of an anomaly alarm by applying one or more counter concepts to the expected result related to the anomaly and comparing each of the counters with a threshold. Thus, the technology according to one embodiment of the present disclosure can provide the user with more optimal and accurate alarms by using the anomaly detection result to determine whether or not to generate an anomaly alarm and / or the intensity of the anomaly alarm. As an example, the presence of an anomaly may be determined in situations where driver distraction is detected, driver drowsiness is detected, and / or the driver is not wearing a seat belt. In this way, by adjusting the type and / or intensity of alarms corresponding to anomalies, more intuitive and clear alarms can be communicated to the user, thereby maximizing the utilization of the DMS.
[0125] Figure 4 illustrates a method for determining whether or not to generate an abnormal alarm in a DMS according to one embodiment of the present disclosure.
[0126] In one embodiment, the computing device 100 may acquire an image including the driver inside the vehicle (410).
[0127] In one embodiment, the image (or first image) can represent a target image for determining an abnormal alarm, as an image that includes the driver.
[0128] In one embodiment, the image means an image acquired from a camera. In one embodiment, the image may mean an image of the driver taken by a camera installed inside the vehicle. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to not wearing a seat belt. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to distraction. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to closing the eyes. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to drowsiness.
[0129] In this disclosure, the term "image" may be used to encompass one or more frames. For example, one image may correspond to one frame. For example, an image may be a still image and may correspond to a frame obtained from a video. For example, an image may include a collection of frames taken multiple times.
[0130] In one embodiment, the image may correspond to an image or frame taken at a specific point in time. In one embodiment, the image may mean a still image or frame extracted or captured from a video of a driver taken by a camera.
[0131] In one embodiment, the computing device 100 may acquire multiple images or videos acquired by a camera. The computing device 100 may acquire or extract specific images (e.g., specific frames) from the acquired images or videos in order to determine an anomaly or to trigger an anomaly alarm. For example, the images subject to an anomaly or anomaly alarm determination may be selected or determined frames from among multiple frames. In such an example, the computing device 100 may extract specific frames randomly from among multiple frames, or in units of a predetermined time period.
[0132] In one embodiment, the computing device 100 may use an artificial intelligence-based model to obtain model output information from an image (420).
[0133] In one embodiment, the model may correspond to a deep learning-based model that is pre-trained to take an image of a driver as input and output the probability of a predetermined object being present in the image of the driver. In such an embodiment, the model can output the probability of a seat belt being present in the image and / or the possibility of wearing a seat belt as model output information. In such an embodiment, the model can output the probability of eyes being closed and / or drowsiness in the image as model output information. In such an embodiment, the model can output the distance between the top and bottom of the eyes in the image as model output information. In such an embodiment, the model can output the probability of inattention to the road ahead and / or the driver's gaze information in the image as model output information.
[0134] In other embodiments, the model may correspond to a pre-trained deep learning-based model that takes an image of a driver as input and outputs the possibility that a predetermined anomaly that impedes safety is present in the image of the driver. The model may be pre-trained on a training dataset in which the presence or absence of anomalies in the images is labeled. The model can be pre-trained on a training dataset in which the presence or absence of anomalies in the images is labeled. The model can be pre-trained on a training dataset in which the location and / or type of anomaly in the images is labeled.
[0135] In one embodiment, the model output information may include a value that quantitatively indicates the likelihood of a predetermined object being present in the image. For example, the model output information may include a seat belt detection score in the image. For example, the model output information may include quantitative information related to seat belt detection in the image. For example, the model output information may include an eye-closing score in the image. For example, the model output information may include a score related to drowsiness in the image. For example, the model output information may include a score related to the distance between the upper and lower parts of the eyes in the image. For example, the model output information may include a score related to forward non-gaze or distraction in the image.
[0136] In one embodiment, the first model output information may include a value that quantitatively indicates the possibility of an abnormality in the first image. In one embodiment, the first model output information may include a value that indicates whether or not an abnormality exists in the first image. For example, the first model output information may include quantitative information related to the failure to detect a seat belt in the image. For example, the first model output information may include quantitative information related to the driver's eyes being closed in the image. For example, the first model output information may include quantitative information related to the driver's drowsiness in the image. For example, the first model output information may include quantitative information related to the driver's forward gaze or lack thereof in the image, or quantitative information related to the driver's distraction.
[0137] In one embodiment, a first image acquired from a camera may be input to an artificial intelligence-based model to obtain first model output information indicating the possibility of a predetermined object or predetermined action being present in the first image.
[0138] In one embodiment, model output information may mean the output of the model. For example, the model may generate model output information that quantitatively indicates the possibility of an anomaly (e.g., not wearing a seat belt, distraction, closed eyes, and / or drowsiness) in the first image. For example, the model may generate model output information that indicates whether or not an anomaly is present in the first image. For example, the model may generate model output information that quantitatively represents the possibility of a predetermined object or predetermined action in the first image, such as a seat belt, closed eyes, drowsiness, smoking, looking forward, and / or not looking forward.
[0139] In one embodiment, the model output information may mean the result of applying post-processing to the model output. In one embodiment, the model output information may include the result of processing the model output. For example, the model generates bounding boxes corresponding to anomalies or specific objects and / or outputs indicating the possibility of anomalies in the bounding boxes or the possibility of corresponding to specific objects, and the computing device 100 may, by post-processing the model output, generate model output information indicating the possibility of anomalies or the presence or absence of anomalies, or the detectability of specific objects. For example, the model outputs results related to the detection of seat belts in an image, and the computing device 100 may generate results from the results regarding whether the driver is wearing a seat belt or not. For example, the model outputs results corresponding to faces, and the computing device 100 may calculate yaw and / or pitch values from the results corresponding to faces, and generate model output information indicating the presence or absence of attention distraction or the possibility of attention distraction based on the calculated yaw and pitch values. For example, if the model outputs the result of detecting an object in the first space, the computing device 100 may generate model output information indicating the presence or absence of drowsiness or the possibility of drowsiness by generating a conversion result that converts the result to the converted space.
[0140] In one embodiment, the model output information may be determined for each image (for example, each frame).
[0141] In one embodiment, the computing device 100 may compare the first model output information with a first threshold to obtain a first anomaly prediction result indicating whether or not an anomaly exists in the first image (430).
[0142] In one embodiment, the computing device 100 can compare first model output information with a predetermined threshold for anomaly detection. In one embodiment, the first model output information may include a quantitative value for comparison with the threshold. In one embodiment, the threshold may represent a variable threshold that can be changed for each of the images being compared. In one embodiment, the threshold may be dynamically changed based on other information.
[0143] In one embodiment, the computing device 100 may obtain a first anomaly prediction result indicating whether or not an anomaly exists in the first image, based on the results of the comparison. For example, the first anomaly prediction result may be obtained by comparing a quantitative value obtained by processing the output of the model or the output of the model with a threshold. The first anomaly prediction result may have a value indicating whether or not an anomaly exists in the image in question. For example, the first anomaly prediction result may have a value of 1 if an anomaly exists and a value of 0 if an anomaly does not exist, and the opposite situation is also possible depending on the implementation. In such an example, if the first model output information has a value of 0.7 and the threshold has a value of 0.6, the first anomaly prediction result may be set to have a value of 1, indicating that an anomaly exists. If the first model output information has a value of 0.5 and the threshold has a value of 0.6, the first anomaly prediction result may be set to have a value of 0, indicating that an anomaly does not exist. In one embodiment, the first anomaly prediction result may be determined for each image (for example, each frame). For example, the primary prediction result for the first abnormality may be set to have a value of 1 if the seat belt is not detected, and a value of 0 if the seat belt is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if eye closure is not detected, and a value of 1 if eye closure is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if drowsiness is not detected, and a value of 1 if drowsiness is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if forward inattention is not detected, and a value of 1 if forward inattention is detected. In such examples, the first model output information may be set to have a higher value the more likely it is that the seat belt will not be detected. In such examples, the first model output information may be set to have a higher value the more likely it is that eye closure, drowsiness, and / or forward inattention will be detected.
[0144] The technology according to one embodiment of the present disclosure can ensure the accuracy and reliability of the model output by dynamically adjusting the thresholds compared with the model output information in various ways.
[0145] In one embodiment, the first threshold may be determined based on at least one primary prediction result of a previous anomaly corresponding to at least one previous image acquired before the first image. In one embodiment, the first threshold compared with the first model output information may be determined based on second model output information generated by the model in response to a second image acquired before the first image. In one embodiment, the first threshold compared with the first model output information may be modified based on the threshold for anomaly determination of the second image acquired immediately before the first image, based on the primary prediction result of a previous anomaly or the previous model output information. The current threshold corresponding to the current image may be variable based on a comparison between the previous image and the previous threshold. For example, if the model output information of the previous image exceeds the previous threshold, the current threshold corresponding to the current image may be determined to a lower threshold among several threshold options or decreased from the previous threshold. For example, if the model output information of the previous image does not exceed the previous threshold, the current threshold corresponding to the current image may be determined to a higher threshold among several threshold options or increased from the previous threshold.
[0146] In other embodiments, the first threshold, which is compared with the first model output information, may be determined based on a comparison between a primary prediction result of an anomaly obtained from a previous image and a specific threshold. The specific threshold can be considered as a threshold for determining other thresholds. In such embodiments, in determining the current threshold for the current image, the computing device 100 may compare the ratio value of 1 with the specific threshold and the ratio value of 0 with the specific threshold for a plurality of previous primary prediction results of anomalies (e.g., having values of 0 or 1) corresponding to a plurality of previous images, to determine which primary prediction result exceeds the specific threshold. If the primary prediction result exceeding the specific threshold is 1, the computing device 100 may decide to set the threshold to a lower value among a plurality of threshold options, or to decrease the threshold relative to a previous threshold. If the primary prediction result exceeding the specific threshold is 0, the computing device 100 may decide to set the threshold to a higher value among a plurality of threshold options, or to increase the threshold relative to a previous threshold. If the primary prediction result exceeding the specific threshold is neither 0 nor 1, the computing device 100 may decide to maintain the threshold.
[0147] In one embodiment, the threshold change range may be predetermined. The threshold can be sequentially increased, sequentially decreased, or maintained in a counter-like manner based on the predicted results of previous images. The increase or decrease of the threshold can be performed within the threshold change range.
[0148] In other embodiments, the first threshold, which is compared with the first model output information, may be determined based on the secondary prediction result of anomalies in the previous image. In such embodiments, if the secondary prediction result of anomalies in the previous image is 1, the threshold for the current image may be decreased compared to the previous threshold, or set to a lower value among the multiple threshold options. If the secondary prediction result of anomalies in the previous image is 0, the threshold for the current image may be increased compared to the previous threshold, or set to a higher value among the multiple threshold options.
[0149] In one embodiment, the expression "determined based on specific information" may include the fact that the threshold can be quantitatively modified in magnitude depending on the magnitude and / or type of value of the specific information. The computing device 100 may acquire multiple images over time. A primary prediction result of an anomaly may be acquired for each of these multiple images. For example, a primary prediction result of a second anomaly corresponding to a previously acquired second image may be acquired for the first image, and a primary prediction result of a first anomaly corresponding to the first image may be acquired. In such an example, the first threshold used to acquire the primary prediction result of the first anomaly corresponding to the first image may be determined based on the primary prediction result of the second anomaly corresponding to the second image. For example, the first threshold may be changed in accordance with the primary prediction result of the second anomaly. For example, the first threshold may be determined to a preset value such as 0.3 if the primary prediction result of the second anomaly is 1, and 0.5 if the primary prediction result of the second anomaly is 0.
[0150] In one embodiment, the variable reference value here may be a threshold corresponding to the second image acquired immediately before the first image. For example, if the primary prediction result for the second anomaly includes the result that an anomaly exists, the first threshold corresponding to the first image following the second image may be set to increase compared to past thresholds (for example, the threshold corresponding to the second image). When the magnitude of the threshold increases, it may be determined that an anomaly exists in the first image if the quantitative value of the model output information corresponding to the image is relatively high. For example, if the primary prediction result for the second anomaly includes the result that an anomaly does not exist, the first threshold corresponding to the first image following the second image may be set to decrease compared to past thresholds. When the magnitude of the threshold decreases, it may be determined that an anomaly exists in the first image even if the quantitative value of the model output information corresponding to the image is relatively low.
[0151] In other embodiments, the first threshold, which is compared with the first model output information, may be modified based on the second model output information corresponding to a previously acquired second image of the first image. For example, the first threshold may have a negative correlation with the value of the model output information of the previous image. In such an example, if the model output information of the previous image is relatively large, the magnitude of the first threshold may be set to be small. If the model output information of the previous image is relatively small, the magnitude of the first threshold may be set to be large.
[0152] In one embodiment, the first threshold may be determined based on a primary prediction result of a second anomaly obtained by comparing a second threshold, which is determined based on a primary prediction result of a third anomaly corresponding to a previously acquired third image of the second image, with second model output information. Here, the second threshold may be a threshold for determining whether or not an anomaly exists in the second image.
[0153] Thus, the technology according to one embodiment of the present disclosure can dynamically adjust the threshold for anomaly prediction or judgment of the current image for each of a plurality of sequential images by utilizing anomaly-related results corresponding to previous images. This adjustment of the threshold can increase the accuracy and / or confidence of the model's output.
[0154] In one embodiment, the first threshold may be determined based on the ratio of result values indicating the presence of an anomaly from previous primary prediction results of anomalies corresponding to a predetermined first number of previously acquired images of the first image. For example, if the ratio of result values indicating the presence of an anomaly in the previous primary prediction results of anomalies (i.e., multiple primary prediction results of anomalies) is greater than or equal to a first ratio, the first threshold may be set to a first value; if the ratio of result values indicating the presence of an anomaly from the previous primary prediction results of anomalies is less than the first ratio, the first threshold may be set to a second value higher than the first value.
[0155] In one embodiment, the first threshold may be determined or modified depending on what the majority value is among the result values indicating the presence or absence of an anomaly from the previous primary prediction results of anomalies corresponding to a predetermined first number of previously acquired images of the first image. For example, if the majority value of the previous primary prediction results of anomalies is a value indicating the presence of an anomaly, the first threshold corresponding to the current image may be set to decrease. For example, if the mainstream value of the previous primary prediction results of anomalies is a value indicating the absence of an anomaly, the first threshold corresponding to the current image may be set to increase.
[0156] In one embodiment, the model output information corresponding to each of the sequentially acquired images may be structured in the form of a queue. For example, one queue may consist of a predetermined number of units. For example, a first unit in one queue may be assigned first model output information corresponding to the first image, a second unit may be assigned second model output information corresponding to the second image, and a third unit may be assigned third model output information corresponding to the third image.
[0157] In one embodiment, the primary prediction results for anomalies corresponding to each of a plurality of sequentially acquired images may be structured in the form of a queue. Each value corresponding to a single image may be assigned to the same position in the plurality of queues. The queue composed of primary prediction results for anomalies and the queue composed of model output information may have corresponding positions for a single image. For example, the first image acquired at a first time point may be assigned to the same position in the first queue composed of model output information and the second queue composed of primary prediction results for anomalies.
[0158] In one embodiment, a queue may consist of a predetermined number of units. For example, a predetermined number of units in a queue may contain data corresponding to images acquired over time. For example, a first unit in a queue may be assigned the first predicted result of an anomaly corresponding to a first image, a second unit the first predicted result of an anomaly corresponding to a second image, and a third unit the first predicted result of an anomaly corresponding to a third image. In such an example, the threshold to be compared with the model output result corresponding to the currently acquired image may be determined based on the ratio of result values indicating the presence of anomalies among the first, second, and third predicted results of anomalies contained in a queue. For example, if both the first and second predicted results of anomalies include result values indicating the presence of anomalies, and the third predicted result includes result values indicating the absence of anomalies, the ratio of result values may be (2 / 3 × 100). By comparing such a ratio of result values with a specific threshold, the threshold corresponding to the current image can be determined or modified. The ratio indicating the presence of anomalies in previous anomaly prediction results and the magnitude of the threshold corresponding to the current image may have a negative correlation. In this disclosure, the expression that the first value and the second value have a negative correlation means that as the first value increases, the second value tends to decrease. In this disclosure, the expression that the first value and the second value have a negative correlation may also mean that as the first value becomes relatively larger, the second value becomes relatively smaller.
[0159] In one embodiment, the computing device 100 may obtain a secondary prediction result for the first anomaly by performing a first voting using the primary prediction result for the first anomaly (440).
[0160] In one embodiment, the secondary prediction result of the first anomaly may represent a quantitative value used as a parameter for determining an anomaly alarm corresponding to the first image. In another embodiment, the secondary prediction result of the first anomaly may have a value indicating the presence or absence of an anomaly, and is used as a parameter for determining whether or not to generate an anomaly alarm corresponding to the first image.
[0161] In another embodiment, the first threshold used for anomaly prediction of the first image may be determined based on a secondary prediction result of a second anomaly corresponding to a previously acquired second image of the first image.
[0162] In one embodiment, the first voting may be used to correct the accuracy and / or reliability of the model's output. In one embodiment, the first voting may utilize the first predicted results of anomalies in previously acquired images of the first image and the first predicted results of anomalies in the currently acquired first image.
[0163] In one embodiment, the first voting may include a process of determining the majority value of the primary predicted results of anomalies corresponding to sequential images. In one embodiment, the first voting may include a process of determining the ratio of primary predicted results of anomalies corresponding to sequential images. In one embodiment, the first voting may include a process of comparing the ratio of primary predicted results of anomalies corresponding to sequential images with a predetermined threshold.
[0164] In one embodiment, the first voting may utilize values on a voting queue composed of multiple units corresponding to sequential images. These values are primary prediction results for anomalies corresponding to sequential images. For example, the first unit in the voting queue may be assigned the primary prediction result for the first anomaly corresponding to the first image, the second unit may be assigned the primary prediction result for the second anomaly corresponding to the second image, and the third unit may be assigned the primary prediction result for the third anomaly corresponding to the third image. Here, the first image may correspond to the most recently acquired image or the current image, the second image may be an image acquired prior to the first image, and the third image may be an image acquired prior to the second image.
[0165] In one embodiment, there may be multiple boating queues. For example, a first boating queue may include model output information acquired over time, a second boating queue may include primary prediction results acquired over time, and a third boating queue may include secondary prediction results acquired over time.
[0166] In one embodiment, the secondary prediction result for the first anomaly corresponding to the first image may be determined based on what the mainstream values of the primary prediction result for the first anomaly, the primary prediction result for the second anomaly, and the primary prediction result for the third anomaly are.
[0167] In one embodiment, the secondary prediction result for the first anomaly corresponding to the first image may be determined based on a comparison between the ratio of result values representing anomalies among the primary prediction results for the first anomaly, the second anomaly, and the third anomaly, and a specific threshold. For example, if the primary prediction result for the first anomaly indicates the presence of an anomaly, the primary prediction result for the second anomaly indicates the presence of an anomaly, and the primary prediction result for the third anomaly indicates the absence of an anomaly, the secondary prediction result for the first anomaly corresponding to the first image may be set to the mainstream value of the three primary prediction results for anomalies or their representative value indicating the presence of an anomaly (e.g., 1). This allows the first unit corresponding to the first image in the queue composed of secondary prediction results for anomalies to have a value of 1. For example, if the primary prediction result for the first anomaly indicates the absence of an anomaly, the primary prediction result for the second anomaly indicates the presence of an anomaly, and the primary prediction result for the third anomaly indicates the absence of an anomaly, the secondary prediction result for the first anomaly corresponding to the first image may be set to the mainstream value of the three primary prediction results for anomalies or their representative value indicating the absence of an anomaly (e.g., 0). As a result, the first unit corresponding to the first image in the queue, which is composed of secondary prediction results for anomalies, can have a value of 0.
[0168] In one embodiment, the secondary anomaly prediction results corresponding to each of a plurality of sequentially acquired images may be structured in the form of a voting queue. Each value corresponding to a single image may be assigned to the same position in the plurality of queues. The queue composed of secondary anomaly prediction results, the queue composed of primary anomaly prediction results, and the queue composed of model output information may have positions corresponding to each other for a single image. For example, the values related to the anomaly prediction of a first image acquired at a first time point may be assigned to the same position (e.g., corresponding positions) in the queue composed of model output information, the queue composed of primary anomaly prediction results, and the queue composed of secondary anomaly prediction results, respectively.
[0169] In one embodiment of the present disclosure, the first voting may generate a predicted result of a set of anomalies that represents an image group consisting of a first image and a predetermined second number of previously acquired images of the first image, in order to ensure the accuracy of the first model output information. Such a predicted result of a set of anomalies may represent a result that represents the primary predicted result of anomalies corresponding to a plurality of images, including the first image.
[0170] In one embodiment, the computing device 100 may determine the majority value of the primary prediction result of the anomaly corresponding to the first image and a predetermined second number of previously acquired images of the first image, and use the determined majority value to generate the secondary prediction result of the first anomaly. For example, the majority value may be determined as the result value that accounts for a higher proportion of the primary prediction result of the anomaly, among the result value indicating the presence of the anomaly and the result value indicating the absence of the anomaly.
[0171] As another example, the mainstream value may be determined by comparing the result value present in the primary prediction of anomalies with a predetermined second threshold. For example, if the second threshold is 45% and the proportion of a particular result value in the primary prediction of anomalies is 50%, the secondary prediction of anomalies may be set to that particular result value.
[0172] As described above, the technology according to one embodiment of the present disclosure can further improve the reliability and accuracy of the anomaly determination result by performing a first voting that utilizes the primary prediction result of the anomaly.
[0173] In one embodiment, the secondary prediction result for an anomaly may include a quantitative value used as a parameter for determining an anomaly alarm corresponding to the acquired image. For example, the secondary prediction result for an anomaly may include a counter value. For instance, if the secondary prediction result for a second anomaly corresponding to a second image acquired previously for the first image has a value of 2, it may be determined that a value of 1 is added according to the result of the first voting corresponding to the first image. In this case, a value of 2+1=3 may be included in the secondary prediction result for the first anomaly corresponding to the first image. In another example, if the secondary prediction result for a second anomaly corresponding to a second image acquired previously for the first image has a value of 2, it may be determined that a value of 1 is decreased according to the result of the first voting corresponding to the first image. In such a case, a value of 2-1=1 may be included in the secondary prediction result for the first anomaly corresponding to the first image.
[0174] In one embodiment, the unit of the counter value to be increased or decreased may be determined based on the difference between the image acquisition times. For example, if the counter value corresponding to the second image is 0.7, and the first image is acquired 500ms or more after the acquisition time of the second image, and the result of the first voting corresponding to the first image is determined to add a value, the counter value corresponding to the first image may be set to 0.7 + 0.5 = 1.2.
[0175] In one embodiment, the secondary prediction result of the first anomaly may be compared with a predetermined counter threshold. For example, if the counter value corresponding to the secondary prediction result of the first anomaly is greater than or equal to the predetermined counter threshold, the computing device 100 may decide to generate an alarm (e.g., turn the alarm ON). For example, if the counter value corresponding to the secondary prediction result of the second anomaly is greater than or equal to the counter threshold, and the counter value corresponding to the secondary prediction result of the first anomaly changes to less than the counter threshold, the computing device 100 may decide to turn the alarm OFF.
[0176] In one embodiment, the secondary prediction result for the first anomaly may include multiple counters. Including multiple counters in this way may generate multiple types of alarms. For example, a first counter and a second counter may be included in the secondary prediction result for the first anomaly. As a result of the first voting, the values corresponding to the first and second counters can be changed independently. Each of the first and second counters is compared with pre-assigned first and second counter thresholds, and if the counter value is greater than or equal to the counter threshold, an alarm corresponding to each counter may be generated. As an example, each counter may have a pre-defined minimum and maximum range, and if it falls outside the minimum and maximum range according to the result of the first voting, the counter value may be set to have a minimum and maximum range.
[0177] In one embodiment, the secondary prediction result of the first anomaly may be configured to generate multiple alarms by being compared with a plurality of counter thresholds. For example, the first counter value corresponding to the first image may be compared with a first counter threshold and a second counter threshold, respectively. If any one of these counter thresholds is met, a first alarm corresponding to that counter threshold may be generated. Furthermore, if any other of the counter thresholds is met, a second alarm corresponding to that counter threshold may be generated.
[0178] In one embodiment, the computing device 100 may determine an anomaly alarm corresponding to the first image by performing a second voting using the secondary prediction result of the first anomaly (450).
[0179] In this disclosure, the first voting may utilize the primary prediction result of a previous anomaly corresponding to at least one previously acquired image of the first image subject to anomaly determination, and the primary prediction result of a first anomaly corresponding to the first image. The second voting may utilize the secondary prediction result of a previous anomaly corresponding to at least one previously acquired image, and the secondary prediction result of a first anomaly corresponding to the first image.
[0180] In one embodiment, the second voting may be performed after the first voting. In one embodiment, the second voting may utilize the results of the first voting.
[0181] In one embodiment, the second voting can determine whether the abnormality detection results for a predetermined number of images are continuous. In one embodiment, if the abnormality detection results for a predetermined number of images are continuous with values indicating the presence of an abnormality, the second voting may decide to generate an abnormality alarm.
[0182] In one embodiment, the second voting may determine a set of images consisting of a first image corresponding to the current image and previously acquired sequential images of the first image, in order to ensure accuracy in the generation of anomaly alarms. The second voting may determine whether or not there is continuity in the secondary prediction results of anomalies corresponding to the images constituting the set of images. The second voting is a process that utilizes whether or not there is continuity in the secondary prediction results of anomalies. For example, if all secondary prediction results of anomalies in the set of images consisting of sequential images including the first image indicate the presence of an anomaly, the computing device 100 may decide to generate an anomaly alarm corresponding to the first image. For example, if some of the secondary prediction results of anomalies in the set of images consisting of sequential images including the first image indicate the presence of an anomaly, and other parts indicate the absence of an anomaly, the computing device 100 may decide that there is no continuity in the secondary prediction results of anomalies.
[0183] For example, suppose the number of criteria for determining continuity is 3. Under this assumption, if the secondary prediction results for the first anomaly, the second anomaly, and the third anomaly corresponding to the three sequential images, including the first image currently being judged for anomaly, do not indicate the presence of anomalies, the result value of the anomaly alarm corresponding to the first image may be set to 0 via the second voting. In such an example, no anomaly alarm is generated for the first image currently being acquired.
[0184] In one embodiment, the second voting may include comparing a secondary prediction result of an anomaly, including a counter value, with a counter threshold. For example, suppose the counter threshold is 3. In such an example, when the secondary prediction result of an anomaly reaches 3, the alarm corresponding to the counter threshold can be turned ON. Also, when the secondary prediction result of an anomaly changes from 3 to 2, the alarm can be turned OFF. As described above, the second voting may be performed by comparing each of several counter values with each of the counter thresholds. The second voting may be performed by comparing each of a single counter value with several counter thresholds.
[0185] As described above, the technology according to one embodiment of the present disclosure may utilize one or more boats to generate alarms in response to anomalies. By utilizing one or more boats, the sensitivity, accuracy, and reliability of alarms in response to anomalies can be increased.
[0186] Figure 5 illustrates an exemplary method for determining driver distraction according to one embodiment of the present disclosure.
[0187] In one embodiment, the computing device 100 may acquire images including the driver inside the vehicle (510).
[0188] In one embodiment, the computing device 100 may acquire an image including the driver's face inside the vehicle.
[0189] In one embodiment, the image refers to an image acquired from a camera. In another embodiment, the image may refer to an image of the driver taken by a camera installed inside the vehicle.
[0190] In one embodiment, the image may correspond to an image or frame taken at a specific point in time. In one embodiment, the image may mean a still image or frame extracted or captured from a video of a driver taken by a camera.
[0191] In one embodiment, the computing device 100 may receive multiple images or videos acquired by a camera. The computing device 100 may acquire or extract specific images (e.g., specific frames) from the acquired images or videos in order to determine the driver's gaze or the driver's distraction. For example, the images subject to the distraction determination may be selected or determined frames from among multiple frames. In such an example, the computing device 100 may extract specific frames randomly from among multiple frames or in units of predetermined time periods.
[0192] Images containing the driver's face in this disclosure may include, for example, an image in which the driver's face is shown as a bounding box, an image in which the contour of the driver's face is segmented, an image containing the driver's face, and / or an image in which the feature points of the driver's face are shown.
[0193] In one embodiment, the computing device 100 may, in response to image acquisition, use the first model to determine which of a plurality of gaze classes corresponds to the image (520).
[0194] In one embodiment, the first model may correspond to a pre-trained artificial intelligence-based model. The first model may correspond to a pre-trained artificial intelligence model using a training dataset generated based on clustering of reference images from a plurality of images that satisfy the condition that the vehicle's speed is above a predetermined critical speed. Here, the vehicle's speed may be mapped to the images. The reference image may mean an image that satisfies the condition corresponding to the mapped speed. For example, the reference image may mean an image in which the vehicle's speed is above a predetermined critical speed. For example, the reference image may mean an image in which the vehicle's speed at the time of acquisition is 30 km / h or higher.
[0195] In one embodiment, the first model may be trained to output a gaze class to which an image or face angle belongs in response to an input related to an image or an input related to the angle of a face in an image. In one embodiment, the first model may be trained to output the distance between an image or face angle and a plurality of clusters (groups) in response to an input related to an image or an input related to the angle of a face in an image.
[0196] In one embodiment, the gaze class may include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking in a direction other than straight ahead. In this disclosure, the forward gaze can be defined as the range of angles that can be identified as facing the direction of vehicle movement with respect to the driver's seat. In this disclosure, the forward gaze may be defined as a predetermined range of angles with the direction of vehicle movement as the reference axis. Such a forward gaze may encompass a two-dimensional or three-dimensional range of angles. In this disclosure, the range of angles outside the forward gaze range may be defined as a direction other than straight ahead.
[0197] In one embodiment, the computing device 100 may determine the gaze class (e.g., frontal or non-frontal) corresponding to the input image by extracting yaw values and pitch values from the input image and calculating the distance between the extracted combination of yaw and pitch values and the clustering result of the reference image. For example, the gaze angle may be determined by the combination of yaw and pitch values. For example, the computing device 100 may determine, by calculating the distance, the gaze cluster among a plurality of gaze clusters that is closest to the combination of yaw and pitch values. The gaze class of the image may be determined by the gaze class to which such gaze cluster belongs. Herein, gaze clusters may be generated by clustering or grouping the reference image using the driver's gaze angle in the reference image. In this specification, clustering and grouping may be used interchangeably with each other. In this specification, clusters and groups may be used interchangeably with each other.
[0198] In one embodiment, the first model may correspond to a classification model. The computing device 100 may operate to distinguish between forward and non-forward based on input values when yaw and pitch values are input using the first model.
[0199] In this disclosure, distraction and forward inattention may be used interchangeably.
[0200] In one embodiment, the computing device 100 may determine whether or not driver distraction is present in the image based on the gaze class corresponding to the image (530).
[0201] For example, the computing device 100 may determine a gaze class corresponding to an input image and determine whether or not driver distraction exists depending on the type of gaze class. In such an example, if the gaze class corresponding to the input image is determined to be a first gaze class representing a forward view, the computing device 100 may determine that there is no driver distraction. In such an example, if the gaze class corresponding to the input image is determined to be a second gaze class representing a non-forward view, the computing device 100 may determine that driver distraction exists.
[0202] For example, the computing device 100 may extract yaw and pitch values from the input image and use the distance between the gaze cluster corresponding to the input image and the extracted yaw and pitch values to determine whether or not driver distraction is present in the input image.
[0203] For example, the computing device 100 may set the likelihood of driver distraction higher the greater the distance between at least one cluster belonging to the first line-of-sight class representing the front and the yaw and pitch values extracted from the input image. For example, the computing device 100 may set the likelihood of driver distraction lower the greater the distance between at least one cluster belonging to the first line-of-sight class representing the front and the yaw and pitch values extracted from the input image.
[0204] For example, the computing device 100 may set the likelihood of driver distraction higher the closer the distance is between at least one cluster belonging to the second gaze class representing non-forward views and the yaw and pitch values extracted from the input image. For example, the computing device 100 may set the likelihood of driver distraction lower the further the distance is between at least one cluster belonging to the second gaze class representing non-forward views and the yaw and pitch values extracted from the input image.
[0205] Figure 6 illustrates the line of sight angles based on the camera's installation position in a DMS.
[0206] The camera product for achieving DMS operation can be installed at various locations 610a, 630a, 650a, and 670a inside the vehicle.
[0207] Reference numeral 610a indicates the position of the windshield inside the vehicle. When a camera is installed at the position of reference numeral 610a, an image of the driver, such as that shown by reference numeral 610b, may be acquired.
[0208] Reference numeral 630a indicates the location of the rearview mirror inside the vehicle. When a camera is installed at the location of reference numeral 630a, an image of the driver, such as that shown in reference numeral 630b, may be acquired.
[0209] Reference numeral 650a indicates the location of the dashboard inside the vehicle. When a camera is installed at the location of reference numeral 650a, an image of the driver, such as that shown by reference numeral 650b, may be obtained.
[0210] Reference numeral 670a indicates the location of the center fascia inside the vehicle. When a camera is installed at the location of reference numeral 670a, an image of the driver, such as that shown by reference numeral 670b, may be acquired.
[0211] As mentioned above, the driver's line of sight angle in the image may be determined to differ depending on the camera's installation position. In particular, when aftermarket DMS products are used, where the installation position within the vehicle cannot be determined, it can be difficult to determine the driver's reference front position and thus difficult to capture a standard for determining the driver's distraction.
[0212] The technology according to one embodiment of the disclosed material can achieve the technical effect of efficiently distinguishing between forward gaze and non-forward gaze from the driver's image, even when the camera's installation location is not specified.
[0213] Figure 7 illustrates an example of a method for collecting training data to train a model according to one embodiment of the present disclosure.
[0214] In one embodiment, the computing device 100 may acquire images of the driver taken inside a moving vehicle. The computing device 100 may utilize the fact that when the driving speed is above a predetermined reference speed, the driver's viewing angle is likely to be straight ahead. By determining images that meet these conditions from the collected images, the computing device 100 can build training data for a model to detect driver distraction.
[0215] As shown in Figure 7, images of drivers at various driving speeds can be collected. The two-dimensional space in Figure 7 represents the driver's line of sight angle in the collected images. For example, in the two-dimensional space, the X-axis or horizontal axis may represent the yaw value, and the Y-axis or vertical axis may represent the pitch value. The driver's line of sight angle may be determined based on a combination of the yaw and pitch values. For example, the line of sight angle may represent a two-dimensional vector obtained from the yaw and pitch values.
[0216] In one embodiment, the computing device 100 may map the collected images onto a two-dimensional space using the yaw and pitch values of each of the collected images. The computing device 100 can map the collected images onto a two-dimensional space using the viewing angle of each of the collected images. A single point displayed in the two-dimensional space may correspond to a single image.
[0217] In one embodiment, the vehicle's speed corresponding to the acquisition time of each collected image may be determined. The determined speed may be mapped to each of the collected images. For example, when a particular image is acquired, the vehicle's speed at the time that image was acquired may also be acquired. For example, the speed for each image may be stored together with the image.
[0218] In one embodiment, the computing device 100 may distinguish between images in which the vehicle's speed exceeds a predetermined reference speed and images in which the vehicle's speed is below a predetermined reference speed. In Figure 7, the reference speed is exemplified as 30 km / h. Reference numeral 710 (710a and 710b) indicates images in which the vehicle's speed is below the reference speed. Reference numerals 720, 730 and 740 indicate images in which the vehicle's speed exceeds the reference speed.
[0219] In one embodiment, the computing device 100 can assign reference images to multiple clusters by clustering reference images that exceed a certain speed. When a certain number of images (or reference images) or more have been collected, the computing device 100 may, after proceeding with clustering of the collected reference images, determine that viewing angles accounting for a proportion greater than or equal to a certain standard ratio of the total number are front clusters. After proceeding with clustering of the collected reference images, the computing device 100 may determine that viewing angles accounting for a proportion less than a certain standard ratio of the total number are non-front clusters.
[0220] In the example in Figure 7, 62 images are collected and mapped into a two-dimensional space. Of these images, images 710:710a and 710b, which are below the reference velocity, do not meet the clustering criteria and can therefore be excluded from clustering. A total of 46 reference images are collected, excluding images 710:710a and 710b, which are below the reference velocity. Clustering can be performed on these reference images. Such clustering may include grouping points in the two-dimensional space into multiple groups based on a combination of yaw and pitch values or face angle. The computing device 100 may generate multiple groups or clusters 720, 730, and 740 as a result of clustering on the reference images. The computing device 100 may determine quantitative values of the images (i.e., reference images) included in each of the clusters 720, 730, and 740. The computing device 100 may determine the number of images (i.e., reference images) contained in each of the clusters 720, 730, and 740. In the example in Figure 7, the first cluster 720 contains 4 reference images, the second cluster 730 contains 18 reference images, and the third cluster 740 contains 28 reference images. As an example, the computing device 100 can distinguish between frontal and non-frontal clusters by comparing the ratio of the number of images in each cluster to the overall image or overall reference image with a predetermined critical ratio. As another example, the computing device 100 can distinguish between frontal and non-frontal clusters by comparing the number of images in each cluster with a predetermined critical number. In the example in Figure 7, the first cluster 720 may be determined to be a non-frontal cluster because it contains fewer reference images than the critical ratio or critical number. The second cluster 730 and the third cluster 740 may be determined to be frontal clusters because they contain more reference images than or equal to the critical ratio or critical number.Therefore, the computing device 100 can construct a training dataset by applying non-frontal labeling to images contained in the first cluster 720, and frontal labeling to images contained in the second cluster 730 and the third cluster 740. A model for determining attentional distraction may be trained using such a training dataset.
[0221] As described above, the computing device 100 can collect gaze angles (e.g., yaw and pitch values) collected under conditions of a certain speed or higher while the vehicle is in motion. The computing device 100 can determine the ground truth through clustering for each of the collected gaze angles and construct a model training dataset in a manner that assigns the ground truth to the collected gaze angles. The ground truth here may be given for each clustered group, and if the total number of images in the group is equal to or greater than a baseline ratio, the ground truth corresponding to forward views may be assigned to the images, and if the total number of images in the group is not equal to or greater than a baseline ratio, the ground truth corresponding to non-forward views may be assigned to the images. A model for determining distraction or forward inattention using the training dataset may be trained to output a gaze class result and the distance to the corresponding class when a new gaze angle (e.g., yaw and pitch values) is input, using the training dataset which includes the collected images and the ground truth corresponding to the images.
[0222] Figures 8A, 8B, and 8C illustrate a model training and inference method according to one embodiment of the present disclosure.
[0223] In one embodiment, Figure 8(A) illustrates the model's learning process. A clustering model 820a may be used to group multiple data into multiple groups. The reference image 810a can represent a set of images that satisfy a critical number condition. The clustering model 820a may cluster the reference image 810a to produce a grouping result 830a containing multiple groups or multiple clusters. The clustering model 820a may correspond to a pre-trained artificial intelligence-based model that clusters input data, for example, by including similar data in one cluster based on gaze angle values. As an unrestrictive example, the clustering model 820a may operate using a rule-based algorithm that groups or distinguishes data based on data values. Clustering of the reference image 810a may include grouping the reference image 810a into multiple gaze clusters based on the driver's gaze angle (e.g., a combination of yaw and pitch values) within the reference image 810a. The grouping result 830a may correspond to the illustrative content shown in Figure 7.
[0224] In one embodiment, the ground truth (forward or non-forward) for each cluster within the grouped result 830a may be determined based on the quantitative values of the images contained in each of the multiple groups or multiple clusters. A training dataset 840a can be constructed that includes the images within the clusters and the ground truth for each cluster.
[0225] The training dataset 840a may be generated by determining at least one cluster corresponding to a major viewing angle (e.g., a major viewing angle range) as a forward class, and at least one cluster corresponding to an angle other than the major viewing angle as a non-forward class. For example, the major may be used to represent the proportion or number exceeding a certain threshold. Another example is that the major may be used to represent the cluster with the most images among the clusters. Yet another example is that the major may be used to represent a cluster with a predetermined rank based on the number of images among the clusters.
[0226] The training dataset 840a may be generated by labeling each of the multiple gaze clusters, generated by clustering the reference images 810a, into a first gaze class or a second gaze class based on the quantitative information of the images contained in each of the multiple clusters. For example, the training dataset may be generated in a manner in which, among the multiple clusters, at least one cluster containing reference images with a predetermined critical ratio or higher relative to the total number of reference images 810a is labeled into a first gaze class, and among the multiple clusters, at least one gaze cluster containing reference images with a predetermined critical ratio or lower relative to the total number of reference images 810a is labeled into a second gaze class. For illustrative purposes, two gaze classes (forward or non-forward) are described, but embodiments in which gaze classes are labeled into three or more gaze classes may also be included within the scope of this disclosure, depending on the implementation.
[0227] In one embodiment, a first model 850a for determining attentional distraction may be trained using a training dataset 840a. The first model 850a can be trained to take the training dataset 840a as input and output a result 860a. For example, result 860a may include a result indicating whether the input data has a forward class or not. For example, result 860a may include a result indicating whether or not attentional distraction is present. For example, result 860a may include an attentional distraction score indicating the likelihood of attentional distraction. For example, result 860a may include a gaze angle extracted from the input data and a value representing the distance between each of several clusters. For example, result 860a may include a value indicating the distance between a gaze angle extracted from the input data and the cluster to which that gaze angle belongs.
[0228] In one embodiment, the first model 850a may be updated according to a predetermined period or condition. For example, when a reference image 810a that meets the condition of being above a certain travel speed is retrieved by the computing device 100, the reference image 810a may be added within a predetermined queue size. If data exceeding the queue size is retrieved, the data in the queue may be updated in a manner in which the oldest data in the queue is deleted. If data exceeding the queue size is retrieved, the data in the queue may be updated in a manner in which a specific data determined probabilistically in the queue is deleted. As an example, the training dataset 840a may be updated based on the above-described data update method, thereby further advancing the training of the first model 850a.
[0229] In one embodiment, the training of the first model 850a may be performed in predetermined periodic units. As the training of the first model 850a progresses, the weights of the first model 50a are updated.
[0230] In one embodiment, Figure 8(B) illustrates the inference process of a pre-trained model.
[0231] In one embodiment, the first model 850a may correspond to a pre-trained classification model so that when new yaw and pitch values are input, it outputs the distance between the classification results of the corresponding values as result 860a.
[0232] As an unrestrictive example, the first model 820b in Figure 8(B) may correspond to the first model 850a trained in Figure 8(A).
[0233] As an unrestrictive example, image 810b in Figure 8(B) may be a different image from the reference image 810a in Figure 8(A). In such an example, image 810b may correspond to an image collected regardless of the condition that the driving speed is above the reference speed.
[0234] In one embodiment, the first model 820b may determine the gaze angle of the face from the image 810b. The first model 820b may determine the gaze angle of the face by extracting yaw and pitch values from the image 810b. The gaze angle of the face may be included in the result 830b.
[0235] In one embodiment, the first model 820b may determine the distance between the extracted yaw and pitch values and the grouping result 830a during the learning process. For example, the first model 820b may determine the distance between a plurality of clusters included in the grouping result 830a and the face angles determined from the extracted yaw and pitch values. The said distance may be included in the result 830b.
[0236] In one embodiment, the first model 820b may determine the cluster corresponding to image 810b from among a plurality of clusters included in the grouping result 830a of the reference image 810a. The cluster corresponding to image 810b may be included in the result 830b. In one embodiment, the first model 820b may determine the gaze class corresponding to image 810b from among a plurality of gaze classes to belong to the gaze class to which the cluster corresponding to image 810b belongs. Such a gaze class may be included in the result 830b.
[0237] In one embodiment, the first model 820b may output a result 830b indicating whether or not driver distraction is present in image 810b. In one embodiment, the first model 820b may extract yaw values and pitch values from image 810b and use the distance between the cluster corresponding to image 810b and the extracted yaw and pitch values from among a plurality of learned clusters to determine whether or not driver distraction is present in image 810b. The first model 820b may set a higher distraction score indicating the possibility of driver distraction the greater the distance between at least one cluster belonging to the gaze class corresponding to forward and the yaw and pitch values extracted from image 810b, or the greater the distance between at least one cluster belonging to the second gaze class corresponding to non-forward and the yaw and pitch values extracted from image 810b. Such a distraction score may also be included in result 830b.
[0238] Figure 8(C) shows an example in which a second model 820c for recognizing another face is used in the inference process. The first model 830d and the first model 820b may correspond to each other, and such first models 830d and 830b may correspond to the first model 850a after training is complete. Images 810c and 810b in Figure 8 may correspond to each other.
[0239] In one embodiment, the yaw and pitch values 830c corresponding to the driver's face in image 810c may be generated by a second model 820c different from the first model 830d. The second model 820c may correspond to an artificial intelligence-based model pre-trained to output the face angle corresponding to the input image 810c. The second model 820c may correspond to an artificial intelligence-based model pre-trained to output the yaw and pitch values 830c corresponding to the input image 810c. The second model 810c may correspond to an artificial intelligence-based model pre-trained to output the yaw and pitch values 830c corresponding to the driver's face in image 810c from image 810c.
[0240] In one embodiment, the first model 830d may generate result 830e using the yaw and pitch values 830c generated by the second model 810c as input. Result 830e here may correspond to result 830b in Figure 8(B) described above.
[0241] Figure 9 illustrates an exemplary method for determining a driver distraction alarm according to one embodiment of the present disclosure.
[0242] In the explanation of Figure 9, any content that overlaps with the explanation given above will be omitted to avoid repetition.
[0243] For example, the methods related to the first and second voting in Figure 4 may correspond to the methods related to the first and second voting in Figure 9. For example, the model output information in Figure 4 may correspond to the attention distraction score in Figure 9. For example, the thresholds in Figure 4 may correspond to the thresholds in Figure 9. For example, the primary and secondary prediction results for anomalies in Figure 4 may correspond to the primary and secondary prediction results for attention distraction in Figure 9, respectively.
[0244] A technology according to one embodiment of the present disclosure may determine the driver's gaze, forward inattention, and / or distraction from an image of the driver, and determine an alarm based on the determined result. Determining an alarm is used to include determining whether or not to generate an alarm, and determining the type and / or intensity of the alarm. Such a technology according to one embodiment of the present disclosure can ensure the accuracy and reliability of alarms provided to the driver in the DMS.
[0245] The technology according to one embodiment of the Disclosure may use one or more thresholds, one or more counters, and / or one or more voting in the process of determining the driver's gaze, forward inattention, and / or distraction from the driver's image. Such technology according to one embodiment of the Disclosure enables a more accurate determination in the DMS whether or not driver distraction is present.
[0246] In one embodiment, the computing device 100 may acquire an attention distraction score corresponding to the target image (910).
[0247] In one embodiment, the distraction score may include a result value that quantitatively indicates the likelihood of driver distraction in the image. In one embodiment, the distraction score may include a result value that quantitatively indicates the likelihood of the driver in the image looking away from the road. In one embodiment, the distraction score may include a result value that quantitatively indicates the likelihood of the driver in the image not looking straight ahead or forward. In one embodiment, the distraction score may include a result value that quantitatively indicates the likelihood of the driver in the image not concentrating on driving. Such a distraction score may be generated based on the output of the artificial intelligence-based classification model described above.
[0248] In one embodiment, the target image represents an object for determining whether or not attention distraction is present. In another embodiment, the target image represents an object for determining an attention distraction alarm.
[0249] In one embodiment, the computing device 100 may compare the attention distraction score with a threshold to obtain a primary prediction result of attention distraction indicating whether or not attention distraction is present in the target image (920).
[0250] For example, a threshold is a threshold used as a criterion for determining whether or not attention is distracted. For example, a threshold is a threshold used as a criterion for determining the primary prediction result of attention distraction.
[0251] In one embodiment, the threshold may be determined based on at least one previous result corresponding to at least one previously acquired image of the target image. The threshold may be variable based on at least one previous result corresponding to at least one previously acquired image of the target image. The threshold may be variable based on at least one previous result corresponding to at least one previously acquired image of the target image, with respect to the previous threshold of the image immediately preceding the one acquired immediately before the target image. The expression that the threshold is variable may be used to include having the same value as the previous threshold or having a different value from the previous threshold.
[0252] In one embodiment, the previous result can represent a judgment or prediction of driver distraction with respect to previously acquired images of the target image. For example, if the ratio of result values indicating the presence of distraction in the previous result corresponding to previously acquired images of the target image (e.g., sequentially acquired previous images) is greater than or equal to a first ratio, the threshold is set to a first value. If the ratio of result values indicating the presence of distraction in the previous result corresponding to the previously acquired images of the target image is less than the first ratio, the threshold may be set to a second value higher than the first value. In such an example, assume that the first, second, and third images of the target image have been acquired prior to the target image, and that the result (1) indicating the presence of distraction is derived from the first image, the result (0) indicating the absence of distraction is derived from the second image, and the result (1) indicating the presence of distraction is derived from the third image. Under these assumptions, the threshold for determining attention distraction in the target image may be set lower than the threshold for determining the result of the first image (i.e., the image acquired immediately before the target image), depending on the results of the previous images (first, second, and third images) that indicate the presence of attention distraction (e.g., a percentage of 66.6%). In such an example, if the previous threshold is low, the threshold corresponding to the target image may be set to the same value as the previous threshold. In an additional example, the threshold may be variably determined by considering both the previous results and the results corresponding to the target image.
[0253] In one embodiment, the primary prediction result for attention distraction may correspond to the primary prediction result for anomalies in Figure 4. The primary prediction result for attention distraction may be determined based on a comparison between an attention distraction score obtained from a model or an attention distraction score obtained by a rule-based algorithm and a threshold. If the attention distraction score exceeds the threshold, it is determined that attention distraction is present in the target image. If attention distraction is present in the target image, the primary prediction result for attention distraction may have a value of 1. If the attention distraction score does not exceed the threshold, it is determined that attention distraction is not present in the target image. If attention distraction is not present in the target image, the primary prediction result for attention distraction may have a value of 0.
[0254] In one embodiment, the computing device 100 may obtain a secondary prediction result for attention distraction by performing a first voting using the primary prediction result for attention distraction (930).
[0255] In one embodiment, the computing device 100 may determine a distraction alarm corresponding to an input target image by performing a first voting using the primary distraction prediction result.
[0256] In one embodiment, the computing device 100 may use the primary distraction prediction results for the target image and previous images to generate a secondary distraction prediction result corresponding to the target image. For example, the secondary distraction prediction result may indicate whether or not a state of driver distraction exists in the image. For example, the secondary distraction prediction result may indicate whether or not a distraction alarm corresponding to the image will be generated. For example, the secondary distraction prediction result may be used as a factor for determining the type of distraction alarm corresponding to the image. For example, the secondary distraction prediction result may be used as a factor for determining the intensity of the distraction alarm corresponding to the target image. The secondary distraction prediction result may be used as a factor for determining the occurrence of a distraction alarm corresponding to the image.
[0257] In one embodiment, the first voting may determine which of the primary distraction prediction result for the target image subject to attention distraction judgment and the primary distraction prediction result for previous images is the dominant value, and use the determined dominant value to generate a secondary distraction prediction result corresponding to the target image.
[0258] In one embodiment, the first voting may generate a collective attention-distraction prediction result representing an image group consisting of a target image and a predetermined first number of sequential images previously acquired of the target image. In one embodiment, the first voting may be used to generate a secondary attention-distraction prediction result indicating whether or not attention-distraction exists in the target image, utilizing the collective attention-distraction prediction result. The collective attention-distraction prediction result here may include a result value that represents the group, among result values indicating the presence of attention-distraction and result values indicating the absence of attention-distraction. Such a collective attention-distraction prediction result may represent a single prediction result that represents multiple primary attention-distraction prediction results. For example, if images with results indicating attention-distraction have a high proportion among multiple images, the prediction result that represents the group (i.e., the prediction result for each image included in the group) may be determined to be in an attention-distraction state.
[0259] In additional embodiments, the attention distraction score may be additionally reflected in the outcome value that represents the set. In such embodiments, the attention distraction score may be used as a weight or factor when determining the outcome value that represents the set.
[0260] In one embodiment, it is assumed that there are primary distraction prediction results corresponding to five sequential images, including a target image, and that these have values of 0, 0, 1, 1, and 0, respectively. Under this assumption, 0 can represent the absence of distraction, and 1 can represent the presence of distraction. A secondary distraction prediction result corresponding to the target image may be determined by a combination of the multiple primary distraction prediction results. Under the above assumption, the collective distraction prediction result representing the five primary distraction prediction results may have a value of 0. Thus, the first secondary distraction prediction result corresponding to the target image may have a value of 0.
[0261] In one embodiment, the secondary prediction result of attention distraction may be represented by a counter value. The counter value may be determined using the result of a first voting. Depending on the result of the first voting, it may be determined whether the counter value is added, decreased, or maintained. Based on the first voting and a critical range of the counter value, it may be determined whether the counter value is added, decreased, or maintained. The critical range of the counter value can define a maximum value beyond which the counter value will not increase, but will be maintained or decreased, and a minimum value beyond which it will not decrease, but will be maintained or increased. The computing device 100 may generate one or more current counter values corresponding to the current image in a manner that maintains, increases, or decreases one or more previous counter values corresponding to previously acquired images of the target image, based on the result of the first voting, and may also obtain a secondary prediction result of attention distraction including the one or more current counter values.
[0262] In one embodiment, the unit of increment or decrement of the counter may be set in various forms. For example, a predetermined fixed value may be used as the unit of increment or decrement. For example, the unit of increment or decrement of the counter may be determined based on the difference between the acquisition time of the previous image and the acquisition time of the current image. In such an example, if the difference between the acquisition time of the previous image and the acquisition time of the current image is 15 ms, and the primary prediction result of attention distraction includes the presence of attention distraction, the counter value of the target image can be increased by 15 or 0.15 compared to the counter value of the previous image.
[0263] In one embodiment, the computing device 100 may determine an attention distraction alarm corresponding to the target image by performing a second voting using the secondary prediction result of attention distraction (940).
[0264] In one embodiment, the computing device 100 may determine whether or not to generate a distraction alarm corresponding to the target image, or whether or not driver distraction is present in the target image, by performing a second voting using the secondary distraction prediction results. In one embodiment, the computing device 100 may determine the intensity or type of the distraction alarm corresponding to the target image by performing a second voting using the secondary distraction prediction results.
[0265] In one embodiment, the second voting may determine whether there is a continuity of secondary distraction prediction results within an image set consisting of a target image subject to attention distraction judgment and a predetermined second number of sequential images of the target image previously acquired. The second voting may utilize combinations of secondary distraction prediction results. Based on whether there is a continuity of secondary distraction prediction results, the second voting may determine whether to generate an alarm corresponding to the target image, or whether user distraction is present in the target image.
[0266] In one embodiment, the second voting may include comparing a counter value with a counter threshold. The second voting may be used to determine whether one or more distraction alarms corresponding to an image are ON or OFF by comparing one or more current counter values with one or more predetermined counter thresholds. For example, if the current counter value reaches a certain threshold, the distraction alarm may be set ON, and if the current counter value falls below a certain threshold, the distraction alarm may be set OFF. Based on the comparison of each of a plurality of counter values with each of a plurality of thresholds, if a certain counter value exceeds a certain threshold, the type or intensity of the distraction alarm may be determined by a combination of a certain counter value and a certain threshold. In one embodiment, the degree to which the counter value exceeds the alarm threshold and the intensity of the alarm may be interrelated. The intensity of the alarm may be determined to differ depending on the degree to which the counter value exceeds the alarm threshold. The excess value of the counter value exceeding the alarm threshold and the intensity of the alarm may have a positive correlation. In such embodiments, the greater the degree to which the counter value exceeds the alarm threshold, the greater the intensity of the alarm may be. The intensity of the alarm here can represent the magnitude of an alarm image, alarm sound, alarm light, and / or alarm vibration.
[0267] In one embodiment, the second voting may use multiple alarm thresholds. For example, the multiple alarm thresholds may have different alarm intensities and / or alarm types. For another example, some of the multiple alarm thresholds may have the same alarm intensity and / or alarm type. The second voting may determine the type of alarm, the intensity of the alarm, and / or whether an alarm occurs by comparing each of the multiple alarm thresholds with a counter corresponding to a specific image. For example, a first alarm, which occurs when the counter exceeds one of the multiple alarm thresholds, may have a lower intensity than a second alarm, which occurs when the counter exceeds two or more alarm thresholds. For example, a first alarm, which occurs when the counter exceeds one of the multiple alarm thresholds, and a second alarm, which occurs when the counter exceeds two or more alarm thresholds, may have different types of alarms. For example, when the counter exceeds multiple alarm thresholds, an alarm with a higher priority or intensity may be selected from among the alarms corresponding to each of the multiple alarm thresholds. In this case, alarms with lower priority or intensity will not occur even if the alarm threshold is exceeded.
[0268] In one embodiment, multiple counters may be assigned to the target image. The second voting may determine whether an alarm is generated for each of the multiple counters by comparing each of the multiple counters with an alarm threshold. For each of the counters, the type or intensity of the alarm may be the same or different from one another. For example, a first alarm, which occurs when one counter exceeds the alarm threshold, may have a lower intensity than a second alarm, which occurs when multiple counters exceed the alarm threshold.
[0269] In one embodiment, the second voting may determine whether an alarm has occurred, the type of alarm, and / or the intensity of an alarm by comparing each of the multiple counters with each of the multiple alarms.
[0270] As described above, the technology according to one embodiment of the present disclosure may utilize one or more boatings, one or more counters, and one or more thresholds to generate an alarm in response to distraction and / or to determine whether or not the driver is distracted. This can increase the sensitivity, accuracy, and reliability of the alarm in response to distraction. By utilizing one or more boatings, the sensitivity, accuracy, and reliability of the judgment of distraction can all be increased.
[0271] Figure 10 illustrates an exemplary method for determining a driver abnormality alarm according to one embodiment of the present disclosure.
[0272] Figure 10 illustrates a methodology for generating an abnormal alarm according to one embodiment of the present disclosure.
[0273] The example shown in Figure 10 will be replaced with the previously explained content to avoid repetition of the explanation.
[0274] As shown in Figure 10, a voting technique according to one embodiment of the present disclosure may utilize a plurality of queues to determine an anomaly alarm from the output or results of a model (e.g., classification results, distance values and / or attention distraction scores), or to determine whether or not to generate an anomaly alarm. The queues in the present disclosure may be used interchangeably with a voting queue.
[0275] In one embodiment, the boating queue may include a first queue 1010, a second queue 1030, a third queue 1050, and a fourth queue 1070. Such first queue 1010, second queue 1030, third queue 1050, and fourth queue 1070 may be used to determine anomaly alarms according to one embodiment of the present disclosure. Although four queues are shown in Figure 10, it will be apparent to those skilled in the art that various numbers of queues are available depending on various implementations, such as adding new boatings or deleting existing boatings.
[0276] In one embodiment, the X-axis direction (i.e., the lateral direction) of the first queue 1010, second queue 1030, third queue 1050, and fourth queue 1070 is the direction for representing images acquired in time. The first queue 1010, second queue 1030, third queue 1050, and fourth queue 1070 are configured such that moving to the right approaches the current point in time, and moving to the left approaches a past point in time. The first queue 1010, second queue 1030, third queue 1050, and fourth queue 1070 contain multiple units. In Figure 10, the first queue 1010, second queue 1030, third queue 1050, and fourth queue 1070 are shown to contain five units corresponding to each other. It will be apparent to those skilled in the art that, depending on the mode of implementation, a single queue may contain a variety of units.
[0277] In one embodiment, each of the multiple units in the queue may correspond to each of the acquired images. In the first queue 1010, the second queue 1030, the third queue 1050, and the fourth queue 1070, the first units 1010a, 1030a, 1050a, and 1070a may include results corresponding to the first image from the first viewpoint; the second units 1010b, 1030b, 1050b, and 1070b may include results corresponding to the second image from the second viewpoint; the third units 1010c, 1030c, 1050c, and 1070c may include results corresponding to the third image from the third viewpoint; the fourth units 1010d, 1030d, 1050d, and 1070d may include results corresponding to the fourth image from the fourth viewpoint; and the fifth units 1010e, 1030e, 1050e, and 1070e may include results corresponding to the fifth image from the fifth viewpoint. Here, the first viewpoint is the viewpoint prior to the second viewpoint, the second viewpoint is the viewpoint prior to the third viewpoint, the third viewpoint is the viewpoint prior to the fourth viewpoint, and the fifth viewpoint is the viewpoint prior to the fourth viewpoint.
[0278] In one embodiment, the units corresponding to the first queue 1010, the second queue 1030, the third queue 1050, and the fourth queue 1070 may each contain the judgment result for the same image acquired at the same time. For example, the first image acquired at the first time may be processed in the following order: first unit 1010a of the first queue 1010, first unit 1030a of the second queue 1030, first unit 1050a of the third queue 1050, and first unit 1070a of the fourth queue 1070. As a result of such processing, the fourth queue 1070, which is the last queue, may contain the result of turning the abnormal alarm ON or OFF for each image.
[0279] In one embodiment, the first queue 1010 may include a first unit 1010a, a second unit 1010b, a third unit 1010c, a fourth unit 1010d, and a fifth unit 1010e. For example, the fifth unit 1010e may include a decision result corresponding to the currently (or most recently) retrieved image. In one embodiment, the first queue 1010 is a queue representing the output of the model or a queue showing the result of post-processing the output of the model. For example, the first queue 1010 may include model output information for each of a plurality of images. For example, the first queue 1010 is a queue showing the detection result or classification result of the model. The first unit 1010a includes first model output information corresponding to the first image of the first viewpoint, the second unit 1010b includes second model output information corresponding to the second image of the second viewpoint, the third unit 1010c includes third model output information corresponding to the third image of the third viewpoint, the fourth unit 1010d includes fourth model output information corresponding to the fourth image of the fourth viewpoint, and the fifth unit 1010e may include fifth model output information corresponding to the fifth image of the fifth viewpoint. Here, the first viewpoint is the earliest viewpoint, and viewpoints can be represented as you move from the first viewpoint to the fifth viewpoint. In addition, the sixth unit 1010f is assigned sixth model output information corresponding to the sixth image acquired at the sixth viewpoint, which is a future viewpoint. In one embodiment, the first queue 1010 may be set to include a predetermined number of units in a manner in which the oldest unit is removed when a new unit is drawn in. For example, if the sixth model output information is assigned to the sixth unit 1010f, the first unit 1010a is removed from the queue.
[0280] In one embodiment, the second queue 1030 may include a first unit 1030a, a second unit 1030b, a third unit 1030c, a fourth unit 1030d, and a fifth unit 1030e. The values contained in each of the units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030 may be determined based on a comparison between the information contained in the unit corresponding to the first queue 1010 and the threshold assigned to the corresponding unit of the first queue 1010. The values contained in each of the units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030 may include a primary prediction result of an anomaly.
[0281] In one embodiment, the second queue 1030 may include a primary prediction result for an anomaly determined based on a comparison between model output information and a threshold. The first unit 1030a includes a primary prediction result for a first anomaly corresponding to a first image of a first viewpoint, the second unit 1030b includes a primary prediction result for a second anomaly corresponding to a second image of a second viewpoint, the third unit 1030c includes a primary prediction result for a third anomaly corresponding to a third image of a third viewpoint, the fourth unit 1030d includes a primary prediction result for a fourth anomaly corresponding to a fourth image of a fourth viewpoint, and the fifth unit 1030e includes a primary prediction result for a fifth anomaly corresponding to a fifth image of a fifth viewpoint.
[0282] The expression “the threshold is variable” as used herein may be used to encompass embodiments in which the threshold is set to one of several (e.g., two) threshold options, and embodiments in which the threshold is increased, decreased, or kept equal to a previous threshold.
[0283] In one embodiment, the threshold assigned to the current unit of the first queue 1010 may be variable based on values included in previous units of other queues. For example, the threshold assigned to the current unit of the first queue 1010 may be determined based on the primary prediction result of the first anomaly of the previous unit of the first queue 1010. For example, in Figure 10, if the model output information exceeds the threshold, the primary prediction result of the anomaly may have a value of 1, and if the model output information does not exceed the threshold, the primary prediction result of the anomaly may have a value of 0. The primary prediction result of the first anomaly may be determined by comparing the first model output information of 0.6 with the threshold of 0.5, and may have a value of 1. The primary prediction result of the second anomaly may be determined by comparing the second model output information of 0.6 with the threshold of 0.5, and may have a value of 1. The primary prediction result of the third anomaly may be determined by comparing the third model output information of 0.6 with the threshold of 0.5, and may have a value of 1. The primary prediction result of the fourth anomaly can be determined by comparing the fourth model output information of 0.3 with the threshold of 0.4, and may have a value of 0. The primary prediction result for the fifth anomaly can be determined by comparing the fifth model output information of 0.6 with a threshold of 0.4, and may have a value of 1. As shown in Figure 10, if the value of a specific unit in the first queue 1010 is greater than or equal to the threshold assigned to that specific unit in the first queue 1010, the value included in that specific unit in the second queue 1030 is determined to be 1. If the value of a specific unit in the first queue 1010 is less than the threshold assigned to that specific unit in the first queue 1010, the value included in that specific unit in the second queue 1030 is determined to be 0. For example, the threshold assigned to the current unit of the first queue 1010 (e.g., the Nth unit) may be variable depending on the value included in the unit immediately preceding the current unit of the second queue 1030 (e.g., the N-1th unit). In such an example, if the value of the second unit 1030b of the second queue 1030 is set to 1, the threshold assigned to the third unit 1010c of the first queue 1010 may be determined to be 0.4 by subtracting 0.1 from the threshold of the previous unit, the second unit 1010b, which is 0.5.In another embodiment, if the value of the second unit 1030b of the second queue 1030 is set to 0, the threshold assigned to the third unit 1010c of the first queue 1010 may be kept equal to the threshold of the previous unit, the second unit 1010b, which is 0.5, or 0.1 may be added to determine a threshold of 0.6. In this specification, N may mean a natural number.
[0284] In one embodiment, the threshold corresponding to each unit of the first queue 1010 may be set to one of a first value and a second value. In such an embodiment, the threshold corresponding to the current unit of the first queue 1010 may be set to one of a first value and a second value using the result value corresponding to the previous unit. In such an embodiment, if the result value corresponding to the previous unit indicates the presence of an anomaly, the threshold corresponding to the current unit of the first queue 1010 may be set to the lower of the first value and the second value. In such an embodiment, if the result value corresponding to the previous unit indicates the absence of an anomaly, the threshold corresponding to the current unit of the first queue 1010 may be set to the higher of the first value and the second value.
[0285] In one embodiment, the current threshold value corresponding to the current unit may be maintained or changed based on the previous threshold value corresponding to the previous unit. In an example where the threshold value changes, for example, when it is determined that the predicted result of an abnormality corresponding to the previous unit indicates the presence of an abnormality, the current threshold value may be set to have a downward trend. Here, the downward trend may include keeping the current threshold value the same as the previous threshold value or setting the current threshold value lower than the previous threshold value. Setting the current threshold value low may mean that the possibility of determining an abnormality for the current image is increased. In an example where the threshold value changes, for example, when it is determined that the predicted result of an abnormality corresponding to the previous unit represents the absence of an abnormality, the current threshold value may be set to have an upward trend. Here, the upward trend may include keeping the current threshold value the same as the previous threshold value or setting the current threshold value higher than the previous threshold value. Setting the current threshold value high may mean that the possibility of determining an abnormality for the current image is decreased. As described above, the technique according to an embodiment of the present disclosure can improve the accuracy of abnormality determination by determining the predicted result of an abnormality of the current image so as to have a similar (corresponding) trend to the predicted result of an abnormality of the previous image.
[0286] In one embodiment, the threshold assigned to the current unit of the first queue 1010 may be determined based on values included in a plurality of previous units of the second queue 1030. In one embodiment, the threshold assigned to a particular unit of the first queue 1010 may be variable depending on the values of previous units of other queues other than the first queue 1010. For example, the threshold assigned to the current unit of the first queue 1010 may be determined based on a ratio value obtained from the values of previous units of the second queue 1030, or on a comparison of the average value with a particular threshold. Here, the particular threshold is a threshold for determining the threshold assigned to the current unit, and may mean a threshold in the form of a ratio, such as 30%, 40%, 50%, and 80%. For example, the threshold of the second unit 1010b of the first queue 1010 may be determined based on the values of the first unit 1030a of the second queue 1030 and the previous units of the first unit 1030a. If it is determined that the ratio of values of 1 (i.e., the ratio judged to be abnormal) among the values of the first unit 1030a and the previous units of the first unit 1030a of the second queue 1030 does not exceed a certain threshold, then the threshold for the second unit 1010b of the first queue 1010 can be kept at 0.5. The threshold for the third unit 1010c of the first queue 1010 may be determined based on the values of the first unit 1030a and the second unit 1030b of the second queue 1030. Alternatively, the threshold for the third unit 1010c of the first queue 1010 may be determined based on the values of the first unit 1030a, the second unit 1030b, and the previous units of the first unit 1030a of the second queue 1030. For example, if it is determined that the ratio of values of 1 among the values of the first unit 1030a of the second queue 1030 and the previous units of the first unit 1030a (i.e., the ratio judged to be abnormal) exceeds a certain threshold, the threshold for the third unit 1010c of the first queue 1010 may be set to 0.4.For example, if it is determined that the ratio of values of 1 (i.e., the ratio judged to be abnormal) among the values of the first unit 1030a and the previous units of the first unit 1030a of the second queue 1030 does not exceed a certain threshold, then the threshold for the third unit 1010c of the first queue 1010 may be set to 0.5. As another example, the threshold assigned to the fifth unit 1010e of the first queue 1010 may be determined based on a comparison between the values representing abnormality among the values of the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 and a predetermined threshold. For example, if it is determined that the ratio of values resulting in the presence of abnormality (i.e., the value of 1) among the values of the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 is 40% or more, then the threshold assigned to the fifth unit 1010e of the first queue 1010 may be set to 0.4. For example, if it is determined that the proportion of values indicating an anomaly (i.e., a value of 1) among the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 is less than 40%, the threshold assigned to the fifth unit 1010e of the first queue 1010 may be set to 0.5.
[0287] In one embodiment, the threshold value assigned to the current unit (e.g., the Nth unit) of the first queue 1010 may be variable according to the values included in the previous units (e.g., the N-1th unit, the N-2th unit, the N-3th unit, and the N-4th unit) of the current unit of the second queue 1030. In such an example, the threshold value assigned to the fifth unit 1010e of the first queue 1010 may be determined based on the mainstream value or representative value of the values of the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030. The mainstream value or representative value may be determined as the value with the highest proportion among the values included in a predetermined number of units. For example, among the values of 1 or 0 included in five sequential units, the mainstream value or representative value may be determined. In such an example, since four of the units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 have a value of 1 and one unit has a value of 0, the mainstream value or representative value of the five units of the second queue 1030 may be determined as 1. When the mainstream value or representative value of the values of the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 is 1, the threshold value assigned to the fifth unit 1010e of the first queue 1010 may be set to 0.4. In such an example, when the mainstream value or representative value of the values of the previous units 1030a, 1030b, 1030c, and 1030d of the second queue 1030 is 0, the threshold value assigned to the fifth unit 1010e of the first queue 1010 may be set to 0.5. As shown in FIG. 10, the threshold value assigned to the sixth unit 1010f of the first queue 1010 may be determined based on the mainstream value, representative value, or ratio of the values of the previous units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030.
[0288] In this specification, for the sake of convenience of explanation, the corresponding content is omitted, but it is obvious to those skilled in the art that according to the characteristics of the queue in which the oldest unit is excluded when a new unit is added to the queue, the values of the previous units of the first unit 1010a of the first queue 1010 can also be considered.
[0289] In one embodiment, the threshold assigned to the current unit of the first queue 1010 may be determined based on a value included in at least one previous unit of the third queue 1050. The threshold assigned to the fifth unit 1010e of the first queue 1010 may be determined based on the value of a previous unit of the third queue 1050 (e.g., the fourth unit 1050d). For example, if the value of the fourth unit 1050d of the third queue 1050 is 1, the threshold for the fifth unit 1010e of the first queue 1010 may be determined to be 0.4. Alternatively, the threshold may be kept the same as the previous threshold, or the threshold may be decreased compared to the previous threshold. For example, if the value of the first unit 1050a of the third queue 1050 is 0, the threshold for the second unit 1010b of the first queue 1010 may be determined to be 0.5. Alternatively, the threshold may be kept the same as the previous threshold, or the threshold may be increased compared to the previous threshold.
[0290] In one embodiment, referring to the example in Figure 10, the threshold corresponding to the next unit of the first queue 1010 may be determined according to the mainstream values (for example, values having three or more values of 0 and 1) of the five units of the second queue 1030. In the example in Figure 10, the value of the first unit 1030a of the second queue 1030 may be determined to be 1. Assume that the four preceding units of the first unit 1030a of the second queue 1030 have values of 0, 0, 0, and 0, respectively. Under such assumptions, the mainstream value of the first unit 1030a and the four preceding units of the second queue 1030 may be determined to be 0. This may result in the threshold corresponding to the second unit 1010b of the first queue 1010 being determined to be 0.5. Next, it may be determined that the second unit 1030b of the second queue 1030 has a value of 1, and the mainstream value of the second unit 1030b and the four preceding units of the second queue 1030 may be determined to be 0. As a result, the threshold corresponding to the third unit 1010c of the first queue 1010 may be determined to be 0.5. Next, the third unit 1030c of the second queue 1030 may be determined to have a value of 1, and the mainstream value of the third unit 1030c of the second queue 1030 and the previous four units may be determined to be 1. As a result, the threshold corresponding to the fourth unit 1010d of the first queue 1010 may be determined to be 0.4. Next, the fourth unit 1030d of the second queue 1030 may be determined to have a value of 0, and the mainstream value of the fourth unit 1030d of the second queue 1030 and the previous four units may be determined to be 1. As a result, the threshold corresponding to the fifth unit 1010e of the first queue 1010 may be determined to be 0.4. Next, it is determined that the fifth unit 1030e of the second queue 1030 has a value of 1, and the mainstream value of the fifth unit 1030e of the second queue 1030 and the previous four units (i.e., 1030a, 1030b, 1030c, and 1030d) may be determined to be 1. As a result, the threshold corresponding to the sixth unit 1010f of the first queue 1010 may be determined to be 0.4.
[0291] In one embodiment, the values of units 1050a, 1050b, 1050c, 1050d, and 1050e of the third queue 1050 may be determined based on the results of a first voting that utilizes the values of units 1030a, 1030b, 1030c, 1030d, and 1030e contained in the second queue 1030. The values contained in the units of the third queue 1050 may be determined based on the mainstream, representative, or ratio of the primary predicted results of anomalies. For example, the fifth unit 1150e of the third queue 1050 may be determined based on the representative or mainstream values of units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030. In such an example, the fifth unit 1050e of the third queue 1050 may be set to 1, which is the mainstream or representative value of units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030. The number of units used to determine the mainstream or representative value may vary, not just five, but also three or four, for example. In another example, the fifth unit 1050e of the third queue 1050 may be determined based on a comparison of the ratio of units with a value of 1 from the values of units 1030a, 1030b, 1030c, 1030d, and 1030e of the second queue 1030 with a threshold ratio.
[0292] In one embodiment, the fourth queue 1070 may correspond to a queue for generating alarms. In one embodiment, the values of units 1070a, 1070b, 1070c, 1070d, and 1070e of the fourth queue 1070 may be determined based on a second voting that utilizes the values of the units of the third queue 1050. The fourth queue 1070 may decide whether or not to generate an alarm by considering the continuity of the unit values of the third queue 1050 via the second voting. For example, assume a criterion of continuity of unit values of 3. Under such assumption, since neither the first unit 1050a of the third queue 1050 nor the two preceding units have a value of 1, the first unit 1070a of the fourth queue 1070 may be set to 0 (i.e., alarm off). Since neither the second unit 1050b of the third queue 1050 nor the two preceding units (1050a and the previous unit) have a value of 1, the second unit 1070b of the fourth queue 1070 may be set to 0 (i.e., alarm off). Since neither the third unit 1050c of the third queue 1050 nor the two preceding units 1050a and 1050b have a value of 1, the third unit 1070c of the fourth queue 1070 may be set to 0 (i.e., alarm off). Since neither the fourth unit 1050d of the third queue 1050 nor the two preceding units 1050b and 1050c have a value of 1, the fourth unit 1070c of the fourth queue 1070 may be set to 0 (i.e., alarm off). Since the fifth unit 1050e of the third queue 1050 and the two preceding units 1050c and 1050d all have a value of 1, the fifth unit 1070e of the fourth queue 1070 may be set to 1 (i.e., alarm on).
[0293] As described above, the technology according to one embodiment of the present disclosure may utilize multiple queues to determine the presence or absence of an anomaly or to generate an alarm corresponding to an anomaly. Values contained in other queues (e.g., previous queues) among these multiple queues may be used to determine the values contained in a particular queue or the threshold to be assigned to a particular queue.
[0294] Figure 11 illustrates an example of a method for determining an abnormal driver alarm according to one embodiment of the present disclosure.
[0295] In the examples shown in Figure 11, the previously mentioned examples will be replaced with the previously mentioned explanations to avoid repetition of explanations.
[0296] The first queue 1110 and the units 1110a, 1110b, 1110c, 1110d, 1110e, and 1110f included in the first queue in Figure 11 may correspond to the first queue 1010 and the units 1010a, 1010b, 1010c, 1010d, 1010e, and 1010f included in the first queue in Figure 10, respectively.
[0297] The second queue 1130 and the units 1130a, 1130b, 1130c, 1130d, and 1130e included in the second queue in Figure 11 may correspond to the second queue 1030 and the units 1030a, 1030b, 1030c, 1030d, and 1030e included in the second queue in Figure 10, respectively.
[0298] The third queue 1150 in Figure 11 may include a first unit 1150a, a second unit 1150b, a third unit 1150c, a fourth unit 1150d, and a fifth unit 1150e. In one embodiment, the values of units 1150a, 1150b, 1150c, 1150d, and 1150e of the third queue 1150 may be determined based on the results of a first voting that utilizes the values of units 1130a, 1130b, 1130c, 1130d, and 1130e contained in the second queue 1130. The values included in the units of the third queue 1150 may be determined based on the mainstream value, representative value, or ratio of the primary predicted result of anomalies.
[0299] In one embodiment, an example is shown in which the threshold of the first queue 1110 in Figure 11 is determined based on the value of at least one previous unit of the second queue 1130. Referring to the example in Figure 11, the threshold corresponding to the next unit of the first queue 1110 may be determined by the mainstream value of the five units of the second queue 1130 (for example, a value having three or more values of 0 and 1). In the example in Figure 11, the value of the first unit 1130a of the second queue 1130 may be determined to be 1. Assume that the four units preceding the first unit 1130a of the second queue 1130 have values of 0, 0, 0, and 0, respectively. Under such assumption, the mainstream value of the first unit 1130a and the four preceding units of the second queue 1130 may be determined to be 0. Thus, the threshold corresponding to the second unit 1110b of the first queue 1110 may be determined to be 0.5. Next, it is determined that the second unit 1130b of the second queue 1130 has a value of 1, and the mainstream value of the second unit 1130b of the second queue 1130 and the previous four units may be determined to be 0. This may result in the threshold corresponding to the third unit 1110c of the first queue 1110 being determined to be 0.5. Next, it is determined that the third unit 1130c of the second queue 1130 has a value of 1, and the mainstream value of the third unit 1130c of the second queue 1130 and the previous four units may be determined to be 1. This may result in the threshold corresponding to the fourth unit 1110d of the first queue 1110 being determined to be 0.4. Next, it is determined that the fourth unit 1130d of the second queue 1130 has a value of 0, and the mainstream value of the fourth unit 1130d of the second queue 1130 and the previous four units may be determined to be 1. This may result in the threshold corresponding to the fifth unit 1110e of the first queue 1110 being determined to be 0.4. Next, it is determined that the fifth unit 1130e of the second queue 1130 has a value of 1, and the mainstream value of the fifth unit 1130e of the second queue 1130 and the previous four units (i.e., 1130a, 1130b, 1130c, and 1130d) may be determined to be 1. This may result in the threshold corresponding to the sixth unit 1110f of the first queue 1110 being determined to be 0.4.
[0300] Figure 11 illustrates an embodiment in which, by performing a first voting, the predicted results for secondary anomalies are generated in a form that utilizes a counter or counter value. The example in Figure 11 shows an embodiment in which, as a result of the first voting, the counter value increases as the mainstream value or representative value becomes 1, and maintains its value as a result of the first voting as the mainstream value or representative value becomes 0. However, embodiments in which the counter value decreases as the mainstream value or representative value becomes 0, depending on the mode of implementation, may also be included within the scope of the rights of this disclosure.
[0301] In one embodiment, when a first voting is performed using multiple units, including the first unit 1130a of the second queue 1130, the second unit 1150b of the third queue 1150 may have a value of 0. As a result of the first voting being performed using multiple units, including the first unit 1130a and the second unit 1130b of the second queue 1130, the third unit 1150c of the third queue 1150 may have a value of 1 by adding a counter value of 1. As a result of the first voting being performed using multiple units, including the first unit 1130a, the second unit 1130b, and the third unit 1130c of the second queue 1130, the fourth unit 1150d of the third queue 1150 may have a value of 2 by adding a counter value of 1. As a result of the first voting being performed using multiple units including the first unit 1130a, second unit 1130b, third unit 1130c, and fourth unit 1130d of the second queue 1130, the fifth unit 1150e of the third queue 1150 may have a value of 3, with the counter value of 1 being added. In this way, when generating the secondary prediction result of anomalies, the counter value may not be added if the result of the first voting is 0, and the counter value may be added if the result of the first voting is 1.
[0302] In one embodiment, the fourth queue 1170 may correspond to a queue for generating alarms. In one embodiment, the values of units 1170a, 1170b, 1170c, 1170d, and 1170e of the fourth queue 1170 may be determined based on a second voting that utilizes the values of the units of the third queue 1150.
[0303] In one embodiment, the second voting may include comparing the value of each unit in the third queue 1150 (e.g., a counter value) with one or more thresholds. For example, if the value of a unit in the third queue 1150 is greater than or equal to a threshold, the second voting may decide to generate an alarm (ON) for the corresponding unit in the fourth queue 1170. For example, if the value of a unit in the third queue 1150 is less than a threshold, the second voting may decide not to generate an alarm (OFF) for the corresponding unit in the fourth queue 1170. In the example in Figure 11, since the threshold is set to 3, the computing device 100 may decide to generate an abnormal alarm for the fifth unit 1170e of the fourth queue 1170, which corresponds to the fifth unit 1150e of the third queue 1150 having a value of 3.
[0304] In one embodiment, the counter value may have a predetermined range. When the counter value reaches a boundary value within the predetermined range, the computing device 100 may maintain or change the counter value in a different manner than when it has not reached the boundary value. For example, suppose the range of the counter value is 0 to 5, and the threshold is 3. Under such assumptions, the threshold of 3 and each unit of the third queue 1150 can be compared. If the first unit of the third queue 1150 has a value of 3, the computing device 100 may set the corresponding first unit of the fourth queue 1170 to ON. If the second unit, which is the next unit of the third queue 1150, has a value of 4, the computing device 100 may set the corresponding unit of the fourth queue 1170 to ON. If the third unit, which is the next unit of the third queue 1150, has a value of 5, the computing device 100 may set the corresponding third unit of the fourth queue 1170 to ON. In this situation, if it is determined through the first voting that the fourth unit, which is the next unit in the third queue 1150, will have a value of 1 added to it, the computing device 100 may set the fourth unit of the third queue 1150 to 5 (i.e., not add 1). As another example, in the same situation, if the value of 0 is determined through the first voting to be the dominant or representative value for the fourth unit, which is the next unit in the third queue 1150, the computing device 100 may set the fourth unit of the third queue 1150 to 4 (i.e., subtract 1). In this way, when the secondary prediction result of an anomaly reaches a critical range, the counter value can no longer increase, decrease, or be maintained. In such a situation, if the counter value changes from 3 to 2 as it decreases, the computing device 100 may decide to turn off the anomaly alarm. As described above, the technology according to one embodiment of the present disclosure can set a critical range for the counter value and maintain or increase the counter value when the counter value reaches the minimum value of the critical range.The technology according to one embodiment of the present disclosure can set a critical range for a counter value and maintain or decrease the counter value when it reaches the maximum value of the critical range. The technology according to one embodiment of the present disclosure can set a critical range for a counter value and replace the counter value with the minimum or maximum value when it falls outside the minimum or maximum value of the critical range.
[0305] In one embodiment, the type or intensity of the alarm may be set to differ depending on the difference between the counter value and the threshold. If the counter value is 4 and the threshold is 3, the alarm may be set to the first type or first intensity. If the counter value is 5 and the threshold is 3, the alarm may be set to the second type or second intensity. Here, the second intensity may be greater than the first intensity. Here, the second type of alarm can convey a more intuitive and powerful message to the user than the first type of alarm.
[0306] In one embodiment, the technology according to one embodiment of the present disclosure may utilize one counter or multiple counters. Furthermore, multiple thresholds to be compared with the counters may be used. Alternatively, multiple counters and multiple thresholds may be used.
[0307] An example of using a single counter can be described as follows: In a situation where there is one counter for N alarms, the counter may be incremented, decremented, or maintained by 1 according to the predicted result (predicted value) of the secondary anomaly. A starting value and / or critical range (min and max values) of the counter may be set, and N thresholds between the min and max values for triggering alarms may be set, where N is a natural number. The counter is compared with each of the N thresholds, and an alarm is triggered if the counter is greater than or equal to a specific threshold, and in the case of a combined condition (when the counter is greater than or equal to multiple thresholds), an alarm corresponding to the higher threshold may be triggered. For example, suppose the critical range of the counter is min=0 and max=6, and the first threshold=3 and the second threshold=5. In this situation, when the secondary predicted result of the anomaly in a particular image is 1, the counter has a value of 1, and no alarm is triggered. When the secondary predicted result of the anomaly in the next image is 1, the counter has a value of 2, and no alarm is triggered. When the secondary predicted result of the anomaly in the image after that is 1, the counter has a value of 3, and the first alarm corresponding to the first threshold may be triggered. If the secondary prediction result for an anomaly in the next image is 1, the counter will have a value of 4, and the first alarm corresponding to the first threshold can be generated. If the secondary prediction result for an anomaly in the next image is 1, the counter will have a value of 5, and the second alarm corresponding to the second threshold can be generated. If the secondary prediction result for an anomaly in the next image is 1, the counter will still have a value of 5, and the second alarm corresponding to the second threshold can be generated. If the secondary prediction result for an anomaly in the next image is 0, the counter will have a value of 4, and the first alarm corresponding to the first threshold can be generated. If the secondary prediction result for an anomaly in the next image is 0, the counter will have a value of 3, and the first alarm corresponding to the first threshold can be generated. If the secondary prediction result for an anomaly in the image after that is 0, the counter will have a value of 32, and the alarm can be turned OFF.
[0308] An example of using multiple counters can be explained as follows: In a situation where there are M counters for N alarms, each counter can be increased, decreased, or maintained by 1 according to the predicted result (predicted value) of a secondary anomaly, where N and M are natural numbers. A starting value and / or critical range (min and max values) may be set for each counter, and M thresholds between the min and max values for triggering an alarm may be set for each counter. Each of the N counters is compared with each of the M thresholds, and an alarm can be triggered if a particular counter is greater than or equal to a specific threshold. For example, the critical range of the first counter may be min=0 and max=4, and the first counter may be assigned a first threshold value of 3. Alternatively, the critical range of the second counter may be min=0 and max=6, and the second counter may be assigned a second threshold value of 5. In such an exemplary situation, when the predicted secondary anomaly result for a particular image is 0, both the first and second counters can have a value of 0, and no alarm is triggered. In the following image, when the secondary prediction result for an anomaly is 1, both the first and second counters can have a value of 1, and no alarm is generated. In the following image, when the secondary prediction result for an anomaly is 1, both the first and second counters can have a value of 2, and no alarm is generated. In the image after that, when the secondary prediction result for an anomaly is 1, both the first and second counters can have a value of 3, and since the first counter is determined to be above the first threshold, the first alarm or primary alarm can be generated. In the following image, when the secondary prediction result for an anomaly is 1, both the first and second counters can have a value of 4, and since the first counter is determined to be above the first threshold, the first alarm or primary alarm can persist. In the image after that, when the secondary prediction result for an anomaly is 1, the first counter is maintained at its maximum value of 4, the second counter can have a value of 5, the first counter may be determined to be above the first threshold, and the second counter may be determined to be above the second threshold. In this case, the second alarm or secondary alarm can be generated.When the secondary prediction result of an anomaly in the next image is 0, the first counter can decrease to a value of 3, and the second counter can decrease to a value of 4. Since the first counter is determined to be greater than or equal to the first threshold, a first alarm or a primary alarm can occur. When the secondary prediction result of an anomaly in the next image is 0, the first counter decreases to a value of 2, and the second counter is decreased to a value of 3, whereby the alarm can be turned off.
[0309] FIG. 12 exemplarily shows a method for determining a driver anomaly alarm according to an embodiment of the present disclosure.
[0310] Among the examples shown in FIG. 12, the foregoing examples will be replaced with the foregoing description to prevent duplication of explanation.
[0311] The first queue 1210 in FIG. 12 and the units 1210a, 1210b, 1210c, 1210d, 1210e, and 1210f included in the first queue may respectively correspond to the first queue 1010 in FIG. 10 and the units 1010a, 1010b, 1010c, 1010d, 1010e, and 1010f included in the first queue.
[0312] The second queue 1230 in FIG. 12 and the units 1230a, 1230b, 1230c, 1230d, and 1230e included in the second queue may respectively correspond to the second queue 1030 in FIG. 10 and the units 1030a, 1030b, 1030c, 1030d, and 1030e included in the second queue.
[0313] In one embodiment, an example is shown in which the threshold of the first queue 1210 in Figure 12 is determined based on the value of at least one previous unit of the second queue 1230. Referring to the example in Figure 12, the threshold corresponding to the next unit of the first queue 1210 may be determined by the mainstream value of the five units of the second queue 1230 (for example, a value having three or more values of 0 and 1). In the example in Figure 12, the value of the first unit 1230a of the second queue 1230 may be determined to be 1. Assume that the four previous units of the first unit 1230a of the second queue 1230 have values of 0, 0, 0, and 0, respectively. Under such assumption, the mainstream value of the first unit 1230a and the four previous units of the second queue 1230 may be determined to be 0. Thus, the threshold corresponding to the second unit 1210b of the first queue 1210 may be determined to be 0.5. Next, it is determined that the second unit 1230b of the second queue 1230 has a value of 1, and the mainstream value of the second unit 1230b of the second queue 1230 and the previous four units may be determined to be 0. This means that the threshold corresponding to the third unit 1210c of the first queue 1210 may be determined to be 0.5. Next, it is determined that the third unit 1230c of the second queue 1230 has a value of 1, and the mainstream value of the third unit 1230c of the second queue 1230 and the previous four units may be determined to be 1. This means that the threshold corresponding to the fourth unit 1210d of the first queue 1210 may be determined to be 0.4. Next, it is determined that the fourth unit 1230d of the second queue 1230 has a value of 0, and the mainstream value of the fourth unit 1230d of the second queue 1230 and the previous four units may be determined to be 1. This means that the threshold corresponding to the fifth unit 1210e of the first queue 1210 may be determined to be 0.4. Next, it is determined that the fifth unit 1230e of the second queue 1230 has a value of 1, and the mainstream value of the fifth unit 1230e of the second queue 1230 and the previous four units (i.e., 1230a, 1230b, 1230c, and 1230d) may be determined to be 1. As a result, the threshold corresponding to the sixth unit 1210f of the first queue 1210 may be determined to be 0.4.
[0314] In one embodiment, the values of units 1250a, 1250b, 1250c, 1250d, and 1250e of the third queue 1250 may be determined based on the results of a first voting that utilizes the values of units 1230a, 1230b, 1230c, 1230d, and 1230e contained in the second queue 1230. The values contained in the units of the third queue 1250 may be determined based on the mainstream value, representative value, or ratio of the primary predicted result of anomalies.
[0315] In Figure 12, the third queue 1250 and the first unit 1250a, second unit 1250b, third unit 1250c, fourth unit 1250d, and fifth unit 1250e included in the third queue 1250 may correspond to the first unit 1150a, second unit 1150b, third unit 1150c, fourth unit 1150d, and fifth unit 1150e of the third queue 1150, respectively. That is, Figure 12 shows a boating queue to which a counter is applied. Figure 12 illustrates a method in which the unit of increase or decrease of the counter value is not 1. Figure 12 shows an embodiment in which the difference value at the acquisition time between images is used as the unit of increase or decrease of the counter value.
[0316] In one embodiment, the difference between the acquisition time of the first image corresponding to the first unit 1250a of the third queue 1250 and the acquisition time of the second image corresponding to the second unit 1250b may be determined to be 500ms. This allows the unit of increase or decrease of the counter value between the first unit 1250a and the second unit 1250b to be set to 500ms. The difference between the acquisition time of the second image corresponding to the second unit 1250b of the third queue 1250 and the acquisition time of the third image corresponding to the third unit 1250c may be determined to be 700ms. This allows the unit of increase or decrease of the counter value between the second unit 1250b and the third unit 1250c to be set to 700ms. The difference between the acquisition time of the third image corresponding to the third unit 1250c of the third queue 1250 and the acquisition time of the fourth image corresponding to the fourth unit 1250d may be determined to be 300ms. This allows the unit of increase or decrease of the counter value between the third unit 1250c and the fourth unit 1250d to be set to 300ms. The difference between the acquisition time of the fourth image corresponding to the fourth unit 1250d of the third queue 1250 and the acquisition time of the fifth image corresponding to the fifth unit 1250e may be determined to be 500ms. This allows the unit of increase or decrease in the counter value between the fourth unit 1250d and the fifth unit 1250e to be set to 500ms.
[0317] In one embodiment, the fourth queue 1270 may correspond to a queue for generating alarms. In one embodiment, the values of units 1270a, 1270b, 1270c, 1270d, and 1270e of the fourth queue 1270 may be determined based on a second voting that utilizes the values of the units of the third queue 1250.
[0318] As described above, the technology according to one embodiment of the present disclosure may utilize multiple queues and multiple voting to generate abnormal alarms and / or to determine the type and intensity of abnormal alarms. Furthermore, the technology according to one embodiment of the present disclosure may utilize one or more counters and one or more thresholds to determine user-friendly and more accurate alarms.
[0319] Figure 12 illustrates an embodiment in which the threshold value is 1.1. This may result in an abnormal alarm being turned ON for the fifth unit 1270e of the fourth queue 1270.
[0320] Figure 13 illustrates a methodology for determining a driver distraction alarm according to one embodiment of the present disclosure.
[0321] In one embodiment, the first queue 1310 may include a first unit 1310a, a second unit 1310b, a third unit 1310c, a fourth unit 1310d, and a fifth unit 1310e. The first queue 1310 may include values that quantitatively represent the detection result of attention distraction in the target image, the attention distraction score, and / or the classification result of attention distraction. The corresponding units in the first queue 1310, the second queue 1330, the third queue 1350, and the fourth queue 1370 may include judgment results for the same image acquired at the same time. For example, the first unit 1310a of the first queue 1310, the first unit 1330a of the second queue 1330, the first unit 1350a of the third queue 1350, and the first unit 1370 of the fourth queue 1370 may have values corresponding to the first image acquired at the first time. For example, the first image acquired at the first time point may be processed in the following order: first unit 1310a of the first queue 1310, second unit 1330a of the second queue 1330, first unit 1350a of the third queue 1350, and first unit 1370a of the fourth queue 1370. As a result of this processing, the fourth queue 1370, which is the last queue, may contain the result of turning the attention distraction alarm ON or OFF for each image.
[0322] In one embodiment, the first queue 1310 may represent the attention distraction score in the image (or ROI) obtained from the model. As a non-restrictive example, the first queue 1310 may represent the maximum score related to the detection of attention distraction in the image obtained from the model. For example, the maximum score may mean the result with the highest score among the processing results for a particular image generated by the model. The higher the attention distraction score, the more likely it is that attention distraction is present.
[0323] Each of the units 1310a, 1310b, 1310c, 1310d, and 1310e of the first queue 1310 can be compared to a corresponding threshold. If the value of a unit exceeds the threshold, the corresponding unit of the second queue 1330 may be assigned a value of 1. If the value of a unit does not exceed the threshold, the corresponding unit of the second queue 1330 may be assigned a value of 0. In the example in Figure 13, the value of the first unit 1310a of the first queue 1310 is 1.1 and the threshold is 1.0, so the first unit 1330a of the second queue 1330 corresponding to the first unit 1310a of the first queue 1310 may have a value of 1, indicating the presence of attentional distraction. Thus, each of the units 1330a, 1330b, 1330c, 1330d, and 1330e of the second queue 1330 may contain a primary prediction result of attentional distraction.
[0324] The threshold corresponding to the sixth unit 1310f of the first queue 1330 may be determined using at least one of the units 1330a, 1330b, 1330c, 1330d, and 1330e of the second queue 1330. Thus, the thresholds assigned to each of the units 1310a, 1310b, 1310c, 1310d, and 1310e of the first queue 1330 may be determined using at least a portion of the unit values 1330a, 1330b, 1330c, 1330d, and 1330e of the second queue 1330.
[0325] For example, the threshold of the sixth unit 1310f of the first queue 1310 may be variable depending on the value of the fifth unit 1330e of the second queue 1330. In such an example, if the fifth unit 1330e is 1, the threshold of the sixth unit 1310f may be decreased or maintained at a minimum value compared to the previous threshold of the fifth unit 1310e. In such an example, if the fifth unit 1330e is 0, the threshold of the sixth unit 1310f may be increased or maintained at a maximum value compared to the previous threshold of the fifth unit 1310e.
[0326] For example, the threshold of the sixth unit 1310f of the first queue 1310 may be determined based on a comparison of the values of several previous units 1330a, 1330b, 1330c, 1330d, and 1330e of the second queue 1330 with a critical ratio. If each of the several previous units 1330a, 1330b, 1330c, 1330d, and 1330e has values of 1, 1, 1, 0, and 1, and the critical ratio to the value of 1 is 60%, then the threshold of the sixth unit 1310f of the first queue 1310 can be reduced or kept at a minimum value compared to the previous threshold of the fifth unit 1310e, because the proportion of 1 in the several previous units 1330a, 1330b, 1330c, 1330d, and 1330e exceeds the critical ratio.
[0327] For example, the threshold of the sixth unit 1310f of the first queue 1310 may be variable depending on what the mainstream or representative values of several previous units 1330a, 1330b, 1330c, 1330d, and 1330e of the second queue 1330 are. In the example in Figure 13, since the mainstream or representative value of several previous units 1330a, 1330b, 1330c, 1330d, and 1330e is 1, the threshold of the sixth unit 1310f of the first queue 1310 can be reduced or kept at a minimum value compared to the previous threshold of the fifth unit 1310e.
[0328] For example, the threshold of the sixth unit 1310f of the first queue 1310 may be variable based on the value assigned to the previous unit of the third queue 1350.
[0329] In one embodiment, Figure 13 shows an example in which the threshold of the first queue 1310 is determined based on the value of at least one previous unit of the second queue 1330. Referring to the example in Figure 13, the threshold corresponding to the next unit of the first queue 1310 may be determined by the mainstream value (for example, a value having three or more values of 0 and 1) of the values of the five units of the second queue 1330. In such an example, if the value of 1 is determined to be the mainstream value, the threshold for the next unit of the first queue 1310 is determined to be 0.9, otherwise the threshold for the next unit of the first queue 1310 may be determined to be 1.0. In the example in Figure 13, the value of the first unit 1330a of the second queue 1330 may be determined to be 1. Assume that the four previous units of the first unit 1330a of the second queue 1330 have values of 0, 0, 0, and 0, respectively. Under such assumption, the mainstream value of the first unit 1330a of the second queue 1330 and the four previous units may be determined to be 0. As a result, the threshold corresponding to the second unit 1310b of the first queue 1310 may be determined to be 1.0. Next, the second unit 1330b of the second queue 1330 is determined to have a value of 1, and the mainstream value of the second unit 1330b of the second queue 1330 and the previous four units may be determined to be 0. As a result, the threshold corresponding to the third unit 1310c of the first queue 1310 may be determined to be 1.0. Next, the third unit 1330c of the second queue 1330 is determined to have a value of 1, and the mainstream value of the third unit 1330c of the second queue 1330 and the previous four units may be determined to be 1. As a result, the threshold corresponding to the fourth unit 1310d of the first queue 1310 may be determined to be 0.9. Next, the fourth unit 1330d of the second queue 1330 is determined to have a value of 0, and the mainstream value of the fourth unit 1330d of the second queue 1330 and the previous four units may be determined to be 1. As a result, the threshold corresponding to the fifth unit 1310e of the first queue 1310 may be determined to be 0.9. Next, it is determined that the fifth unit 1330e of the second queue 1330 has a value of 1, and the mainstream value of the fifth unit 1330e of the second queue 1330 and the previous four units (i.e., 1330a, 1330b, 1330c, and 1330d) may be determined to be 1.As a result, the threshold corresponding to the sixth unit 1310f of the first queue 1310 may be determined to be 0.9.
[0330] In one embodiment, the third queue 1350 may include secondary distraction prediction results. The values of units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350 may be determined using first voting. The secondary distraction prediction results may be determined based on the mainstream, representative, or ratio values of the primary distraction prediction results for multiple units of the second queue 1330. For example, if the mainstream, representative, or ratio value of the primary distraction prediction results for multiple units of the second queue 1330 is determined to be 1, the corresponding unit of the third queue 1330 may have a value of 1 added to it, or the value of that unit may be maintained at its maximum value. For example, if the mainstream, representative, or ratio value of the primary distraction prediction results for multiple units of the second queue 1330 is determined to be 0, the corresponding unit of the third queue 1330 may have a value of 1 subtracted from it, or the value of that unit may be maintained at its minimum value.
[0331] Thus, in the example in Figure 13, the secondary expectation result of attention distraction can be represented by the value of the counter. As shown in the example in Figure 13, the first unit 1350a and the second unit 1350b of the third queue 1350 may have a value of 0. To determine the value of the third unit 1350c of the third queue 1350, the first unit 1330a, the second unit 1330b, and the third unit 1330c of the second queue 1330 may be used. Alternatively, to determine the value of the third unit 1350c of the third queue 1350, the previously obtained values of the first unit 1330a other than the first unit 1330a, the second unit 1330b, and the third unit 1330c of the second queue 1330 may be used. For example, in Figure 13, the values of the first unit 1330a, the second unit 1330b, and the third unit 1330c (or any additional previous unit) have been determined to have a mainstream value of 1, so a value of 1 may be added to the third unit 1350c of the third queue 1350. Similarly, a value of 1 may be added to the fourth unit 1350d of the third queue 1350, so that the fourth unit 1350d has a value of 2. Likewise, a value of 1 may be added to the fifth unit 1350e of the third queue 1350, so that the fifth unit 1350e has a value of 3.
[0332] In additional embodiments, methodologies for representing the secondary prediction result of attention distraction as 0 or 1 indicating whether or not attention distraction is present may also be included within the scope of this disclosure. In such embodiments, the third unit 1350c of the third queue 1350 may be assigned a value of 1, the fourth unit 1350d of the third queue 1350 may be assigned a value of 1, and the fifth unit 1350e of the third queue 1350 may be assigned a value of 1.
[0333] In one embodiment, the fourth queue 1370 may correspond to a queue for determining a distraction alarm. Each of the units 1370a, 1370b, 1370c, 1370d, and 1370e of the fourth queue 1370 may have values relating to whether or not to generate an alarm, the type of alarm, and / or the intensity of the alarm. As shown in the example in Figure 13, the values of the units 1370a, 1370b, 1370c, 1370d, and 1370e of the fourth queue 1370 may be determined using a second voting. The second voting may include comparing the values of the units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350 with alarm thresholds. If the values of units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350 are greater than or equal to the alarm threshold, the values of the corresponding units of the fourth queue 1370 may be determined to be ON. If the values of units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350 are less than the alarm threshold, the values of the corresponding units of the fourth queue 1370 may be determined to be OFF. In the example in Figure 13, it is assumed that the alarm threshold is 2, which allows the fourth unit 1370d and the fifth unit 1370e of the fourth queue 1370 to be assigned the ON value for the attention distraction alarm.
[0334] In an additional embodiment, the values of units 1370a, 1370b, 1370c, 1370d, and 1370e of the fourth queue 1370 may be determined based on whether or not there is continuity among at least some of the units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350. In such an embodiment, assuming that the continuity threshold is 3, the third unit 1350c, the fourth unit 1350d, and the fifth unit 1350e of the third queue 1350 all have a value of 1, so the fifth unit 1370e of the fourth queue 1370 may be assigned the ON value for the distraction alarm, and the fourth unit 1370d may be assigned the OFF value for the distraction alarm.
[0335] Embodiments in which one or more counters and / or thresholds are used to determine an attention-distraction alarm, as shown in Figure 13, also fall within the scope of the rights of this disclosure.
[0336] As described above, the technology according to one embodiment of the Disclosure can utilize multiple queues, multiple voting, counter values, continuity values, and / or variable thresholds to determine the presence or absence of attention distraction or to generate an alarm corresponding to attention distraction. Interoperability between these queues can be achieved by configuring the system so that values contained in a particular queue or thresholds assigned to a particular queue are determined using values contained in other queues (e.g., previous queues). In this way, the technology according to one embodiment of the Disclosure can provide more accurate attention distraction alarms by adding multiple decisions regarding alarm generation to maximize the user experience.
[0337] Figure 14 is a schematic diagram of the computing environment of a computing device 100 according to one embodiment of the present disclosure.
[0338] In this disclosure, computing devices, computing apparatus, computers, systems, components, modules, or units include routines, procedures, programs, components, data structures, etc., that perform a particular task or realize a particular type of abstract data. Furthermore, a person skilled in the art will readily recognize that the methods presented in this disclosure can be implemented in other computer system configurations, including single-processor or multi-processor computing devices, minicomputers, mainframe computers, as well as personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, and others (each of which may operate in conjunction with one or more related devices).
[0339] The embodiments described herein can also be implemented in a distributed computing environment in which a task is performed by remote processing units connected via a communication network. In a distributed computing environment, program modules can reside in both local and remote memory storage devices.
[0340] Computing devices typically include various computer-readable media. Any media accessible by a computer can be computer-readable, and such computer-readable media include volatile and non-volatile media, transient and non-transitory media, and mobile and non-mobile media. As an unrestricted example, computer-readable media may include computer-readable storage media and computer-readable transmission media.
[0341] Computer-readable storage media include volatile and non-volatile media, temporary and non-temporary media, mobile and non-mobile media, implemented in any way or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD (digital video disk) or other optical disc storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other media that can be accessed by a computer and used to store desired information.
[0342] Computer-readable transmission media typically include all information transmission media that implement computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism. The term modulated data signal means a signal whose one or more characteristics have been set or modified in order to encode information within the signal. As an unrestricted example, computer-readable transmission media include wired media such as wired networks or direct-wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the aforementioned media is also included within the scope of computer-readable transmission media.
[0343] An exemplary environment 2000 that realizes various aspects of the present invention, including a computer 2002, is shown, the computer 2002 including a processing unit 2004, system memory 2006, and a system bus 2008. The computer 200 in this specification may be used in a manner that is interoperable with computing device 100. The system bus 2008 connects system components, including (but not limited to) system memory 2006, to the processing unit 2004. The processing unit 2004 may be any processor from a variety of commonly used processors. Dual-processor and other multi-processor architectures can also be used as the processing unit 2004.
[0344] The system bus 2008 may be any of several types of bus structures that can be further interconnected to a memory bus, a peripheral bus, and a local bus using any of the various common bus architectures. System memory 2006 includes read-only memory (ROM) 2010 and random access memory (RAM) 2012. The basic input / output system (BIOS) is stored in non-volatile memory 2010, such as ROM, EPROM, or EEPROM, and this BIOS includes basic routines that help transfer information between components within the computer 2002, such as during startup. RAM 2012 may also include high-speed RAM, such as static RAM, for caching data.
[0345] Computer 2002 also includes an internal hard disk drive (HDD) 2014 (e.g., EIDE, SATA), a magnetic floppy disk drive (FDD) 2016 (e.g., for reading from or writing to a portable diskette 2018), an SSD, and an optical disk drive 2020 (e.g., for reading CD-ROM disks 2022, or for reading from or writing to other high-capacity optical media such as DVDs). The hard disk drive 2014, magnetic disk drive 2016, and optical disk drive 2020 may be coupled to the system bus 2008 by a hard disk drive interface 2024, a magnetic disk drive interface 2026, and an optical drive interface 2028, respectively. Interfaces 2024 for implementing external drives include, for example, at least one or both of USB (Universal Serial Bus) and IEEE 1394 interface technologies.
[0346] These drives and associated computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, and the like. In the case of Computer 2002, the drives and media correspond to storing any data in a suitable digital format. While the above description of computer-readable storage media refers to HDDs, portable magnetic disks, and portable optical media such as CDs or DVDs, those skilled in the art will understand that other types of computer-readable storage media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, and the like, can also be used in the exemplary operating environment, and that any such media may contain computer-executable instructions for performing the methods of the present invention.
[0347] Multiple program modules, including an operating system 2030, one or more application programs 2032, other program modules 2034, and program data 2036, can be stored in the drive and RAM 2012. All or part of the operating system, applications, modules, and / or data can also be cached in RAM 2012. It is understood that the present invention can be implemented with various commercially available operating systems or combinations of operating systems.
[0348] The user can input commands and information to the computer 2002 via one or more wired / wireless input devices, such as a keyboard 2038 and a pointing device such as a mouse 2040. Other input devices (not shown) include microphones, IR remote controls, joysticks, gamepads, stylus pens, touchscreens, and others. These and other input devices are often connected to the processing unit 2004 via an input device interface 2042 connected to the system bus 2008, but may also be connected via other interfaces such as parallel ports, IEEE 1394 serial ports, game ports, USB ports, IR interfaces, and others.
[0349] Monitor 2044 or other types of display devices are also connected to system bus 2008 via interfaces such as video adapter 2046. In addition to monitor 2044, the computer generally includes other peripheral output devices (not shown) such as speakers, printers, and others.
[0350] Computer 2002 may operate in a networked environment using logical connections to one or more remote computers, such as remote computer 2048, via wired and / or wireless communication. Remote computer 2048 may be a workstation, server computer, router, personal computer, portable computer, microprocessor-based entertainment device, peer device, or other ordinary network node, and generally include many or all of the components described for computer 2002, except for memory storage device 205, which is shown for simplicity. The illustrated logical connections include wired / wireless connections to a near-field communication network (LAN) 2052 and / or a larger network, such as a far-field communication network (WAN) 2054. Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks such as intranets, all of which may be connected to a global computer network, such as the Internet.
[0351] When used in a LAN networking environment, computer 2002 is connected to the local network 2052 via a wired and / or wireless network interface or adapter 2056. Adapter 2056 facilitates wired or wireless communication to LAN2052, which also includes a wireless access point installed therein for communication with the wireless adapter 2056. When used in a WAN networking environment, computer 2002 is connected to a communication server on WAN2054, including a modem 2058, or has other means of establishing communication over WAN2054, such as via the Internet. The modem 2058, which may be internal or external, and wired or wireless, is connected to the system bus 2008 via a serial port interface 2042. In a networked environment, program modules or parts thereof described for computer 2002 may be stored in remote memory / storage device 2050. The illustrated network connections are illustrative, and it should be understood that other means may be used to establish communication links between computers.
[0352] Computer 1602 operates to communicate with any wireless device or individual that is located and operates wirelessly, such as a printer, scanner, desktop and / or portable computer, PDA (portable data assistant), communication satellite, any device or location associated with a wirelessly discoverable tag, and a telephone. This includes at least Wi-Fi and Bluetooth wireless technologies. Thus, the communication may be a predefined structure, like a conventional network, or simply ad hoc communication between at least two devices.
[0353] It is understood that the specific order or hierarchical structure of the process stages presented is an example of an exemplary approach. It is understood that the specific order or hierarchical structure of the stages within the process can be rearranged within the scope of this disclosure based on design priorities. The claims of the methods in this disclosure provide elements of various stages in sample order, but are not limited to the specific order or hierarchical structure presented.
[0354] These and other modifications may be added to the embodiments in light of the detailed description above. In general, the terms used in the following claims should not be construed as limiting the claims to any specific embodiment disclosed in the specification and claims, but rather as encompassing all possible embodiments, along with the entire scope of the rightsable equivalents that such claims have. Accordingly, the claims are not limited by this disclosure. (Embodiments of the Invention)
[0355] As mentioned above, the relevant information is described in the best mode for carrying out the invention. [Industrial applicability]
[0356] This disclosure can be applied to a method, computer program, and / or computing device that provides alarms based on driver behavior in a driver monitoring system.
Claims
1. A method for detecting driver distraction in a Driver Monitoring System (DMS), which is performed by a computing device, wherein the method is: Receiving a first image including the driver inside the vehicle, In response to acquiring the first image, the first model is used to determine the gaze class corresponding to the first image from among a plurality of gaze classes - the plurality of gaze classes include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking away from the front - and Based on the gaze class corresponding to the first image, it is determined whether or not the driver's attention is distracted in the first image. Includes, Determining whether or not the driver's attention is distracted in the first image is: Extracting the yaw value and pitch value from the first image, The distance between the gaze cluster corresponding to the first image and the extracted yaw and pitch values, from among multiple gaze clusters generated by clustering predetermined reference images, is used to determine whether or not the driver's attention is distracted in the first image. including, method.
2. Determining whether or not the driver's attention is distracted in the first image is: The greater the distance between at least one cluster belonging to the first gaze class and the yaw and pitch values extracted from the first image, or the closer the distance between at least one cluster belonging to the second gaze class and the yaw and pitch values extracted from the first image, the higher the attention distraction score indicating the driver's potential for distraction. including, The method according to claim 1.
3. The yaw value and pitch value corresponding to the driver's face in the first image are generated by a second model different from the first model, and The second model corresponds to a pre-trained artificial intelligence-based model that outputs yaw and pitch values corresponding to the driver's face in the first image from the first image. The method according to claim 2.
4. The first model corresponds to a pre-trained artificial intelligence-based model that, in response to yaw and pitch values extracted from an image, outputs a gaze class corresponding to the yaw and pitch values and the distance between the yaw and pitch values and the gaze class. The aforementioned pre-trained first model is updated by further acquiring reference images extracted according to predetermined conditions, and If the received reference image exceeds the critical size of the queue of the first model, the older reference image is deleted from the queue with respect to the time the reference image was acquired. The method according to claim 1.
5. The first model described above is An artificial intelligence model pre-trained using a training dataset generated based on clustering of reference images from a set of images that satisfy the condition that the vehicle's speed is above a selected critical speed, corresponds to an artificial intelligence model. The method according to claim 1.
6. The aforementioned training dataset is Based on the quantitative information of the images contained in each of the plurality of gaze clusters generated by clustering the reference images, each of the plurality of gaze clusters is labeled into either the first gaze class, which indicates the driver is looking straight ahead, or the second gaze class, which indicates the driver is looking away from the front. The method according to claim 5.
7. Determining whether or not the driver's attention is distracted in the first image is: Based on the gaze clusters corresponding to the first image among a plurality of gaze clusters generated by clustering the selected reference images, and the yaw value and pitch value extracted from the first image, a first attention distraction score corresponding to the first image is determined. The first distraction score is compared with the first threshold to generate a first distraction primary estimation result indicating whether or not the distraction is present in the first image, and By performing a first voting using the first predicted result of the first distraction, the distraction alarm corresponding to the first image is determined. Includes, The first voting uses the primary attention-distraction prediction results for the first image and a selected first number of sequential images previously received for the first image to generate a collective attention-distraction prediction result representing the image group (group) consisting of the first image and the selected first number of sequential images previously received for the first image, and uses the collective attention-distraction prediction result to generate a secondary first attention-distraction prediction result that is used as a parameter for determining the attention-distraction alarm corresponding to the first image. The predicted result of collective attention distraction includes a result value that represents the set, among the result values that indicate the presence of attention distraction and the result values that indicate the absence of attention distraction. The method according to claim 1.
8. Determining whether or not the driver's attention is distracted in the first image is: From among the multiple gaze clusters generated by clustering the selected reference images, a first attention distraction score corresponding to the first image is determined based on the gaze cluster corresponding to the first image and the yaw and pitch values extracted from the first image. The first distraction score is compared with the first threshold to generate a first primary prediction result of the distraction, indicating whether or not the distraction is present in the first image. By performing a first voting using the first predicted result of the first attention distraction, a secondary predicted result of the first attention distraction is generated, and By performing a second voting using the secondary prediction result of the first distraction, the distraction alarm corresponding to the first image is determined. Includes, The first voting generates a secondary distraction prediction result used for the second voting, utilizing the primary distraction prediction result for each of the multiple images, including the first image, and the second voting uses the secondary distraction prediction result for each of the multiple images, including the first image, to determine the distraction alarm. The method according to claim 1.
9. The second voting determines whether or not there is a continuity of secondary prediction results of attention distraction within the image set consisting of the first image and a predetermined second number of sequential images previously acquired for the first image. Determining an attention-distraction alarm corresponding to the first image includes determining to generate the attention-distraction alarm if there is continuity in the secondary prediction results of the attention-distraction. The method according to claim 8.
10. Generating the secondary prediction result of the first attentional distraction is Based on the results of the first voting, generate one or more current counter values corresponding to the first image by increasing or decreasing one or more previous counter values corresponding to a second image previously acquired for the first image, and generate a secondary prediction result of the first attention distraction that includes the one or more current counter values. Includes, Determining the attention-distraction alarm corresponding to the first image is: The ON or OFF status of one or more attention-distraction alarms corresponding to the first image is determined by comparing one or more current counter values with one or more predetermined counter thresholds. including, The method according to claim 8.
11. Determining the secondary prediction result of the first attentional distraction is Based on the results of the first voting, determine at least one current counter value corresponding to the first image in a manner that increases or decreases at least one previous counter value corresponding to a second image previously received for the first image. Includes, The unit of increase or decrease for the at least one previous counter value is determined based on the time difference between the time of reception of the second image and the time of reception of the first image. The method according to claim 8.
12. The first threshold is determined based on at least one prior primary prediction result of attentional distraction corresponding to at least one prior image previously acquired of the first image. The method according to claim 7.
13. Determining the gaze class corresponding to the first image is: Extracting the yaw value and pitch value from the first image, Using the distance between the extracted yaw and pitch values and the clustering results of the reference images extracted according to predetermined conditions, the gaze cluster corresponding to the first image is determined from among the multiple gaze clusters generated by the clustering of the reference images, and Among the plurality of gaze classes, the gaze class to which the gaze cluster corresponding to the first image belongs is determined to be the gaze class corresponding to the first image. including, The method according to claim 1.
14. A computer program stored on a computer-readable storage medium, wherein, when the computer program is executed by at least one processor, the at least one processor allows the Driver Monitoring Systems (DMS) to perform an operation for detecting driver distraction, and the operation is: Receiving a first image including the driver inside the vehicle, In response to receiving the first image, the first model is used to determine the gaze class corresponding to the first image from among a plurality of gaze classes - the plurality of gaze classes include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking away from the front - and Based on the gaze class corresponding to the first image, it is determined whether or not the driver's attention is distracted in the first image. Includes, The operation to determine whether or not the driver's attention is distracted in the first image is: Extracting the yaw value and pitch value from the first image, The distance between the gaze cluster corresponding to the first image and the extracted yaw and pitch values, from among multiple gaze clusters generated by clustering predetermined reference images, is used to determine whether or not the driver's attention is distracted in the first image. including, A computer program stored on a computer-readable storage medium.
15. A computing device, at least one processor; and Memory; Includes, The aforementioned at least one processor is Receiving a first image including the driver inside the vehicle, In response to acquiring the first image, the first model is used to determine the gaze class corresponding to the first image from among a plurality of gaze classes - the plurality of gaze classes include a first gaze class in which the driver is looking straight ahead and a second gaze class in which the driver is looking away from the front - and Based on the gaze class corresponding to the first image, it is determined whether or not the driver's attention is distracted in the first image. Execute, The operation to determine whether or not the driver's attention is distracted in the first image is: Extracting the yaw value and pitch value from the first image, The distance between the gaze cluster corresponding to the first image and the extracted yaw and pitch values, from among multiple gaze clusters generated by clustering predetermined reference images, is used to determine whether or not the driver's attention is distracted in the first image. including, Computing device.