A method for providing alarms in a DMS that respond to driver behavior.

The AI-based DMS model with a voting mechanism enhances anomaly detection accuracy by optimizing alarm generation, addressing frequent false alarms and variable camera positioning issues, thus improving system reliability.

JP2026515618APending Publication Date: 2026-05-19NOTA INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NOTA INC
Filing Date
2024-09-23
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing driver monitoring systems (DMS) face issues with frequent false alarms due to sensitive detection algorithms and variable camera positioning, leading to reduced accuracy and reliability, especially under changing lighting and occlusion conditions.

Method used

Implement an artificial intelligence-based model that uses a pre-trained deep learning model to analyze driver images, applying a voting mechanism to enhance anomaly detection accuracy by considering multiple images and thresholds, thereby optimizing alarm generation.

Benefits of technology

The system provides more accurate and reliable alarms by minimizing false triggers and adapting to varying driving conditions, improving the overall performance and reliability of the DMS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515618000001_ABST
    Figure 2026515618000001_ABST
Patent Text Reader

Abstract

A method for providing alarms in a Driver Monitoring System (DMS) in response to driver behavior is disclosed. The method includes receiving a first image including the driver inside a vehicle, and using an artificial intelligence model to generate first model output information from the first image indicating the possibility that an anomaly that impairs the driver's safety exists in the first image. The method includes comparing the first model output information with a first threshold to generate a primary prediction result of a first anomaly indicating whether or not the anomaly exists in the first image, performing a first voting using the primary prediction result of the first anomaly to generate a secondary prediction result of the first anomaly, and performing a second voting using the secondary prediction result of the first anomaly to determine an anomaly alarm corresponding to the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a driver monitoring system, and more specifically, to an alarm optimization technology based on the results of driver monitoring.

Background Art

[0002] A driver monitoring system (DMS) is a technology that senses the driver's state to assist safe driving. The DMS can execute functions such as sensing the driver's drowsiness, distracted attention, whether the seat belt is worn and / or drunk driving, and warning the driver or controlling the vehicle. Such a DMS is evaluated as one of the core technologies of autonomous driving vehicles to ensure the safe operation of the vehicle, and can also contribute to protecting the driver's safety and preventing traffic accidents.

[0003] To sense the driver's state, the DMS can utilize various sensors such as cameras, infrared sensors, acceleration sensors, and gyroscopes. The camera can be used to judge the driver's drowsiness or distracted attention by sensing the expression of the driver's face and the movement of the eyelids. The infrared sensor can be used to track the driver's line of sight by sensing the movement of the driver's pupils. The acceleration sensor and the gyroscope can be used to judge drowsiness or drunk driving by sensing the driver's posture and movement.

[0004] As a method for improving the accuracy and reliability of the DMS, the sensors have been advanced. When the resolution and image quality of the camera are improved, the expression and eyelid movement of the driver can be detected more accurately. When the performance of the infrared sensor is improved, the movement of the driver's pupils can be tracked more accurately. When the performance of the acceleration sensor and the gyroscope is improved, the posture and movement of the driver can be measured more accurately.

[0005] One factor that improves the performance of DMS (Driver's Monitoring System) is the advancement of artificial intelligence technology. When artificial intelligence technology is applied to DMS, along with the sophistication of sensors, it becomes possible to more accurately perceive the driver's condition. [Overview of the project] [Problems that the invention aims to solve]

[0006] However, even if detection technology advances and the driver's state becomes more sophisticated, if the model used in the DMS is not equipped with a separate alarm-related algorithm, alarms may be triggered for every small action the driver takes, potentially interfering with driving. Furthermore, a sensitive driver detection algorithm may actually increase the frequency of false alarms. A high frequency of false alarms can lead to problems with the performance and reliability of the DMS product.

[0007] Furthermore, since aftermarket DMS products cannot specify their installation location, a problem may arise where the position of the camera must be customized for each vehicle based on the driver's position. Also, the size and angle of the face on the camera can change depending on the driver's driving habits and actions, which can reduce the accuracy of the DMS's sensing results. In addition, depending on the various occlusion and lighting conditions that occur during driving, there may be situations where objects that should be detected (e.g., seat belts) are not correctly detected in the image. [Means for solving the problem]

[0008] The inventors of this disclosure provide various embodiments for addressing various technical problems, including various technical problems in the related art and the problems identified by the inventors. The various embodiments of this disclosure are the result of efforts to optimize the alarms provided by the DMS.

[0009] One embodiment of this disclosure is based on efforts to accurately detect the driver's state or condition by taking into account various circumstances that occur during the driving process. The technical challenges described herein are not limited to those mentioned above, and other technical challenges not explicitly mentioned will be clearly understood by those skilled in the art from the following description.

[0010] In one embodiment of the content of this disclosure, a method for providing alarms in a Driver Monitoring System (DMS) that respond to driver behavior is disclosed.

[0011] This disclosure includes receiving a first image including the driver inside a vehicle; using an artificial intelligence-based model to generate first model output information from the first image indicating the possibility that an anomaly that would impair the driver's safety is present in the first image; comparing the first model output information with a first threshold to generate a primary prediction result of a first anomaly indicating whether or not the anomaly is present in the first image; performing a first voting using the primary prediction result of the first anomaly to generate a secondary prediction result of the first anomaly; and performing a second voting using the secondary prediction result of the first anomaly to determine an anomaly alarm corresponding to the first image. The artificial intelligence model corresponds to a pre-trained deep learning-based model that takes an image of the driver as input and outputs the possibility that a predetermined anomaly that would impair the driver's safety is present in the image of the driver.

[0012] In one embodiment, the predetermined abnormalities may include a first abnormality corresponding to not wearing a seat belt, a second abnormality corresponding to distraction in the driving situation, and a third abnormality corresponding to drowsiness in the driving situation. In one embodiment, the abnormalities can be identified or predetermined by pre-selected user input or automatically by the computing device 100. For example, the abnormalities can be predefined or predetermined as a seat belt non-wearing event, a strain event, a drowsiness event, a fire event, and / or a forward inattention event. For example, if a new type of abnormal event occurs, the abnormalities can be updated and added.

[0013] In one embodiment, if the first model output information is greater than or equal to the first threshold, the first prediction result for the first anomaly indicates that the anomaly exists in the first image, and if the first model output information is less than the first threshold, the first prediction result for the first anomaly indicates that the anomaly does not exist in the first image.

[0014] In one embodiment, the first model output information includes a quantitative value for comparison with the first threshold, and the secondary prediction result of the first anomaly represents a quantitative value used as a parameter for determining an anomaly alarm corresponding to the first image.

[0015] In one embodiment, the first threshold is determined based on a primary prediction result of at least one previous anomaly corresponding to at least one previously acquired previous image of the first image, or based on second model output information generated by the model in response to a second previously acquired image of the first image, or based on a primary prediction result of a second anomaly obtained by comparing the second model output information with a second threshold determined based on a primary prediction result of a third anomaly in response to a third previously acquired image of the second image. The second threshold is used to determine whether or not the anomaly is present in the second image.

[0016] In one embodiment, the first threshold is determined based on a secondary prediction result of a second anomaly corresponding to a previously acquired second image of the first image, and the secondary prediction result of the second anomaly is generated by performing the first voting using the primary prediction result of the second anomaly corresponding to the second image.

[0017] In one embodiment, the first threshold is determined based on the ratio of result values ​​indicating the presence of an anomaly from previous primary prediction results of anomalies corresponding to a predetermined first number of previously received images of the first image. If the ratio of result values ​​indicating the presence of an anomaly in the previous primary prediction results of anomalies is greater than or equal to a first ratio, the first threshold is set to a first value; and if the ratio of result values ​​indicating the presence of an anomaly in the previous primary prediction results of anomalies is less than the first ratio, the first threshold is set to a second value higher than the first value.

[0018] In one embodiment, the first voting generates a predicted result of a set of anomalies on the body surface, which consists of an image group composed of the first image and a predetermined second number of previously acquired images of the first image, in order to ensure the accuracy of the first model output information.

[0019] In one embodiment, generating a secondary prediction result for the first anomaly includes determining the majority value of the primary prediction result for the anomaly corresponding to the first image and a predetermined second number of previously acquired images of the first image, and generating the secondary prediction result for the first anomaly using the determined majority value. The majority value is determined as the result value that accounts for a higher proportion of the result value indicating the presence of the anomaly and the result value indicating the absence of the anomaly, from the primary prediction result for the anomaly.

[0020] In one embodiment, obtaining the secondary prediction result of the first anomaly includes determining a first current counter value and a second current counter value corresponding to the first image by changing each of a plurality of previous counter values ​​corresponding to a second image previously acquired based on the results of the first voting, and generating the secondary prediction result of the first anomaly including the first current counter value and the second current counter value. In one embodiment, determining the anomaly alarm corresponding to the first image includes determining whether the first anomaly alarm corresponding to the first image is ON or OFF by comparing the first current counter value with a first counter threshold, and determining whether the second anomaly alarm corresponding to the first image is ON or OFF by comparing the second current counter value with a second counter threshold.

[0021] In one embodiment, generating the secondary prediction result of the first anomaly includes deciding to increase at least one previous counter value corresponding to a second image previously acquired for the first image if the result of the first voting, which indicates the mainstream value of the primary prediction result of the anomaly or the prediction result of a collective anomaly in a predetermined plurality of images, indicates the presence of an anomaly, or deciding to decrease the at least one previous counter value if the result of the first voting, which indicates the mainstream value of the primary prediction result of the anomaly or the prediction result of a collective anomaly in a predetermined plurality of images, indicates the absence of an anomaly.

[0022] In one embodiment, generating a secondary prediction result for the first anomaly includes determining the current counter value corresponding to the first image by changing a previous counter value corresponding to a previously acquired second image of the first image based on the results of the first voting, and generating a secondary prediction result for the first anomaly that includes the current counter value. Determining the anomaly alarm corresponding to the first image includes determining whether the third anomaly alarm corresponding to the first image is ON or OFF by comparing the current counter value with a third counter threshold, and determining whether the fourth anomaly alarm corresponding to the first image is ON or OFF by comparing the current counter value with a fourth counter threshold.

[0023] In one embodiment, generating the secondary prediction result of the first anomaly includes determining at least one current counter value corresponding to the first image by changing at least one previous counter value corresponding to a previously acquired second image of the first image based on the result of the first voting. If the current counter value falls outside a range defined as a predetermined minimum and maximum value, the current counter value is set to the minimum or maximum value.

[0024] In one embodiment, generating the secondary prediction result of the first anomaly includes determining at least one current counter value corresponding to the first image by changing at least one previous counter value corresponding to a previously acquired second image of the first image, based on the result of the first voting. The unit of change of the at least one previous counter value is determined based on the time difference between the acquisition time of the second image and the acquisition time of the first image.

[0025] In one embodiment, in order to ensure the accuracy of the generation of the alarm, the second voting determines an image set composed of a predetermined third number of sequential images including the first image, and determines whether there is continuity in the secondary prediction results of anomalies corresponding to the images constituting the image set. The third number of sequential images includes the first image and images acquired before the first image.

[0026] In one embodiment, determining the anomaly alarm corresponding to the first image includes determining to generate the anomaly alarm corresponding to the first image when all the secondary prediction results of anomalies in an image set composed of a predetermined third number of sequential images including the first image indicate the presence of the anomaly. The third number of sequential images includes the first image and images acquired before the first image.

[0027] In one embodiment, determining the anomaly alarm corresponding to the first image includes determining whether to generate the anomaly alarm corresponding to the first image based on the secondary prediction result of the first anomaly corresponding to the first image and the secondary prediction results of previous anomalies corresponding to each of a predetermined fourth number of previous images of the first image.

[0028] In one embodiment, the first voting utilizes the primary prediction results of previous anomalies corresponding to at least one previous image received before the first image and the primary prediction result of the first anomaly, and the second voting utilizes a combination of the secondary prediction results of previous anomalies corresponding to the at least one previous image and the secondary prediction result of the first anomaly, or utilizes comparing the secondary prediction result of the first anomaly with a counter threshold value.

[0029] In one embodiment, a computer program stored on a computer-readable storage medium is disclosed. The computer program, when executed by at least one processor, allows the at least one processor to perform operations for providing driver behavior-based alarms in a Driver Monitoring System (DMS), the operations including: acquiring a first image including the driver inside a vehicle; using an artificial intelligence-based model to acquire first model output information from the first image indicating the possibility that an anomaly that would impair the driver's safety exists in the first image; comparing the first model output information with a first threshold to acquire a first primary prediction result of an anomaly indicating whether or not the anomaly exists in the first image; performing a first voting using the first primary prediction result of an anomaly to acquire a first secondary prediction result of an anomaly; and performing a second voting using the first secondary prediction result of an anomaly to determine an anomaly alarm corresponding to the first image.

[0030] The artificial intelligence model corresponds to a pre-trained deep learning-based model that takes an image of the driver as input and outputs the possibility that a predetermined abnormality that impairs the driver's safety exists in the driver's image.

[0031] In one embodiment, a computing device is disclosed. The computing device may include at least one processor and memory. The at least one processor can perform the following operations: acquire a first image including a driver inside a vehicle; use an artificial intelligence-based model to acquire first model output information from the first image indicating the possibility that an anomaly that would impair the driver's safety is present in the first image; compare the first model output information with a first threshold to acquire a first primary prediction result of an anomaly indicating whether or not the anomaly is present in the first image; perform a first voting using the first primary prediction result of the anomaly to acquire a second primary prediction result of the anomaly; and perform a second voting using the second primary prediction result of the anomaly to determine an anomaly alarm corresponding to the first image. The artificial intelligence model corresponds to a pre-trained deep learning-based model that takes an image of the driver as input and outputs the possibility that a predetermined anomaly that would impair the driver's safety is present in the image of the driver. [Effects of the Invention]

[0032] A technology according to one embodiment of the present disclosure can optimize alarms provided by the DMS.

[0033] The technology according to one embodiment of the disclosed information can more accurately sense the driver's condition by taking into account various situations that occur during driving. [Brief explanation of the drawing]

[0034] [Figure 1] A schematic block diagram of a computing device according to one embodiment of the disclosed information is shown. [Figure 2] This disclosure shows an exemplary structure of an artificial intelligence-based model according to one embodiment of the disclosed material. [Figure 3] This disclosure illustrates a method for detecting an anomaly and determining an anomaly alarm in a DMS according to one embodiment of the disclosed information. [Figure 4] A method for determining abnormal alarms in a DMS according to one embodiment of the disclosed information is illustrated as an example. [Figure 5] An exemplary method for determining a driver's drowsiness level according to one embodiment of the content of this disclosure is provided. [Figure 6] This disclosure illustrates an example of using a heatmap in the process of detecting the eye region in a DMS according to one embodiment of the disclosed information. [Figure 7] This disclosure illustrates an example of a method for determining the driver's drowsiness level using a heat map in a DMS according to one embodiment of the disclosed material. [Figure 8] An exemplary method for determining a driver drowsiness alarm according to one embodiment of the content of this disclosure is provided. [Figure 9] An example of a spatial transformation method according to one embodiment of the content of this disclosure is shown. [Figure 10] An exemplary method for determining a driver drowsiness alarm according to one embodiment of the content of this disclosure is provided. [Figure 11] This disclosure illustrates a methodology for generating an abnormal alarm according to one embodiment of the disclosed information. [Figure 12] This disclosure illustrates a methodology for generating an abnormal alarm according to one embodiment of the disclosed information. [Figure 13] This disclosure illustrates a methodology for generating an abnormal alarm according to one embodiment of the disclosed information. [Figure 14] An exemplary method for generating a closed-eyes alarm or drowsiness alarm according to one embodiment of the present disclosure is provided. [Figure 15] An exemplary method for generating a closed-eyes alarm or drowsiness alarm according to one embodiment of the present disclosure is provided. [Figure 16] This is a schematic diagram of a computing environment according to one embodiment of the information disclosed herein. [Modes for carrying out the invention]

[0035] This application claims priority and benefits based on Korean Patent Application No. 10-2023-0142146, filed with the Korea Intellectual Property Office on 23 October 2023, and Korean Patent Application No. 10-2023-0177469, filed with the Korea Intellectual Property Office on 8 December 2023, the entire contents of which are incorporated herein by reference.

[0036] Various embodiments are described with reference to the drawings. Various explanations are provided herein to provide an understanding of the disclosure. Before describing specific details for implementing the disclosure, note that configurations not directly related to the technical essence of the disclosure have been omitted to the extent that they do not obscure the technical essence of the invention. Furthermore, terms or words used herein and in the claims should be interpreted as having meanings and concepts consistent with the technical idea of ​​the invention, based on the principle that inventors may define appropriate terms to best describe their inventions.

[0037] As used herein, terms such as “model,” “system,” and / or “module” refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or software executions, and can be used interchangeably with each other. For example, a module may be, but is not limited to, a process executed on a processor, a processor, an object, an execution thread, a program, an application, and / or a computing device. One or more modules may reside within a processor and / or an execution thread. A module may be localized to one computer. A module may be distributed between two or more computers. Such a module may also be executed from various computer-readable media having various data structures stored within it. A module may communicate via local and / or teleprocessing according to signals having one or more data packets (e.g., data from one component interacting with other components in a local system, a distributed system, and / or data transmitted to other systems via signals over a network such as the Internet).

[0038] Furthermore, the term "or" is intended to mean inclusive "or" rather than exclusive "or". That is, unless otherwise specified or contextually clear, "X uses A or B" is intended to mean one of the natural inclusive substitutions. That is, if X uses A; X uses B; or X uses both A and B, "X uses A or B" can apply to any of these cases. Also, the terms "and / or" and "at least one" as used herein should be understood to refer to and include all possible combinations of one or more of the listed related items. For example, the terms "at least one of A or B" or "at least one of A and B" should be interpreted as meaning "including only A," "including only B," and "a combination of A and B."

[0039] Furthermore, the terms “contains” and / or “includes” should be understood to mean the presence of the feature and / or component in question. However, it should be understood that the terms “contains” and / or “includes” do not exclude the presence or addition of one or more other features, components, and / or groups thereof. Furthermore, where not otherwise specified or where it is not contextually clear that it refers to the singular, in this specification and claims, the singular should generally be interpreted as meaning “one or more.”

[0040] Furthermore, those skilled in the art should recognize that various exemplary logical components described in relation to the embodiments disclosed herein can be implemented in hardware, computer software, or a combination of both.

[0041] The description of the presented embodiments is provided so that a person with ordinary skill in the art of this disclosure may utilize or practice the invention. Various modifications to such embodiments will be obvious to a person with ordinary skill in the art of this disclosure. The general principles defined herein can be applied to other embodiments without departing from the scope of this disclosure. Thus, the invention is not limited by the embodiments presented herein. The invention should be analyzed in the broadest sense, consistent with the principles and novel features presented herein.

[0042] In this disclosure, terms such as 1st, 2nd, or 3rd, represented by the Nth, are used to distinguish at least one entity. For example, the entities represented by 1st and 2nd may be identical or different from each other.

[0043] The terms expressed in this disclosure as primary, secondary, or tertiary are used to distinguish at least one entity. For example, the terms N, such as primary, secondary, and tertiary, may be used to distinguish temporal order. In such examples, a larger value of N may indicate a later entity in time, and a smaller value of N may indicate an earlier entity in time.

[0044] As used in this disclosure, the term “model” can be used to encompass artificial intelligence-based models, AI models, computational models, neural networks, network functions, and neural networks. In one embodiment, a model may mean a model file, model identification information, the model's execution environment, the model's running time, and / or the model's framework.

[0045] In this disclosure, a Driver Monitoring System (DMS) can represent a software entity or hardware entity, or a combination thereof, to which vehicle technology for monitoring the driver's condition is applied. Such a DMS may be executed by a computing device according to one embodiment of this disclosure.

[0046] As used in this disclosure, the term "image" can be used to encompass one or more frames. For example, one image may correspond to one frame. In one embodiment, the image may be a still image and correspond to a frame obtained from a video.

[0047] In one embodiment, the image may be captured data in which the driver is included as an object, and may be acquired via a camera installed in the vehicle. In one embodiment, the image may be analyzed and / or processed in the DMS to detect an anomaly corresponding to the driver and / or determine whether, type, and / or intensity of an alarm corresponding to the anomaly is generated.

[0048] In this specification, the expression "determine an alarm" may be used to encompass determining whether or not to generate an alarm, determining the type of alarm, and / or determining the intensity of the alarm.

[0049] For the sake of clarity, the technology described herein will be referred to as "images" below. It will be apparent to those skilled in the art that the technology described herein can be implemented using the term "frames."

[0050] As used in this disclosure, the term “anomaly” may be used to describe any action, situation, or element within an image or frame that impedes driver safety. For example, an anomaly may include driver distraction, drowsiness, failure to wear a seatbelt, and / or the presence of flames (or fire).

[0051] As used in this disclosure, the term “voting” may mean an algorithm for correcting predicted anomaly results in order to improve the accuracy of anomaly detection and / or the accuracy of anomaly alarm generation. In one embodiment, voting may mean a rule-based algorithm that combines results from multiple images to determine, correct, and / or adjust anomaly results corresponding to the current image. In one embodiment, voting may mean a method for determining a predicted result from the current image based on predicted results from previous images. In this disclosure, predicted anomaly results and predicted results can be used interchangeably.

[0052] In one embodiment, each of the multiple factors used in the voting may correspond to each of the multiple images acquired over time. In one embodiment, each of the multiple factors used in the voting may correspond to each of the predicted results of the multiple images acquired over time (e.g., predicted results obtained from a model, and / or predicted results that reflect previous voting results). For example, a first predicted result for the first image, a second predicted result for the second image, and a third predicted result for the third image can be considered as factors used in the voting.

[0053] The technology according to one embodiment of the present disclosure allows for the sequential use of multiple voting processes. For example, the results of a first voting process may be used in a second voting process, and the results of a second voting process may be used in a third voting process. By sequentially using multiple voting processes that use predicted results corresponding to previous and current images as factors, more accurate anomaly alarms can be provided.

[0054] Figure 1 schematically shows a block diagram of a computing device 100 according to one embodiment of the present disclosure.

[0055] According to embodiments of this disclosure, the computing device 100 may include a processor 110 and memory 130.

[0056] The configuration of the computing device 100 shown in Figure 1 is merely a simplified example. In one embodiment of this disclosure, the computing device 100 may include other configurations for executing the computing environment of the computing device 100, and only a portion of the disclosed configuration may constitute the computing device 100. For example, if the computing device 100 described above includes a user terminal, an output unit (not shown) and an input unit (not shown) may be included within the scope of the computing device 100.

[0057] In this disclosure, computing device 100 can be used to encompass any form of server and any form of terminal. Computing device 100 may be used interchangeably with computing equipment.

[0058] In this disclosure, computing device 100 may mean any form of component that constitutes a system for realizing embodiments of this disclosure.

[0059] In one embodiment, the computing device 100 may mean a device on which the DMS is driven.

[0060] In one embodiment, the computing device 100 may mean a device for detecting anomalies from the driver's image and / or determining whether or not to generate an alarm corresponding to the anomaly.

[0061] In one embodiment, the computing device 100 may mean a device used to train a model for detecting anomalies from images of the driver.

[0062] In one embodiment, the computing device 100 may mean a device used to train a model for detecting anomalies from images of a driver. In one embodiment, the computing device 100 may mean a device used to infer a model for detecting anomalies from images of a driver.

[0063] In one embodiment, the computing device 100 may mean a server located remotely from the in-vehicle device that acquires the driver's image.

[0064] In one embodiment, the computing device 100 may acquire an image of the driver from a device in the vehicle, detect anomalies in the image, and / or determine an alarm corresponding to the anomaly in the image. For example, it may determine whether or not to generate an alarm corresponding to the anomaly, or the intensity of the alarm.

[0065] In one embodiment, the computing device 100 may obtain object detection results from a target image that includes the driver. For example, the computing device 100 may obtain detection results from the target image that relate to seat belts, eye closure, drowsiness, and / or inattention.

[0066] In one embodiment, the processor 110 may perform the overall operation of the computing device 100. The processor 110 may consist of at least one core. The processor 110 may include devices for data analysis and / or processing, such as a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU) of the computing device 100.

[0067] The processor 110 may read a computer program stored in the memory 130 and, in accordance with one embodiment of the present disclosure, detect an anomaly and / or determine an anomaly alarm.

[0068] In one embodiment of this disclosure, the processor 110 may perform calculations for learning a neural network. The processor 110 may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), feature extraction from the input data, error calculation, and weighting updates of the neural network using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor 110 may process the learning of network functions. For example, the CPU and GPGPU may both process the learning of network functions and data classification using network functions. Furthermore, in one embodiment of this disclosure, the processors of multiple computing devices can be used together to process the learning of network functions and data classification using network functions. Furthermore, the computer program executed on the computing device 100 in one embodiment of this disclosure may be a CPU, GPGPU, or TPU executable program.

[0069] Furthermore, the processor 110 may typically handle the overall operation of the computing device 100. For example, the processor 110 may provide the user with appropriate information or functions by processing data, information, or signals that are input or output through components included in the computing device 100, or by driving application programs stored in memory.

[0070] According to one embodiment of the present disclosure, the memory 130 may store any form of information generated or determined by the processor 110 and any form of information received by the computing device 100. According to one embodiment of the present disclosure, the memory 130 may also be a storage medium for storing computer software that causes the processor 110 to perform the operations according to the embodiments of the present disclosure. Thus, the memory 130 may mean a computer-readable medium for storing software code necessary to perform the embodiments of the present disclosure, data on which the code is executed, and the results of the code execution.

[0071] According to one embodiment of this disclosure, memory 130 may mean any type of storage medium. For example, memory 130 may include at least one type of storage medium from among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, magnetic disk, and optical disk. The computing device 100 may also operate in conjunction with web storage that performs the storage functions of memory 130 over the internet. The above descriptions of memory are illustrative, and memory 130 used in this disclosure is not limited to the above examples.

[0072] The communication unit (not shown) of this disclosure can be configured in any manner, including wired and wireless, and may be composed of various communication networks such as Personal Area Networks (PANs) and Wide Area Networks (WANs). Furthermore, the network unit 150 can operate on a known World Wide Web (WWW) basis and may utilize wireless transmission technologies used for short-range communication, such as infrared (IrDA) or Bluetooth®.

[0073] The computing device 100 in this disclosure may include any form of user terminal and / or any form of server. Therefore, embodiments of this disclosure may be performed by a server and / or user terminal.

[0074] In one embodiment, the user terminal may include any form of terminal capable of interacting with a server or other computing device. The user terminal may include, for example, a mobile phone, a smartphone, a laptop computer, a PDA (personal digital assistant), a slate PC, a tablet PC, and an ultrabook. In one embodiment, the user terminal may mean a device including a camera installed in a vehicle.

[0075] In one embodiment, the server may include any type of computing system or computing device, such as a microprocessor, a mainframe computer, a digital processor, a portable device, and a device controller.

[0076] In one embodiment, the server may include a storage unit (not shown) for storing data and / or information used in the present disclosure. Such a storage unit may be contained within the server or reside under the server's control. In another example, the storage unit may reside outside the server and be implemented in a manner that allows it to communicate with the server. In this case, the storage unit may be managed and controlled by another external server different from the server.

[0077] Figure 2 shows an exemplary structure of an artificial intelligence-based model according to one embodiment of the present disclosure.

[0078] In this disclosure, models, artificial intelligence models, artificial intelligence-based models, computational models, neural networks, network functions, and neural networks may be used interchangeably with each other.

[0079] The artificial intelligence-based models described herein may include models applicable to various domains, such as models for image processing including object segmentation, object detection, anomaly detection, and / or object classification, and models for text processing including data prediction, text semantic inference, and / or data classification.

[0080] A neural network can generally be composed of a set of interconnected computational units called nodes. Such nodes are sometimes called neurons. A neural network consists of at least one node. The nodes (or neurons) that make up a neural network may be interconnected by one or more links.

[0081] Nodes within an artificial intelligence model may be used to represent components that make up a neural network; for example, nodes in a neural network may correspond to neurons.

[0082] Within a neural network, one or more nodes connected by links can form a relative input-output node relationship. The concepts of input and output nodes are relative; any node that is an output node to another node is an input node to another node, and vice versa. As mentioned above, the relationship between input and output nodes may be generated around links. One input node may be connected to one or more output nodes via links, and vice versa.

[0083] In a relationship between input and output nodes connected via a single link, the data of the output node may be determined based on the data input to the input node. Here, the link connecting the input and output nodes may have weights. The weights may be variable and can be varied by the user or algorithm in order for the neural network to perform a desired function. For example, if one or more input nodes are interconnected to one output node by their respective links, the output node may determine its output node value based on the values ​​input to the input nodes connected to the output node and the weights set for the links corresponding to each input node.

[0084] As mentioned above, a neural network consists of one or more nodes interconnected via one or more links, forming input-output node relationships within the network. The characteristics of a neural network may be determined by the number of nodes and links within the network, the relationships between nodes and links, and the weighting values ​​assigned to each link. For example, if there are two neural networks with the same number of nodes and links but different link weighting values, the two neural networks can be recognized as distinct from each other. A neural network may consist of a set of one or more nodes. A subset of nodes constituting a neural network can constitute a layer. A portion of the nodes constituting a neural network can constitute a single layer based on their distance from the initial input node. For example, a set of nodes that are n in distance from the initial input node can constitute an n-layer. The distance from the initial input node can be defined by the minimum number of links that must be traversed to reach the node in question from the initial input node. However, such a definition of a layer is arbitrary for illustrative purposes, and the order of layers within a neural network may be defined in ways different from those described above. For example, the layer of nodes may be defined by their distance from the final output node.

[0085] In embodiments of this disclosure, a collection of neurons or nodes may be defined as a “layer.”

[0086] The initial input node may mean one or more nodes in the neural network that receive data directly without links in relation to other nodes. Alternatively, it may mean a node in the neural network that does not have other input nodes connected by links in relation to other nodes based on links. Similarly, the final output node may mean one or more nodes in the neural network that do not have an output node in relation to other nodes. Furthermore, a hidden node may mean a node in the neural network that is neither the initial input node nor the final output node.

[0087] A neural network according to one embodiment of the present disclosure may have the same number of nodes in the input layer as the number of nodes in the output layer, and the number of nodes may decrease as you move from the input layer to the hidden layer, and then increase again. Another neural network according to another embodiment of the present disclosure may have fewer nodes in the input layer than the number of nodes in the output layer, and the number of nodes may decrease as you move from the input layer to the hidden layer. Yet another neural network according to yet another embodiment of the present disclosure may have more nodes in the input layer than the number of nodes in the output layer, and the number of nodes may increase as you move from the input layer to the hidden layer. Another neural network according to another embodiment of the present disclosure may be a neural network that is a combination of the neural networks described above.

[0088] A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to input and output layers. Deep neural networks can be used to understand the latent structures of data. For example, the latent structures of photographs, text, videos, audio, protein sequence structures, gene sequence structures, peptide sequence structures, and / or music may be understood via a deep neural network. Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, Generative Adversarial Networks (GANs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siam networks, and Generative Adversarial Networks (GANs). The above descriptions of deep neural networks are illustrative and this disclosure is not limited thereto.

[0089] The artificial intelligence models described herein can be represented by a network structure of any of the aforementioned structures, including an input layer, a hidden layer, and an output layer.

[0090] The neural networks that can be used in the clustering models of this disclosure may be trained in at least one of the following ways: supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Training of a neural network may be the process of applying knowledge to the neural network so that it can perform a particular action.

[0091] Neural networks can be trained to minimize the error in their output. Training a neural network involves repeatedly inputting training data, calculating the network's output and target error for the training data, and updating the weights of each node in the neural network by backpropagating the error from the output layer to the input layer in a way that reduces the error. In guided learning, training data with the correct answer labeled is used (i.e., labeled training data), while in unguided learning, the training data may not be labeled. For example, in guided learning for data classification, the training data may be data with a category labeled for each data point. The labeled training data may be input to the neural network, and the error may be calculated by comparing the neural network's output (category) with the labels on the training data. As another example, in unguided learning for data classification, the error may be calculated by comparing the input training data with the output of the neural network. The calculated errors are backpropagated in the reverse direction of the neural network (i.e., from the output layer to the input layer), and this backpropagation can update the connection weights of each node in each layer of the neural network. The amount of change in the connection weights of each node being updated may be determined by the learning rate. The computation of the neural network on the input data and the backpropagation of errors can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of iterations of the neural network's learning cycle. For example, a high learning rate can be used in the early stages of learning to increase efficiency by allowing the neural network to quickly achieve a certain level of performance, while a lower learning rate can be used in the later stages of learning to improve accuracy.

[0092] In neural network training, training data is generally a subset of real-world data (i.e., data that the trained neural network intends to process). Therefore, there can be training cycles where errors on training data decrease, but errors on real-world data increase. Overfitting is this phenomenon where the network over-trains on training data, leading to increased errors on real-world data. For example, a neural network trained to recognize cats by being shown yellow cats may fail to recognize cats that are not yellow; this is a type of overfitting. Overfitting can act as a cause of increased errors in machine learning algorithms. Various optimization methods can be used to prevent such overfitting. To prevent overfitting, methods such as increasing the amount of training data, regularization, dropout (deactivating some of the network nodes during the training process), and the use of a batch normalization layer can be applied.

[0093] One embodiment of the present disclosure discloses a computer-readable medium storing a data structure including an artificial intelligence-based model. The aforementioned data structure may be stored in a storage unit (not shown) of the present disclosure, executed by a processor 110, and transmitted and received by a communication unit (not shown).

[0094] A data structure may mean the organization, management, and storage of data that enables efficient access to and modification of data. A data structure may also mean the organization of data to solve a specific problem (e.g., data retrieval, data storage, data modification in the shortest time). A data structure can also be defined as physical or logical relationships between data elements designed to support specific data processing functions. Logical relationships between data elements may include user-defined linking relationships between data elements. Physical relationships between data elements may include actual relationships between data elements physically stored in a computer-readable storage medium (e.g., persistent storage). Specifically, a data structure may include a collection of data, relationships between data, and functions or instructions that can be applied to data. A well-designed data structure may enable a computing device to perform operations with minimal use of its resources. Specifically, a well-designed data structure can improve the efficiency of operations such as arithmetic, reading, insertion, deletion, comparison, exchange, and retrieval.

[0095] Data structures can be divided into linear and non-linear data structures depending on their form. A linear data structure is one in which only one piece of data is linked after another. Linear data structures may include lists, stacks, queues, and decks. A list may refer to a set of data that has an internal order. A list may also include linked lists. A linked list may be a data structure in which data is linked in a linear fashion, with each piece of data having a pointer. In a linked list, the pointer may contain linking information to the next or previous piece of data. Linked lists can be expressed as single linked lists, double linked lists, or circular linked lists depending on their form. A stack is a data sequence structure in which data can be accessed in a restricted manner. A stack may be a linear data structure in which data can only be processed (e.g., inserted or deleted) at one end of the data structure. Data stored in a stack may be a LIFO (Last In First Out) data structure, where data entered later comes out earlier. A queue is a data arrangement structure that restricts access to data, and unlike a stack, it may be a data structure where data stored later is retrieved later (FIFO - First in First Out). A deck may be a data structure that allows data to be processed from both ends of the data structure.

[0096] A nonlinear data structure is a structure in which multiple data are linked after a single data item. Nonlinear data structures may include graph data structures. Graph data structures can be defined by vertices and edges, and edges may include lines connecting two different vertices. Graph data structures may also include tree data structures. A tree data structure may be a data structure in which a path connecting two different vertices among the multiple vertices included in the tree is a single data structure. In other words, a graph data structure may not form a loop.

[0097] The data structure may include a neural network. Furthermore, the data structure including a neural network may be stored on a computer-readable medium. The data structure including a neural network may also include pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. The data structure including a neural network may include any of the components of the disclosed configuration. That is, the data structure including a neural network may consist of all or any combination thereof of pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. In addition to the configurations described above, the data structure including a neural network may include any other information that determines the properties of the neural network. Furthermore, the data structure may include, but is not limited to, any form of data used or generated during the computational process of the neural network. Computer-readable media may include computer-readable recording media and / or computer-readable transmission media. A neural network can generally consist of a collection of interconnected computational units called nodes. Such nodes are sometimes called neurons. A neural network consists of at least one node.

[0098] The data structure may include data to be input to a neural network. The data structure including data to be input to a neural network may be stored on a computer-readable medium. The data to be input to a neural network may include training data input during the neural network's learning process and / or input data to be input to a neural network after training is complete. The data to be input to a neural network may include pre-processed data and / or data subject to pre-processing. Pre-processing may include data processing processes for inputting data to a neural network. Therefore, the data structure may include data subject to pre-processing and data generated during pre-processing. The data structures described above are illustrative and the disclosure is not limited thereto.

[0099] The data structure may include weights for the neural network (in this specification, weights and parameters may be used interchangeably). The data structure including the weights for the neural network can be stored on a computer-readable medium. The neural network may include multiple weights. The weights are variable and may be varied by the user or algorithm in order for the neural network to perform a desired function. For example, if one or more input nodes are interconnected to an output node by their respective links, the output node may determine the data values ​​output from the output node based on the values ​​input to the input nodes connected to the output node and the weights set for the links corresponding to each input node. The data structures described above are illustrative and the disclosure is not limited thereto.

[0100] As a non-limiting example, the weights may include weights that change during the neural network learning process and / or weights after the neural network has finished learning. The weights that change during the neural network learning process may include weights at the start of a learning cycle and / or weights that change during a learning cycle. The weights after the neural network has finished learning may include weights after the learning cycle has finished. Therefore, a data structure containing neural network weights may include a data structure containing weights that change during the neural network learning process and / or weights after the neural network has finished learning. Accordingly, the weights and / or each combination of weights described above shall be included in the data structure containing neural network weights. The data structures described above are illustrative and the disclosure is not limited thereto.

[0101] A data structure containing neural network weights can be stored on a computer-readable storage medium (e.g., memory, hard disk) after undergoing a serialization process. Serialization may be a process of converting the data structure into a format that can be stored on the same or different computing devices and later reconfigured for use. Computing devices can serialize the data structure to send and receive data over a network. A serialized data structure containing neural network weights may be reconfigured on the same or other computing devices through deserialization. A data structure containing neural network weights is not limited to serialization. Furthermore, a data structure containing neural network weights may include data structures that enhance computational efficiency while minimizing the use of computing device resources (e.g., in nonlinear data structures, B-trees, R-trees, tries, m-way search trees, AVL trees, Red-Black Trees). The foregoing is illustrative, and this disclosure is not limited thereto.

[0102] The data structure may include the hyperparameters of the neural network. The data structure containing the neural network hyperparameters can be stored on a computer-readable medium. The hyperparameters may be variable variables controlled by the user. Examples of hyperparameters may include the learning rate, cost function, number of iterations in the learning cycle, weight initialization (e.g., setting the range of weights to be initialized), and number of Hidden Units (e.g., number of hidden layers, number of nodes in hidden layers). The aforementioned data structure is illustrative and the disclosure is not limited thereto.

[0103] Figure 3 illustrates an example of a method for detecting an anomaly and determining an anomaly alarm in a DMS according to one embodiment of the present disclosure.

[0104] As shown in Figure 3, the computing device 100 may detect the driver's face (310).

[0105] In one embodiment, the computing device 100 may acquire images from a camera installed inside the vehicle. For example, the images may include the driver inside the vehicle.

[0106] In one embodiment, the computing device 100 may use a face detection model to detect the driver's face from the acquired image. For example, the model may include an object detection model and / or an object segmentation model. For example, the model may correspond to a pre-trained artificial intelligence-based model to detect and / or segment human faces in the image.

[0107] In one embodiment, the model for face detection may correspond to the detection model.

[0108] In one embodiment, the face detection model may output the result of segmenting contours that define faces in an image. In another embodiment, the face detection model may output bounding boxes containing faces from an image. In yet another embodiment, the face detection model may be configured to output together a region in an image corresponding to a face and a plurality of feature points that can identify a face within the region. In yet another embodiment, the face detection model may be configured to output together a region in an image corresponding to a face, a plurality of feature points that can identify a face within the region, and a plurality of feature points that can identify eyes within the face.

[0109] In one embodiment, the computing device 100 may detect a face landmark from the detected face (320).

[0110] In one embodiment, the computing device 100 may use a model for facial landmark detection to acquire feature points contained in the driver's face on the driver's face. In this disclosure, feature points and landmarks can be used interchangeably.

[0111] In one embodiment, the model for facial landmark detection may correspond to a pre-trained artificial intelligence-based model that determines a plurality of feature points for identifying facial features (e.g., eye features, nasal features, and / or mouth features) on the driver's face. In another embodiment, the model for facial landmark detection may be configured to output together a plurality of feature points that can identify a face within a facial region, and a plurality of feature points that can identify eyes within the face.

[0112] In one embodiment, the model for detecting facial landmarks may correspond to a detection model.

[0113] In one embodiment, the driver's identity can be identified based on the detection of the driver's facial landmark. Identification of the driver's identity may be performed based on a comparison between a pre-stored driver's facial landmark and the detected driver's facial landmark.

[0114] In one embodiment, the computing device 100 may detect eye landmarks from the image (330).

[0115] For example, eye landmark detection may be performed more efficiently using the results of face detection. For example, eye landmark detection may be included in the results of face detection. Based on eye landmark detection, the computing device 100 can determine whether the driver is drowsy and / or distracted. For example, eye landmark detection may detect the position of the eyes within the face, whether the eyes are closed, and / or the direction the eyes are looking.

[0116] For example, a model for detecting eye landmarks may correspond to a detection model.

[0117] In one embodiment, the computing device 100 may detect anomalies from the image (340).

[0118] In one embodiment, the computing device 100 may use the results of eye landmark detection to detect anomalies from the image. For example, the computing device 100 may use the results of eye landmark detection to determine whether the driver is distracted while driving (e.g., not paying attention to the road ahead). For example, the computing device 100 may use the results of eye landmark detection to determine whether the driver is drowsy while driving.

[0119] In this disclosure, "abnormality" may be used to describe abnormal behavior while driving. In this disclosure, "abnormality" may be used to describe situations and / or behaviors that impede safety while driving. For example, an abnormality may include a first abnormality corresponding to not wearing a seat belt, a second abnormality corresponding to distraction in the driving situation, a third abnormality corresponding to drowsiness in the driving situation, a fourth abnormality corresponding to the driver smoking, a fifth abnormality corresponding to a fire in the vehicle, and / or a sixth abnormality corresponding to closing one's eyes while driving.

[0120] In other embodiments, the computing device 100 may use one model to detect multiple anomalies. In other embodiments, the computing device 100 may operate to detect multiple anomalies using multiple models (for example, models dedicated to specific anomaly detection).

[0121] In other embodiments, the computing device 100 may detect the driver's body from the image. For example, after detecting the driver's face, the driver's body in the image may be detected based on the driver's face. In another example, driver body detection may be performed independently of driver face detection. Based on such body detection, it may be determined whether the driver is wearing a seat belt, whether the driver is smoking, and / or where the driver's hands are. Based on the driver's body detection, abnormalities related to the driver's body (e.g., not wearing a seat belt, smoking, and / or a fire inside the vehicle) may be detected.

[0122] For example, a model for detecting a body may correspond to a detection model.

[0123] In other embodiments, the computing device 100 may use one model to detect multiple anomalies. In other embodiments, the computing device 100 may operate to detect multiple anomalies using multiple models (for example, models dedicated to specific anomaly detection).

[0124] In one embodiment, the computing device 100 may use a classification model that utilizes eye detection results and / or face detection results to determine whether an anomaly (e.g., distraction) corresponds to the input image. As a non-restrictive example, such a classification model may operate to divide the input image into forward or non-forward sections. As a non-restrictive example, such a classification model may operate to output quantitative values ​​indicating whether the input image is forward or non-forward.

[0125] In one embodiment, the computing device 100 may use the anomaly detection result to determine an alarm corresponding to the anomaly (350).

[0126] In one embodiment, an alarm or abnormal alarm corresponding to an anomaly may include various forms of output to cause the user to recognize the anomaly, such as sound, images, vibrations and / or light.

[0127] For example, if computing device 100 determines that an anomaly exists in an image using one or more models, it may decide whether or not to generate an alarm corresponding to the anomaly. If anomaly detection directly leads to an anomaly alarm, there is a problem that the driver may receive unnecessary or inaccurate alarms while driving. Therefore, the technology according to one embodiment of the present disclosure can provide the user with more optimal and accurate alarms by deciding whether or not to generate an anomaly alarm using the anomaly detection result. For example, the presence of an anomaly may be determined to be a situation in which driver distraction is detected, a situation in which driver drowsiness is detected, and / or a situation in which the driver is not wearing a seat belt.

[0128] For example, if computing device 100 determines that an anomaly exists in the image, it may determine the type of alarm corresponding to the anomaly and / or the intensity of the alarm. Computing device 100 may also use one or more models to determine the intensity of an anomaly alarm in the image. For example, computing device 100 may determine multiple anomaly alarms or the intensity of an anomaly alarm by comparing the expected result associated with the anomaly with each of several thresholds. For example, computing device 100 may determine multiple anomaly alarms or the intensity of an anomaly alarm by applying one or more counter concepts to the expected result associated with the anomaly and comparing each of the counters with a threshold. Thus, the technology according to one embodiment of the present disclosure can provide the user with more optimal and accurate alarms by using the anomaly detection result to determine whether or not to generate an anomaly alarm and / or the intensity of the anomaly alarm. As an example, the presence of an anomaly may be determined in situations where driver distraction is detected, driver drowsiness is detected, and / or the driver is not wearing a seat belt. In this way, by adjusting the type and / or intensity of alarms corresponding to anomalies, more intuitive and clear alarms can be communicated to the user, thereby maximizing the utilization of the DMS.

[0129] Figure 4 illustrates a method for determining whether or not to generate an abnormal alarm in a DMS according to one embodiment of the present disclosure.

[0130] In one embodiment, the computing device 100 may acquire an image including the driver inside the vehicle (410).

[0131] In one embodiment, the image (or first image) can represent a target image for determining an abnormal alarm, as an image that includes the driver.

[0132] In one embodiment, the image means an image acquired from a camera. In one embodiment, the image may mean an image of the driver taken by a camera installed inside the vehicle. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to not wearing a seat belt. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to distraction. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to closed eyes. In one embodiment, the image may correspond to a target image for determining an alarm corresponding to drowsiness.

[0133] In this disclosure, the term "image" may be used to encompass one or more frames. For example, one image may correspond to one frame. For example, an image may be a still image and may correspond to a frame obtained from a video. For example, an image may include a collection of frames taken multiple times.

[0134] In one embodiment, the image may correspond to an image or frame taken at a specific point in time. In one embodiment, the image may mean a still image or frame extracted or captured from a video of a driver taken by a camera.

[0135] In one embodiment, the computing device 100 may acquire multiple images or videos acquired by a camera. The computing device 100 may acquire or extract specific images (e.g., specific frames) from the acquired images or videos in order to determine an anomaly or to trigger an anomaly alarm. For example, the images subject to an anomaly or anomaly alarm determination may be selected or determined frames from among multiple frames. In such an example, the computing device 100 may extract specific frames randomly from among multiple frames, or in units of a predetermined time period.

[0136] In one embodiment, the computing device 100 may use an artificial intelligence-based model to obtain model output information from an image (420).

[0137] In one embodiment, the model may correspond to a deep learning-based model that is pre-trained to take an image of a driver as input and output the probability of a predetermined object being present in the image of the driver. In such an embodiment, the model can output the probability of a seat belt being present in the image and / or the possibility of wearing a seat belt as model output information. In such an embodiment, the model can output the probability of eyes being closed and / or drowsiness in the image as model output information. In such an embodiment, the model can output the distance between the top and bottom of the eyes in the image as model output information. In such an embodiment, the model can output the probability of inattention to the road ahead and / or the driver's gaze information in the image as model output information.

[0138] In other embodiments, the model may correspond to a pre-trained deep learning-based model that takes an image of a driver as input and outputs the possibility that a predetermined anomaly that impedes safety is present in the image of the driver. The model may be pre-trained on a training dataset in which the presence or absence of anomalies in the images is labeled. The model can be pre-trained on a training dataset in which the presence or absence of anomalies in the images is labeled. The model can be pre-trained on a training dataset in which the location and / or type of anomaly in the images is labeled.

[0139] In one embodiment, the model output information may include a value that quantitatively indicates the likelihood of a predetermined object being present in the image. For example, the model output information may include a seat belt detection score in the image. For example, the model output information may include quantitative information related to seat belt detection in the image. For example, the model output information may include an eye-closing score in the image. For example, the model output information may include a score related to drowsiness in the image. For example, the model output information may include a score related to the distance between the upper and lower parts of the eyes in the image. For example, the model output information may include a score related to forward non-gaze or distraction in the image.

[0140] In one embodiment, the first model output information may include a value that quantitatively indicates the possibility of an abnormality in the first image. In one embodiment, the first model output information may include a value that indicates whether or not an abnormality exists in the first image. For example, the first model output information may include quantitative information related to the failure to detect a seat belt in the image. For example, the first model output information may include quantitative information related to the driver's eyes being closed in the image. For example, the first model output information may include quantitative information related to the driver's drowsiness in the image. For example, the first model output information may include quantitative information related to the driver's forward gaze or lack thereof in the image, or quantitative information related to the driver's distraction.

[0141] In one embodiment, a first image acquired from a camera may be input to an artificial intelligence-based model to obtain first model output information indicating the possibility of a predetermined object or predetermined action being present in the first image.

[0142] In one embodiment, model output information may mean the output of the model. For example, the model may generate model output information that quantitatively indicates the possibility of an anomaly (e.g., not wearing a seat belt, distraction, closed eyes, and / or drowsiness) in the first image. For example, the model may generate model output information that indicates whether or not an anomaly is present in the first image. For example, the model may generate model output information that quantitatively represents the possibility of a predetermined object or predetermined action in the first image, such as a seat belt, closed eyes, drowsiness, smoking, looking forward, and / or not looking forward.

[0143] In one embodiment, the model output information may represent the result of applying post-processing to the model output.

[0144] In one embodiment, the model output information may include the result of processing the model output. For example, if the model generates a bounding box corresponding to an anomaly or a specific object and / or an output indicating the possibility of an anomaly in the bounding box or the possibility of corresponding to a specific object, the computing device 100 may, by post-processing the model output, generate model output information indicating the possibility of an anomaly or the presence or absence of an anomaly, or the detectability of a specific object. For example, if the model outputs a result related to the detection of a seat belt in an image, the computing device 100 may generate a result from the result regarding whether the driver is wearing a seat belt or not. For example, if the model outputs a result corresponding to a face, the computing device 100 may calculate yaw and / or pitch values ​​from the result corresponding to the face, and based on the calculated yaw and pitch values, generate model output information indicating the presence or absence of distraction or the possibility of distraction. For example, if the model outputs a result of detecting an object in a first space, the computing device 100 may generate model output information indicating the presence or absence of drowsiness or the possibility of drowsiness by generating a conversion result that converts the result to a converted space.

[0145] In one embodiment, the model output information may be determined for each image (for example, each frame).

[0146] In one embodiment, the computing device 100 may compare the first model output information with a first threshold to obtain a first anomaly prediction result indicating whether or not an anomaly exists in the first image (430).

[0147] In one embodiment, the computing device 100 can compare first model output information with a predetermined threshold for anomaly detection. In one embodiment, the first model output information may include a quantitative value for comparison with the threshold. In one embodiment, the threshold may represent a variable threshold that can be changed for each of the images being compared. In one embodiment, the threshold may be dynamically changed based on other information.

[0148] In one embodiment, the computing device 100 may obtain a first anomaly prediction result indicating whether or not an anomaly exists in the first image, based on the results of the comparison. For example, the first anomaly prediction result may be obtained by comparing a quantitative value obtained by processing the output of the model or the output of the model with a threshold. The first anomaly prediction result may have a value indicating whether or not an anomaly exists in the image in question. For example, the first anomaly prediction result may have a value of 1 if an anomaly exists and a value of 0 if an anomaly does not exist, and the opposite situation is also possible depending on the implementation. In such an example, if the first model output information has a value of 0.7 and the threshold has a value of 0.6, the first anomaly prediction result may be set to have a value of 1, indicating that an anomaly exists. If the first model output information has a value of 0.5 and the threshold has a value of 0.6, the first anomaly prediction result may be set to have a value of 0, indicating that an anomaly does not exist. In one embodiment, the first anomaly prediction result may be determined for each image (for example, each frame). For example, the primary prediction result for the first abnormality may be set to have a value of 1 if the seat belt is not detected, and a value of 0 if the seat belt is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if eye closure is not detected, and a value of 1 if eye closure is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if drowsiness is not detected, and a value of 1 if drowsiness is detected. For example, the primary prediction result for the first abnormality may be set to have a value of 0 if forward inattention is not detected, and a value of 1 if forward inattention is detected. In such examples, the first model output information may be set to have a higher value the more likely it is that the seat belt will not be detected. In such examples, the first model output information may be set to have a higher value the more likely it is that eye closure, drowsiness, and / or forward inattention will be detected.

[0149] The technology according to one embodiment of the present disclosure can ensure the accuracy and reliability of the model output by dynamically adjusting the thresholds compared with the model output information in various ways.

[0150] In one embodiment, the first threshold may be determined based on at least one primary prediction result of a previous anomaly corresponding to at least one previous image acquired before the first image. In one embodiment, the first threshold compared with the first model output information may be determined based on second model output information generated by the model in response to a second image acquired before the first image. In one embodiment, the first threshold compared with the first model output information may be modified based on the threshold for anomaly determination of the second image acquired immediately before the first image, based on the primary prediction result of a previous anomaly or the previous model output information. The current threshold corresponding to the current image may be variable based on a comparison between the previous image and the previous threshold. For example, if the model output information of the previous image exceeds the previous threshold, the current threshold corresponding to the current image may be determined to a lower threshold among several threshold options or decreased from the previous threshold. For example, if the model output information of the previous image does not exceed the previous threshold, the current threshold corresponding to the current image may be determined to a higher threshold among several threshold options or increased from the previous threshold.

[0151] In other embodiments, the first threshold, which is compared with the first model output information, may be determined based on a comparison between a primary prediction result of an anomaly obtained from a previous image and a specific threshold. The specific threshold can be considered as a threshold for determining other thresholds. In such embodiments, in determining the current threshold for the current image, the computing device 100 may compare the ratio value of 1 with the specific threshold and the ratio value of 0 with the specific threshold for a plurality of previous primary prediction results of anomalies (e.g., having values ​​of 0 or 1) corresponding to a plurality of previous images, to determine which primary prediction result exceeds the specific threshold. If the primary prediction result exceeding the specific threshold is 1, the computing device 100 may decide to set the threshold to a lower value among a plurality of threshold options, or to decrease the threshold relative to a previous threshold. If the primary prediction result exceeding the specific threshold is 0, the computing device 100 may decide to set the threshold to a higher value among a plurality of threshold options, or to increase the threshold relative to a previous threshold. If the primary prediction result exceeding the specific threshold is neither 0 nor 1, the computing device 100 may decide to maintain the threshold.

[0152] In one embodiment, the threshold change range may be predetermined. The threshold can be sequentially increased, sequentially decreased, or maintained in a counter-like manner based on the predicted results of previous images. The increase or decrease of the threshold can be performed within the threshold change range.

[0153] In other embodiments, the first threshold, which is compared with the first model output information, may be determined based on the secondary prediction result of anomalies in the previous image. In such embodiments, if the secondary prediction result of anomalies in the previous image is 1, the threshold for the current image may be decreased compared to the previous threshold, or set to a lower value among the multiple threshold options. If the secondary prediction result of anomalies in the previous image is 0, the threshold for the current image may be increased compared to the previous threshold, or set to a higher value among the multiple threshold options. In one embodiment, the expression "determined based on specific information" may include the fact that the threshold can be quantitatively modified in magnitude depending on the magnitude and / or type of value of the specific information. The computing device 100 may acquire multiple images over time. A primary prediction result of an anomaly may be acquired for each of these multiple images. For example, a primary prediction result of a second anomaly corresponding to a previously acquired second image may be acquired for the first image, and a primary prediction result of a first anomaly corresponding to the first image may be acquired. In such an example, the first threshold used to acquire the primary prediction result of the first anomaly corresponding to the first image may be determined based on the primary prediction result of the second anomaly corresponding to the second image. For example, the first threshold may be changed in accordance with the primary prediction result of the second anomaly. For example, the first threshold may be determined to a preset value such as 0.3 if the primary prediction result of the second anomaly is 1, and 0.5 if the primary prediction result of the second anomaly is 0.

[0154] In one embodiment, the variable reference value here may be a threshold corresponding to the second image acquired immediately before the first image. For example, if the primary prediction result for the second anomaly includes the result that an anomaly exists, the first threshold corresponding to the first image following the second image may be set to increase compared to past thresholds (for example, the threshold corresponding to the second image). When the magnitude of the threshold increases, it may be determined that an anomaly exists in the first image if the quantitative value of the model output information corresponding to the image is relatively high. For example, if the primary prediction result for the second anomaly includes the result that an anomaly does not exist, the first threshold corresponding to the first image following the second image may be set to decrease compared to past thresholds. When the magnitude of the threshold decreases, it may be determined that an anomaly exists in the first image even if the quantitative value of the model output information corresponding to the image is relatively low.

[0155] In other embodiments, the first threshold, which is compared with the first model output information, may be modified based on the second model output information corresponding to a previously acquired second image of the first image. For example, the first threshold may have a negative correlation with the value of the model output information of the previous image. In such an example, if the model output information of the previous image is relatively large, the magnitude of the first threshold may be set to be small. If the model output information of the previous image is relatively small, the magnitude of the first threshold may be set to be large.

[0156] In one embodiment, the first threshold may be determined based on a primary prediction result of a second anomaly obtained by comparing a second threshold, which is determined based on a primary prediction result of a third anomaly corresponding to a previously acquired third image of the second image, with second model output information. Here, the second threshold may be a threshold for determining whether or not an anomaly exists in the second image.

[0157] Thus, the technology according to one embodiment of the present disclosure can dynamically adjust the threshold for anomaly prediction or judgment of the current image for each of a plurality of sequential images by utilizing anomaly-related results corresponding to previous images. This adjustment of the threshold can increase the accuracy and / or confidence of the model's output.

[0158] In one embodiment, the first threshold may be determined based on the ratio of result values ​​indicating the presence of an anomaly from previous primary prediction results of anomalies corresponding to a predetermined first number of previously acquired images of the first image. For example, if the ratio of result values ​​indicating the presence of an anomaly in the previous primary prediction results of anomalies (i.e., multiple primary prediction results of anomalies) is greater than or equal to a first ratio, the first threshold may be set to a first value; if the ratio of result values ​​indicating the presence of an anomaly from the previous primary prediction results of anomalies is less than the first ratio, the first threshold may be set to a second value higher than the first value.

[0159] In one embodiment, the first threshold may be determined or modified depending on what the majority value is among the result values ​​indicating the presence or absence of an anomaly from the previous primary prediction results of anomalies corresponding to a predetermined first number of previously acquired images of the first image. For example, if the majority value of the previous primary prediction results of anomalies is a value indicating the presence of an anomaly, the first threshold corresponding to the current image may be set to decrease. For example, if the mainstream value of the previous primary prediction results of anomalies is a value indicating the absence of an anomaly, the first threshold corresponding to the current image may be set to increase.

[0160] In one embodiment, the model output information corresponding to each of the sequentially acquired images may be structured in the form of a queue. For example, one queue may consist of a predetermined number of units. For example, a first unit in one queue may be assigned first model output information corresponding to the first image, a second unit may be assigned second model output information corresponding to the second image, and a third unit may be assigned third model output information corresponding to the third image.

[0161] In one embodiment, the primary prediction results of anomalies corresponding to each of a plurality of sequentially acquired images may be structured in the form of a queue. Each value corresponding to a single image may be assigned to the same position in the plurality of queues. The queue composed of primary prediction results of anomalies and the queue composed of model output information may have corresponding positions for a single image. For example, the first image acquired at a first time point may be assigned to the same position in the first queue composed of model output information and the second queue composed of primary prediction results of anomalies.

[0162] In one embodiment, a queue may consist of a predetermined number of units. For example, a predetermined number of units in a queue may contain data corresponding to images acquired over time. For example, a first unit in a queue may be assigned the first predicted result of an anomaly corresponding to a first image, a second unit the first predicted result of an anomaly corresponding to a second image, and a third unit the first predicted result of an anomaly corresponding to a third image. In such an example, the threshold to be compared with the model output result corresponding to the currently acquired image may be determined based on the ratio of result values ​​indicating the presence of anomalies among the first, second, and third predicted results of anomalies contained in a queue. For example, if both the first and second predicted results of anomalies include result values ​​indicating the presence of anomalies, and the third predicted result includes result values ​​indicating the absence of anomalies, the ratio of result values ​​may be (2 / 3 × 100). By comparing such a ratio of result values ​​with a specific threshold, the threshold corresponding to the current image can be determined or modified. The ratio indicating the presence of anomalies in previous anomaly prediction results and the magnitude of the threshold corresponding to the current image may have a negative correlation. In this disclosure, the expression that the first value and the second value have a negative correlation means that as the first value increases, the second value tends to decrease. In this disclosure, the expression that the first value and the second value have a negative correlation may also mean that as the first value becomes relatively larger, the second value becomes relatively smaller.

[0163] In one embodiment, the computing device 100 may obtain a secondary prediction result for the first anomaly by performing a first voting using the primary prediction result for the first anomaly (440).

[0164] In one embodiment, the secondary prediction result of the first anomaly may represent a quantitative value used as a parameter for determining an anomaly alarm corresponding to the first image. In another embodiment, the secondary prediction result of the first anomaly may have a value indicating the presence or absence of an anomaly, and is used as a parameter for determining whether or not to generate an anomaly alarm corresponding to the first image.

[0165] In another embodiment, the first threshold used for anomaly prediction of the first image may be determined based on a secondary prediction result of a second anomaly corresponding to a previously acquired second image of the first image.

[0166] In one embodiment, the first voting may be used to correct the accuracy and / or reliability of the model's output. In one embodiment, the first voting may utilize the first predicted results of anomalies in previously acquired images of the first image and the first predicted results of anomalies in the currently acquired first image.

[0167] In one embodiment, the first voting may include a process of determining the majority value of the primary predicted results of anomalies corresponding to sequential images. In one embodiment, the first voting may include a process of determining the ratio of primary predicted results of anomalies corresponding to sequential images. In one embodiment, the first voting may include a process of comparing the ratio of primary predicted results of anomalies corresponding to sequential images with a predetermined threshold.

[0168] In one embodiment, the first voting may utilize values ​​on a voting queue composed of multiple units corresponding to sequential images. These values ​​are primary prediction results for anomalies corresponding to sequential images. For example, the first unit in the voting queue may be assigned the primary prediction result for the first anomaly corresponding to the first image, the second unit may be assigned the primary prediction result for the second anomaly corresponding to the second image, and the third unit may be assigned the primary prediction result for the third anomaly corresponding to the third image. Here, the first image may correspond to the most recently acquired image or the current image, the second image may be an image acquired prior to the first image, and the third image may be an image acquired prior to the second image.

[0169] In one embodiment, there may be multiple boating queues. For example, a first boating queue may include model output information acquired over time, a second boating queue may include primary prediction results acquired over time, and a third boating queue may include secondary prediction results acquired over time.

[0170] In one embodiment, the secondary prediction result for the first anomaly corresponding to the first image may be determined based on what the mainstream values ​​of the primary prediction result for the first anomaly, the primary prediction result for the second anomaly, and the primary prediction result for the third anomaly are.

[0171] In one embodiment, the secondary prediction result for the first anomaly corresponding to the first image may be determined based on a comparison between the ratio of result values ​​representing anomalies among the primary prediction results for the first anomaly, the second anomaly, and the third anomaly, and a specific threshold. For example, if the primary prediction result for the first anomaly indicates the presence of an anomaly, the primary prediction result for the second anomaly indicates the presence of an anomaly, and the primary prediction result for the third anomaly indicates the absence of an anomaly, the secondary prediction result for the first anomaly corresponding to the first image may be set to the mainstream value of the three primary prediction results for anomalies or their representative value indicating the presence of an anomaly (e.g., 1). This allows the first unit corresponding to the first image in the queue composed of secondary prediction results for anomalies to have a value of 1. For example, if the primary prediction result for the first anomaly indicates the absence of an anomaly, the primary prediction result for the second anomaly indicates the presence of an anomaly, and the primary prediction result for the third anomaly indicates the absence of an anomaly, the secondary prediction result for the first anomaly corresponding to the first image may be set to the mainstream value of the three primary prediction results for anomalies or their representative value indicating the absence of an anomaly (e.g., 0). As a result, the first unit corresponding to the first image in the queue, which is composed of secondary prediction results for anomalies, can have a value of 0.

[0172] In one embodiment, the secondary anomaly prediction results corresponding to each of a plurality of sequentially acquired images may be structured in the form of a voting queue. Each value corresponding to a single image may be assigned to the same position in the plurality of queues. The queue composed of secondary anomaly prediction results, the queue composed of primary anomaly prediction results, and the queue composed of model output information may have positions corresponding to each other for a single image. For example, the values ​​related to the anomaly prediction of a first image acquired at a first time point may be assigned to the same position (e.g., corresponding positions) in the queue composed of model output information, the queue composed of primary anomaly prediction results, and the queue composed of secondary anomaly prediction results, respectively.

[0173] In one embodiment of the present disclosure, the first voting may generate a predicted result of a set of anomalies that represents an image group consisting of a first image and a predetermined second number of previously acquired images of the first image, in order to ensure the accuracy of the first model output information. Such a predicted result of a set of anomalies may represent a result that represents the primary predicted result of anomalies corresponding to a plurality of images, including the first image.

[0174] In one embodiment, the computing device 100 may determine the majority value of the primary prediction result of the anomaly corresponding to the first image and a predetermined second number of previously acquired images of the first image, and use the determined majority value to generate the secondary prediction result of the first anomaly. For example, the majority value may be determined as the result value that accounts for a higher proportion of the primary prediction result of the anomaly, among the result value indicating the presence of the anomaly and the result value indicating the absence of the anomaly.

[0175] As another example, the mainstream value may be determined by comparing the result value present in the primary prediction of anomalies with a predetermined second threshold. For example, if the second threshold is 45% and the proportion of a particular result value in the primary prediction of anomalies is 50%, the secondary prediction of anomalies may be set to that particular result value.

[0176] As described above, the technology according to one embodiment of the present disclosure can further improve the reliability and accuracy of the anomaly determination result by performing a first voting that utilizes the primary prediction result of the anomaly.

[0177] In one embodiment, the secondary prediction result for an anomaly may include a quantitative value used as a parameter for determining an anomaly alarm corresponding to the acquired image. For example, the secondary prediction result for an anomaly may include a counter value. For instance, if the secondary prediction result for a second anomaly corresponding to a second image acquired previously for the first image has a value of 2, it may be determined that a value of 1 is added according to the result of the first voting corresponding to the first image. In this case, a value of 2+1=3 may be included in the secondary prediction result for the first anomaly corresponding to the first image. In another example, if the secondary prediction result for a second anomaly corresponding to a second image acquired previously for the first image has a value of 2, it may be determined that a value of 1 is decreased according to the result of the first voting corresponding to the first image. In such a case, a value of 2-1=1 may be included in the secondary prediction result for the first anomaly corresponding to the first image.

[0178] In one embodiment, the unit of the counter value to be increased or decreased may be determined based on the difference between the image acquisition times. For example, if the counter value corresponding to the second image is 0.7, and the first image is acquired 500ms or more after the acquisition time of the second image, and the result of the first voting corresponding to the first image is determined to add a value, the counter value corresponding to the first image may be set to 0.7 + 0.5 = 1.2.

[0179] In one embodiment, the secondary prediction result of the first anomaly may be compared with a predetermined counter threshold. For example, if the counter value corresponding to the secondary prediction result of the first anomaly is greater than or equal to the predetermined counter threshold, the computing device 100 may decide to generate an alarm (e.g., turn the alarm ON). For example, if the counter value corresponding to the secondary prediction result of the second anomaly is greater than or equal to the counter threshold, and the counter value corresponding to the secondary prediction result of the first anomaly changes to less than the counter threshold, the computing device 100 may decide to turn the alarm OFF.

[0180] In one embodiment, the secondary prediction result for the first anomaly may include multiple counters. Including multiple counters in this way may generate multiple types of alarms. For example, a first counter and a second counter may be included in the secondary prediction result for the first anomaly. As a result of the first voting, the values ​​corresponding to the first and second counters can be changed independently. Each of the first and second counters is compared with pre-assigned first and second counter thresholds, and if the counter value is greater than or equal to the counter threshold, an alarm corresponding to each counter may be generated. As an example, each counter may have a pre-defined minimum and maximum range, and if it falls outside the minimum and maximum range according to the result of the first voting, the counter value may be set to have a minimum and maximum range.

[0181] In one embodiment, the secondary prediction result of the first anomaly may be configured to generate multiple alarms by being compared with a plurality of counter thresholds. For example, the first counter value corresponding to the first image may be compared with a first counter threshold and a second counter threshold, respectively. If any one of these counter thresholds is met, a first alarm corresponding to that counter threshold may be generated. Furthermore, if any other of the counter thresholds is met, a second alarm corresponding to that counter threshold may be generated.

[0182] In one embodiment, the computing device 100 may determine an anomaly alarm corresponding to the first image by performing a second voting using the secondary prediction result of the first anomaly (450).

[0183] In this disclosure, the first voting may utilize the primary prediction result of a previous anomaly corresponding to at least one previously acquired image of the first image subject to anomaly determination, and the primary prediction result of a first anomaly corresponding to the first image. The second voting may utilize the secondary prediction result of a previous anomaly corresponding to at least one previously acquired image, and the secondary prediction result of a first anomaly corresponding to the first image.

[0184] In one embodiment, the second voting may be performed after the first voting. In one embodiment, the second voting may utilize the results of the first voting.

[0185] In one embodiment, the second voting can determine whether the abnormality detection results for a predetermined number of images are continuous. In one embodiment, if the abnormality detection results for a predetermined number of images are continuous with values ​​indicating the presence of an abnormality, the second voting may decide to generate an abnormality alarm.

[0186] In one embodiment, the second voting may determine a set of images consisting of a first image corresponding to the current image and previously acquired sequential images of the first image, in order to ensure accuracy in the generation of anomaly alarms. The second voting may determine whether or not there is continuity in the secondary prediction results of anomalies corresponding to the images constituting the set of images. The second voting is a process that utilizes whether or not there is continuity in the secondary prediction results of anomalies. For example, if all secondary prediction results of anomalies in the set of images consisting of sequential images including the first image indicate the presence of an anomaly, the computing device 100 may decide to generate an anomaly alarm corresponding to the first image. For example, if some of the secondary prediction results of anomalies in the set of images consisting of sequential images including the first image indicate the presence of an anomaly, and other parts indicate the absence of an anomaly, the computing device 100 may decide that there is no continuity in the secondary prediction results of anomalies.

[0187] For example, suppose the number of criteria for determining continuity is 3. Under this assumption, if the secondary prediction results for the first anomaly, the second anomaly, and the third anomaly corresponding to the three sequential images, including the first image currently being judged for anomaly, do not indicate the presence of anomalies, the result value of the anomaly alarm corresponding to the first image may be set to 0 via the second voting. In such an example, no anomaly alarm is generated for the first image currently being acquired.

[0188] In one embodiment, the second voting may include comparing a secondary prediction result of an anomaly, including a counter value, with a counter threshold. For example, suppose the counter threshold is 3. In such an example, when the secondary prediction result of an anomaly reaches 3, the alarm corresponding to the counter threshold can be turned ON. Also, when the secondary prediction result of an anomaly changes from 3 to 2, the alarm can be turned OFF. As described above, the second voting may be performed by comparing each of several counter values ​​with each of the counter thresholds. The second voting may be performed by comparing each of a single counter value with several counter thresholds.

[0189] As described above, the technology according to one embodiment of the present disclosure may utilize one or more boats to generate alarms in response to anomalies. By utilizing one or more boats, the sensitivity, accuracy, and reliability of alarms in response to anomalies can be increased.

[0190] Figure 5 illustrates an exemplary method for determining a driver's drowsiness state according to one embodiment of the present disclosure.

[0191] In one embodiment, the computing device 100 may acquire an image including the driver's face inside the vehicle (510).

[0192] In one embodiment, the image refers to an image acquired from a camera. In another embodiment, the image may refer to an image of the driver taken by a camera installed inside the vehicle.

[0193] In one embodiment, the image may correspond to an image or frame taken at a specific point in time. In one embodiment, the image may mean a still image or frame extracted or captured from a video of a driver taken by a camera.

[0194] In one embodiment, the computing device 100 may acquire multiple images or videos acquired by a camera. The computing device 100 may acquire or extract specific images (e.g., specific frames) from the acquired images or videos in order to determine drowsiness or provide a drowsiness alarm. For example, the images subject to drowsiness determination or drowsiness alarm determination may be selected or determined frames from among multiple frames. In such an example, the computing device 100 may extract specific frames randomly from among multiple frames, or in units of a predetermined time period.

[0195] In one embodiment, the computing device 100 may perform detection of target points that define the driver's eye region in a first space corresponding to the acquired image to determine whether or not the driver is drowsy (520).

[0196] In one embodiment, the first space may mean the space contained in the acquired image. In one embodiment, the first space may mean a space having a coordinate system corresponding to the acquired image.

[0197] In one embodiment, the eye region may mean a bounding box containing the driver's eyes in the image. In one embodiment, the eye region may mean a segmentation region that defines the shape or contour of the driver's eyes in the image. In one embodiment, the eye region may include one or more feature points or landmarks for identifying the driver's eyes in the image.

[0198] For example, the eye region and / or target point may be detected by the detection model.

[0199] In one embodiment, target points are feature points contained within the driver's face, and a set of target points can define the eye region. For example, target points may mean feature points that define the shape or contour of the driver's eyes. For example, target points may be used to define the boundary of the eye region for identifying the eyes within the face. For example, by concatenating target points, the eye region of the driver may be defined.

[0200] In one embodiment, the computing device 100 may use a model to detect target points that define the eye region. For example, the model may correspond to an artificial intelligence-based model that utilizes heatmap regression. The model is a model that detects target points in the eye region and can utilize a heatmap in the process of regression to the target points.

[0201] In one embodiment, the computing device 100 can determine whether the driver has their eyes closed by utilizing the detection result of a target point in the eye area. Based on such a determination of whether the eyes are closed, the computing device 100 may determine whether the driver is drowsy. For example, if the driver is determined to be drowsy (or has their eyes closed) in the first image, and is also determined to be drowsy (or has their eyes closed) in a predetermined number of subsequent images following the first image, the computing device 100 may determine that the driver is drowsy (or has their eyes closed). As an example, based on the determination that the driver is drowsy (or has their eyes closed), the computing device 100 may decide whether or not to provide the driver with a drowsiness alarm.

[0202] In one embodiment, the computing device 100 can transform a first space corresponding to a first image for more accurate detection of the eye region and for more accurate determination of eye closure. Based on the eye region in the transformed space, the computing device 100 can normalize the eye region so that, because one embodiment of the present disclosure utilizes spatial transformation, it can accurately sense the eye region even when input images of faces of various sizes taken at various angles and distances.

[0203] In one embodiment, the computing device 100 may determine first position information corresponding to a first set of points within the driver's eye area in a first space corresponding to an image (530).

[0204] In one embodiment, the target points may consist of a first set of points and a second set of points. The first set of points may represent points that serve as a reference for coordinate transformation or spatial transformation. The first set of points may also represent feature points that serve as a reference when determining the eye region. For example, the first set of points may include the left end point and the right end point in the eye region. For example, the number of points in the first set may correspond to the number of first position information points. For example, if there are eight points in the first set, eight first position information points corresponding to the points in the first set may be determined. The first position information may represent position information (e.g., coordinate values) in a first space corresponding to the points in the first set. The second set of points may represent the remaining points of the target points that correspond to or define the eye region, excluding the points in the first set. For example, the second set of points may include the remaining points of the target points excluding the left end point and the right end point.

[0205] In one embodiment, the computing device 100 may use a first set of points and first position information to determine a transformation matrix for positioning target points that define the eye region in a first space onto a predetermined transformation space (540).

[0206] In one embodiment, the transformation matrix may mean a matrix for performing a coordinate transformation from the first space to the transformed space. In one embodiment, the transformation matrix may mean a matrix for positioning a target point in the first space in the transformed space, using a first set of points (e.g., the leftmost and rightmost points) as a reference axis in the transformed space. The transformation matrix may perform scaling, rotation, translation, shearing, and / or reflection to position the first set of points in the transformed space.

[0207] In other embodiments, the first set of points may include a first and second point that form the longest first straight line among the straight lines that can be formed by connecting two of the target points. In such embodiments, the transformation matrix can transform the positions of the target points such that the center point of the first straight line is located at the center point in the transformation space.

[0208] In one embodiment, the computing device 100 may use a transformation matrix to determine transformed position information corresponding to a target point in the transformation space (550).

[0209] In one embodiment, the computing device 100 can apply a transformation matrix to a second set of points to position the second set of points from the first space to the transformation space. The computing device 100 may determine the transformed position information corresponding to a target point by obtaining second transformed position information corresponding to the second set of points located in the transformation space.

[0210] In one embodiment, the computing device 100 can apply a transformation matrix to a first set of points and a second set of points to position the points of the first set from the first space to the transformation space, and position the points of the second set from the first space to the transformation space. The computing device 100 may determine the transformed position information corresponding to the target point by obtaining first transformed position information corresponding to the points of the first set located in the transformation space, and second transformed position information corresponding to the points of the second set located in the transformation space.

[0211] In the aforementioned method, the computing device 100 can determine a transformation matrix using a first set of points corresponding to the eye region, and use the transformation matrix to transform the eye region from the coordinate system corresponding to the image to a normalized coordinate system. When an image containing the eye region transformed to the normalized coordinate system is input to an artificial intelligence-based model, the accuracy of the output of the artificial intelligence-based model (e.g., the result of eye region detection and / or eye closure detection) can be increased.

[0212] In one embodiment, the computing device 100 may determine the driver's drowsiness state from the image based on the converted location information (560).

[0213] In one embodiment, the computing device 100 may determine from an image whether a driver is drowsy by utilizing a first artificial intelligence-based model that generates output data indicating whether or not the driver's eyes are closed in response to input data including a target point having transformed location information. For example, the first model may correspond to a pre-trained model using a training dataset in which facial images are labeled with whether or not the driver's eyes are closed. For example, the first model may correspond to an artificial intelligence-based model that generates output data indicating whether or not the driver's eyes are closed in response to input data including a target point having transformed location information.

[0214] In one embodiment, the computing device 100 may determine from an image whether the driver has their eyes closed by utilizing an artificial intelligence-based first model that generates output data indicating whether or not the driver has their eyes closed in response to input data including a target point having converted location information.

[0215] Furthermore, the computing device 100 may use an artificial intelligence-based second model to obtain a heatmap from the first image corresponding to points within the first image, and use the heatmap to decide whether or not to perform detection of target points that define the eye region in the first image, or whether or not the driver's eyes are closed from the first image. For example, the computing device 100 may calculate the standard deviation for the values ​​corresponding to each point in the heatmap, and if the standard deviation exceeds a pre-stored critical criterion, it may decide to use the second model to perform detection of target points from the first image, or decide from the first image that the driver's eyes are not closed.

[0216] In one embodiment, the computing device 100 may determine the size of the transformed eye region formed by the target point in the transformed space based on the transformed position information. Based on the size of the transformed eye region, the computing device 100 may determine from the input image whether or not the driver is drowsy. For example, if the size of the eye region is smaller than a predetermined critical size, the computing device 100 can determine from the image that the driver has their eyes closed. For example, if the computing device 100 determines that the driver has their eyes closed, it can determine that the driver is drowsy (or has their eyes closed) in the image. For example, if the computing device 100 determines from a series of sequential images that the driver has their eyes closed, it can determine that the driver is currently drowsy.

[0217] In another embodiment, the computing device 100 may also determine the driver's drowsiness state by utilizing an additional artificial intelligence-based model that is pre-trained to take an image including the transformed eye region as input and output whether or not the eyes are closed from the input image. The additional artificial intelligence-based model may, for example, correspond to a detection model.

[0218] The results of eye-closed state measurements using landmarks in the eye region may vary depending on the size and / or angle of the face on the camera. Although the eyes are actually the same size, the size of the eyes in the image acquired by the camera may differ from the actual size of the eyes depending on the driver's posture and movement. Therefore, as mentioned above, the positional information of the eye region in the normalized space (i.e., transformed space) may be determined by determining a transformation matrix that performs a coordinate system transformation on a predetermined normalized space using the endpoints of both eyes as the reference axis, and applying the determined transformation matrix to the remaining points of the eye region. Using landmarks of the eye region transformed into a normalized space can ensure consistency and accuracy in eye-closed state measurements.

[0219] In one embodiment, the computing device 100 may determine the driver's drowsiness state using a pre-trained second model that outputs a heatmap corresponding to points (e.g., feature points) in an input image. For example, the second model may correspond to an artificial intelligence-based model that is pre-trained to detect eye regions in an image using heatmap regression. For example, the second model is a model that detects target points in the eye region and can utilize a heatmap when executing a regression algorithm on the target points. In such an embodiment, the computing device 100 may use the second model to detect target points corresponding to the eye region. The computing device 100 may use the heatmap output of the second model to determine whether or not the eyes are closed, whether or not sunglasses are being worn, whether or not to perform target point detection, whether or not to provide an eye-closing alarm, and / or whether or not to perform a transformation to a transformation space. Depending on the implementation, the first and second models described above may be the same model or different models.

[0220] Details of embodiments utilizing heatmaps will be described later in Figures 6, 7, and 10.

[0221] Figure 6 illustrates an example of using a heatmap in the process of detecting the eye region in a DMS according to one embodiment of the present disclosure.

[0222] In one embodiment, the computing device 100 may acquire an image including the driver's face inside the vehicle (610).

[0223] Images containing the driver's face in this disclosure may include, for example, an image in which the driver's face is shown as a bounding box, an image in which the contour of the driver's face is segmented, an image containing the driver's face, and / or an image in which the feature points of the driver's face are shown.

[0224] In one embodiment, the computing device 100 may use an artificial intelligence-based second model to obtain a heatmap from the image corresponding to points within the image (620).

[0225] In one embodiment, the second model may correspond to an artificial intelligence-based model pre-trained to detect eye regions within an image using heatmap regression. For example, the second model is a model for detecting target points in eye regions, and can utilize a heatmap when executing a regression algorithm on the target points. The second model may operate to output a heatmap representing the probability of existence for each feature point in a face, for example, to detect facial feature points or facial landmarks. Using such a heatmap, facial feature points (e.g., target points defining eye regions) may be determined based on the probability of existence of each feature point in the face. The heatmap obtained from the second model can be used to decide whether or not to perform target point detection, and if it is decided to perform target point detection, it can be used to detect the target points.

[0226] In such embodiments, the computing device 100 may use the second model to detect target points corresponding to the eye area. The computing device 100 may use the heatmap, which is the output of the second model, to determine whether or not the eyes are closed, whether or not to detect target points, whether or not sunglasses are being worn, whether or not to provide an eye-closing alarm, and / or whether or not to perform a conversion to a transformation space. For example, the output of the second model can be used as a filter to determine whether or not to perform processes related to detecting whether or not the eyes are closed and / or providing an eye-closing alarm.

[0227] In one embodiment, the computing device 100 uses a heatmap to determine whether or not to perform detection of target points that define the eye region from the image (630).

[0228] In one embodiment, the computing device 100 can calculate the standard deviation of the values ​​corresponding to each point in the heatmap. In one embodiment, the computing device 100 may decide to use the second model to perform detection of the target point from the first image if the standard deviation exceeds a pre-stored critical criterion. If the input image is determined to be different from the learned trend, the standard deviation of the heatmap output values ​​of the pre-trained artificial intelligence-based second model may be relatively large. Thus, the technology according to one embodiment of the disclosure may determine whether an area determined by the model to be an eye area is actually an eye area by calculating the magnitude of the standard deviation of the values ​​in the heatmap. For example, a large standard deviation value of the heatmap may indicate a high probability of occlusion occurring in the image.

[0229] For example, a DMS using an RGB camera may not be able to see through sunglasses, and therefore sunglasses may appear black in the image. Models that utilize landmarks in the eye region tend to be edge-detecting, and in such situations, the eye region may be recognized as the boundary of the sunglasses. This increases the likelihood that if a driver is wearing sunglasses, it will be determined that their eyes are closed. In this way, to prevent misidentification of other objects such as sunglasses and / or eyeglasses as closed eyes, the technology according to one embodiment of the disclosure can use a model for landmark detection that uses heatmap regression. The technology according to one embodiment of the disclosure can determine that regions where the standard deviation of the heatmap values ​​of such a model is greater than a certain threshold correspond to regions of sunglasses or similar objects. The technology according to one embodiment of the disclosure can determine that regions where the standard deviation of the heatmap values ​​of such a model is greater than a certain threshold are not the eye region.

[0230] Figure 7 illustrates an example of a method for determining a driver's drowsiness level using a heat map in a DMS according to one embodiment of the present disclosure.

[0231] In one embodiment, the heatmap obtained from the model may perform a filtering function in a process to determine whether or not the driver is drowsy.

[0232] In one embodiment, the computing device 100 can compare the standard deviation of the values ​​in the heatmap corresponding to points in the image with a first threshold (710). Here, the first threshold may mean a reference value compared with the standard deviation. The computing device 100 can determine that areas where the standard deviation is greater than a certain threshold are corresponding areas that are not eye areas. The computing device 100 can calculate the standard deviation of the heatmap corresponding to the areas determined by the model to be eye areas and determine whether the calculated standard deviation exceeds the first threshold.

[0233] In one embodiment, the computing device 100 may determine whether the driver in the image is drowsy by comparing a blindness score determined using the converted position information with a second threshold, provided that the standard deviation is less than or equal to a first threshold (720).

[0234] In one embodiment, the computing device 100 can use the output of a model (e.g., a heatmap) to detect eye regions and calculate an eye-closed score corresponding to the eye regions, provided that the standard deviation is less than or equal to a first threshold. In another example, the computing device 100 can use points corresponding to eye regions obtained from the model to calculate the eye-closed score. For example, the computing device 100 can determine whether the driver has their eyes closed based on the eye-closed score. Based on such a determination, it may be determined whether the driver is drowsy.

[0235] For example, the computing device 100 may detect eye regions from an image via a model. The computing device 100 can use a heatmap corresponding to the eye region obtained during the eye region detection process to determine whether the detected eye region matches the actual eye region. The process of determining whether the detected eye region is the actual eye region may be performed by calculating the standard deviation of the values ​​in the heatmap corresponding to the eye region.

[0236] For example, in one embodiment, the computing device 100 can determine a transformation matrix if the standard deviation is less than or equal to a first threshold, and transform the detected eye region into a normalized space using the transformation matrix. If the standard deviation of the heatmap values ​​is less than or equal to the first threshold, the computing device 100 can detect the eye region using the method illustrated in Figure 5, position target points corresponding to the eye region in the transformed space, and use the transformed position information to calculate an eye-closed score corresponding to the input image. The computing device 100 can use the eye-closed score to determine whether the driver in the image is drowsy or not. For example, the eye-closed score is a value that quantitatively represents the degree of eye-closing. Such an eye-closed score may be determined by using the position information of the points corresponding to the eye region (e.g., transformed position information) to calculate the size of the eye region using a rule-based algorithm. As another example, the eye-closed score may be calculated using an artificial intelligence-based model (the aforementioned model or another model).

[0237] In one embodiment, the computing device 100 may decide to generate a drowsiness alarm if it determines that the driver is drowsy.

[0238] In one embodiment, the computing device 100 may determine that the driver in the image is not drowsy if the standard deviation exceeds a first threshold (730).

[0239] In one embodiment, the computing device 100 may determine that there are other objects in the image besides the eye region if the standard deviation exceeds a first threshold.

[0240] In one embodiment, the computing device 100 may determine that there are no closed eyes in the image if the standard deviation exceeds a first threshold.

[0241] In one embodiment, the computing device 100 may determine that the detected eye region does not match the actual eye region if the standard deviation exceeds a first threshold.

[0242] In one embodiment, the computing device 100 may determine that the driver in the image is not drowsy without calculating an eye-closed score if the standard deviation exceeds a first threshold.

[0243] For example, computing device 100 can determine that sunglasses are present in an image if the standard deviation exceeds a first threshold.

[0244] For example, computing device 100 can determine that the detected eye area corresponds to sunglasses if the standard deviation exceeds a first threshold.

[0245] For example, computing device 100 may decide not to perform target point detection or target point transformation if the standard deviation exceeds a first threshold.

[0246] For example, computing device 100 may decide not to provide a blindfold alarm if the standard deviation exceeds a first threshold.

[0247] For example, computing device 100 may decide not to convert target points corresponding to the eye region into the transformation space if the standard deviation exceeds a first threshold.

[0248] Figure 8 illustrates an example of a method for determining whether or not to generate a driver drowsiness alarm according to one embodiment of the present disclosure.

[0249] In the explanations in Figure 8, any content that overlaps with the explanations in Figure 4 will be omitted to avoid duplication of explanations, and will be replaced by the explanations in Figure 4. For example, the methods related to the first and second boating in Figure 4 may correspond to the methods related to the first and second boating in Figure 8.

[0250] One embodiment of the technology described herein may determine from an image of the driver whether the driver is closed-eyed and / or drowsy, and decide whether or not to generate an alarm based on the determined result. Such a technology described herein can ensure the accuracy and reliability of alarms provided to the driver in the DMS.

[0251] The technology according to one embodiment of the present disclosure can utilize thresholds and / or voting in the process of determining whether the driver is closed-eyed and / or drowsy from an image of the driver. Such technology according to one embodiment of the present disclosure enables a more accurate determination of whether the driver is drowsy in the DMS.

[0252] In one embodiment, the computing device 100 may determine whether the driver is drowsy in the first image based on at least one of the standard deviation of values ​​in a heatmap corresponding to points in the first image to be judged, and a closed-eyes score determined using the transformed positional information of the first image.

[0253] In one embodiment, the computing device 100 may use the converted location information to obtain a first closed-eyes score (810).

[0254] In one embodiment, the computing device 100 can use the standard deviation of the heatmap to make a preliminary determination of whether the driver is drowsy or has their eyes closed.

[0255] In one embodiment, the computing device 100 can compare the standard deviation of the values ​​in the heatmap corresponding to points in the first image with a first threshold (for example, a threshold that serves as a reference for the magnitude of the standard deviation). Based on such a comparison between the first threshold and the standard deviation, the computing device 100 can make a preliminary determination of whether the driver is drowsy or closed-eyed. For example, if the standard deviation exceeds the first threshold, the computing device 100 may determine that the driver is not drowsy in the first image. Furthermore, if the standard deviation exceeds the first threshold, the computing device 100 may determine that the driver is not drowsy.

[0256] In one embodiment, if the standard deviation is less than or equal to the first threshold, the computing device 100 may determine whether the driver is drowsy or has their eyes closed in the first image by comparing an eye-closed score determined using the converted or unconverted location information with a second threshold (for example, a threshold used to determine whether or not the eyes are closed). The eye-closed score may be obtained by determining the size of the eye area using points corresponding to the eye area.

[0257] In one embodiment, the second threshold may be determined based on at least one previous result corresponding to at least one previously acquired image of the first image. The second threshold may be variable based on at least one previous result corresponding to at least one previously acquired image of the first image. The second threshold may be variable based on the threshold of the image immediately preceding the first image, based on at least one previous result corresponding to at least one previously acquired image of the first image. The previous drowsiness result can indicate the judgment or prediction result of the driver's eye-closing or drowsiness state in relation to the previously acquired image of the first image. For example, if the ratio of result values ​​indicating the presence of eye-closing or drowsiness in the previous result corresponding to the previously acquired image of the first image (e.g., sequentially acquired previous image) is greater than or equal to a first ratio, the second threshold is set to a first value, and if the ratio of result values ​​indicating the presence of eye-closing or drowsiness in the previous result corresponding to the previously acquired image of the first image is less than the first ratio, the second threshold may be set to a second value higher than the first value. In such an example, assume that images 2, 3, and 4 were acquired prior to image 1, and that image 2 yielded a result indicating the presence of closed eyes or drowsiness, image 3 yielded a result indicating the absence of closed eyes or drowsiness, and image 4 yielded a result indicating the presence of closed eyes or drowsiness. Under these assumptions, the threshold for determining the result of closed eyes or drowsiness in image 1 may be set to a value lower than the threshold for determining the result of image 2 (i.e., the image acquired immediately before image 1), depending on the percentage of results indicating the presence of closed eyes or drowsiness among the previous results of images 2, 3, and 4 (e.g., a ratio of 66.6%).

[0258] In one embodiment, the computing device 100 may compare a first eye-closed score with a second threshold to obtain a first anomaly prediction result indicating whether or not there is eye closure in the image (820).

[0259] In one embodiment, the first predicted result for the first abnormality may correspond to the first predicted result for the first abnormality in Figure 4. The first predicted result for the first abnormality may be determined based on a comparison between a blindfold score obtained from a model or a blindfold score obtained by a rule-based algorithm and a threshold. If drowsiness is present in the image, or if it is determined that the driver has their eyes closed in the image, the first predicted result for the first abnormality may have a value of 1, and if drowsiness is not present in the image, or if it is determined that the driver has not their eyes closed in the image, the first predicted result for the first abnormality may have a value of 0.

[0260] In one embodiment, the primary prediction result for an anomaly in Figure 8 may correspond to the primary prediction result with eyes closed.

[0261] In one embodiment, the first predicted result for an anomaly may be determined using the standard deviation of the values ​​in the heatmap corresponding to the points in the image. The standard deviation may be compared to a specific threshold to determine the first predicted result indicating whether or not the eyes are closed. If the standard deviation exceeds the specific threshold, the computing device 100 may determine the first predicted result in the image to be an eye-opening state, regardless of the eye-opening score. That is, if the standard deviation is large enough to exceed the specific threshold, it can be determined to be another exceptional situation, such as wearing sunglasses. As a result, the computing device 100 may determine that the eyes are open in the image, or, depending on the implementation, may not determine whether or not the eyes are closed.

[0262] In one embodiment, the computing device 100 may obtain a secondary prediction result for the first anomaly by performing a first voting using the primary prediction result for the first anomaly (830).

[0263] In one embodiment, the computing device 100 may determine a drowsiness alarm corresponding to an input first image by performing a first voting using the primary prediction result of a first anomaly. The secondary prediction result of the first anomaly may indicate whether or not a driver drowsy state is present in the image. The secondary prediction result of the first anomaly may indicate whether or not a drowsiness alarm corresponding to the image will be generated. The secondary prediction result of the first anomaly can be used as a factor for determining the type of drowsiness alarm corresponding to the image. The secondary prediction result of the first anomaly can be used as a factor for determining the intensity of the drowsiness alarm corresponding to the image. The secondary prediction result of the first anomaly can be used as a factor for determining whether or not a drowsiness alarm will be generated corresponding to the image.

[0264] In one embodiment, the secondary prediction result for an anomaly in Figure 8 may correspond to the secondary prediction result with eyes closed.

[0265] In one embodiment, the first voting may determine which of the primary prediction result for the first anomaly of the first image subject to anomaly judgment and the primary prediction result for the anomaly of a previous image is the dominant value, and use the determined dominant value to generate a secondary prediction result for the first anomaly corresponding to the first image.

[0266] In one embodiment, the first voting may generate a predicted result of an image anomaly that represents an image group consisting of a first image and a predetermined first number of sequential images previously acquired for the first image. In one embodiment, the first voting may be used to generate a secondary predicted result of a first anomaly that indicates whether or not drowsiness is present in the first image. The predicted result of the image anomaly here may include a result value that represents the group, among result values ​​that indicate the presence of drowsiness or closed eyes and result values ​​that indicate the absence of drowsiness or closed eyes.

[0267] In one embodiment, the first voting may include result values ​​representing a closed-eye state and result values ​​representing an open-eye state, which represent a particular set of images. For example, if images with results indicating closed eyes have a high proportion among a set of images, the predicted result representing that set (i.e., the predicted result for each image included in that set) may be determined to be a closed-eye state.

[0268] In an additional embodiment, the value of the closed-eyes score may be further reflected in the result value that represents the set.

[0269] In one embodiment, it is assumed that there are primary prediction results for anomalies corresponding to five images (including the first image), and that these have values ​​of 0, 0, 1, 1, and 0, respectively. Under such assumptions, 0 can represent a closed-eye state, and 1 can represent an open-eye state. Under such assumptions, the prediction result for the collective anomaly can have a value of 0. Consequently, the secondary prediction result for the first anomaly corresponding to the first image can have a value of 0.

[0270] In one embodiment, the secondary prediction result for closed eyes or the secondary prediction result for an anomaly can be represented as a value of a counter. The value of the counter may be determined using the result of a first voting. Depending on the result of the first voting, it may be determined whether the value of the counter is increased, decreased, or maintained. It may also be determined whether the value of the counter is increased, decreased, or maintained based on the first voting and a critical range of the counter value. The critical range of the counter value can define the maximum value at which the counter value does not increase, but is maintained or decreased, and the minimum value at which it is maintained or increases without decreasing. The computing device 100 may generate one or more current counter values ​​corresponding to the current image by increasing or decreasing one or more previous counter values ​​corresponding to previously acquired images of the current image based on the result of the first voting, and may also acquire a closed eyes secondary prediction result including the one or more current counter values.

[0271] In one embodiment, the unit of increment or decrement of the counter can be set in various forms. For example, a predetermined fixed value may be used as the unit of increment or decrement. For example, the unit of increment or decrement of the counter may be determined based on the difference between the acquisition time of the previous image and the acquisition time of the current image. In such an example, if the difference between the acquisition time of the previous image and the acquisition time of the current image is 15 ms, and the primary prediction result includes the presence of a closed eye, the counter value of the target image can be increased by a value such as 15 or 0.15 compared to the counter value of the previous image.

[0272] In one embodiment, the computing device 100 may determine a drowsiness alarm corresponding to an image by performing a second voting using the secondary prediction result of the first anomaly (840).

[0273] In one embodiment, the computing device 100 may determine whether or not to generate a drowsiness alarm corresponding to the first image, or whether or not the driver is drowsy in the first image, by performing a second voting using the secondary prediction result of the first anomaly. In one embodiment, the computing device 100 may determine the intensity or type of drowsiness alarm corresponding to the first image by performing a second voting using the secondary prediction result of the first anomaly.

[0274] In one embodiment, the second voting may determine whether there is continuity in the secondary prediction results of an anomaly within an image set consisting of the first image subject to anomaly determination and a predetermined second number of sequential images previously acquired for the first image. Based on whether there is continuity in the secondary prediction results of anomaly, the second voting may determine whether to generate an alarm corresponding to the first image, or whether the first image shows user drowsiness.

[0275] In one embodiment, the second voting may include comparing a counter value with a threshold. The second voting can be used to determine whether to turn on or off one or more drowsiness alarms corresponding to an image by comparing one or more current counter values ​​with one or more predetermined counter thresholds. For example, the drowsiness alarm can be set to ON when the current counter value reaches a specific threshold, and to OFF when the current counter value falls below a specific threshold. Based on the comparison of each of a plurality of counter values ​​with each of a plurality of thresholds, the type or intensity of the alarm may be determined by a combination of a specific counter value and a specific threshold if a specific counter value exceeds a specific threshold. In one embodiment, the degree to which the counter value exceeds the alarm threshold and the intensity of the alarm may be interrelated. The intensity of the alarm may be determined to differ depending on the degree to which the counter value exceeds the alarm threshold. The excess value of the counter value exceeding the alarm threshold and the intensity of the alarm may have a positive correlation. In such embodiments, the greater the degree to which the counter value exceeds the alarm threshold, the greater the intensity of the alarm may be. The intensity of the alarm here may represent the magnitude of the alarm image, alarm sound, alarm light, and / or alarm vibration.

[0276] In one embodiment, the second voting can use multiple alarm thresholds. For example, the multiple alarm thresholds may have different alarm intensities and / or alarm types. In another example, some of the multiple alarm thresholds may have the same alarm intensity and / or alarm type. The second voting may determine the type of alarm, the intensity of the alarm, and / or whether an alarm has occurred by comparing each of the multiple alarm thresholds with a counter corresponding to a particular image. For example, a first alarm, which occurs when the counter exceeds one of the multiple alarm thresholds, may have a lower alarm intensity than a second alarm, which occurs when the counter exceeds two or more alarm thresholds.

[0277] In one embodiment, multiple counters may be assigned to the target image. The second voting may determine whether an alarm is generated for each of the multiple counters by comparing each of the multiple counters with an alarm threshold. For each of the counters, the type or intensity of the alarm may be the same or different from one another. For example, a first alarm, which occurs when one counter exceeds the alarm threshold, may have a lower intensity than a second alarm, which occurs when multiple counters exceed the alarm threshold.

[0278] In one embodiment, the second voting may determine whether an alarm has occurred, the type of alarm, and / or the intensity of an alarm by comparing each of the multiple counters with each of the multiple alarms.

[0279] As described above, the technology according to one embodiment of the present disclosure can utilize one or more votings to generate alarms corresponding to drowsiness and / or to determine the drowsy state. By utilizing one or more votings, the sensitivity, accuracy, and reliability of alarms corresponding to drowsiness conditions can be increased. By utilizing one or more votings, the sensitivity, accuracy, and reliability of drowsiness judgment can be increased.

[0280] Figure 9 illustrates a method of spatial transformation according to one embodiment of the present disclosure.

[0281] As an example, the spatial transformation method illustrated in Figure 9 may be used in the process of determining the driver's drowsiness state in Figure 5.

[0282] As shown in FIG. 9, in an image 910 for determining a drowsy state, a first set of points 910a and 910b may be determined. The image 910 may correspond to a first space 910. The image 910 may correspond to a cropped bounding box 910 for the eye region in an image acquired from a camera. The image 910 may represent a space 910 corresponding to the eye region or the face region in an image acquired from a camera. The image 910 may include a plurality of points (e.g., target points) defining the eye region. The image 910 may have a first coordinate system for describing the first space 910.

[0283] In one embodiment, the computing device 100 may determine a first set of points 910a and 910b on the first coordinate system, and / or first position information 910a and 910b of the first set of points. The computing device 100 may determine a first set of points 910a and 910b for determining a transformation matrix from among the detected target points. As an example, the first set of points 910a and 910b may correspond to the points at both ends of both detected eyes. The computing device 100 may use the first set of points 910a and 910b to determine a transformation matrix for positioning the first set of points 910a and 910b in a transformed space 930.

[0284] In one embodiment, when a transformation matrix is determined, transformed position information 930a and 930b of the first set of points 910a and 910b in the transformed space 930 may be determined.

[0285] In one embodiment, the computing device 100 may determine the transformed position information 930c of a second set of points excluding a first set of points among the target points in the transformation space 930. The transformed position information 930c of the second set of points may be determined with reference to the transformed position information 930a and 930b of the first set of points. The transformed position information 930c of the second set of points may be obtained by applying a transformation matrix to the second set of points on the first space 910. Combining the transformed position information 930a and 930b of the first set of points and the transformed position information 930c of the second set of points may generate position information corresponding to the target points.

[0286] In one embodiment, the computing device 100 can more accurately calculate the eyelid score for the eye region by transforming the target points corresponding to the eye region in a normalized space (i.e., the transformation space).

[0287] FIG. 10 illustratively shows a method for determining whether to generate a driver drowsiness alarm according to an embodiment of the present disclosure.

[0288] In one embodiment, the computing device 100 may detect eye landmarks from the acquired image (1010).

[0289] For example, the eye landmarks may correspond to one or more feature points corresponding to the eye region in the image. For example, the eye landmarks may be generated based on information obtained from a model that outputs feature points corresponding to the eye region in the input image using heatmap regression. In other embodiments, whether to perform detection on the eye landmarks may be determined based on the magnitude of the standard deviation of the heatmap that is the output of the model.

[0290] In one embodiment, the computing device 100 can calculate the standard deviation of values ​​in a heatmap included in the results of eye landmark detection 1010. For example, the computing device 100 can utilize the heatmap provided during the eye landmark detection 1010 process to calculate the standard deviation of values ​​included in the heatmap. For example, the computing device 100 may determine the standard deviation of values ​​included in the heatmap by utilizing a heatmap included in or obtained from the output of the model during the eye landmark detection 1010 process. A large standard deviation in such a heatmap may indicate that there is a high probability that objects different from the eye region used during the model training process exist, or that they are not eye regions. A small standard deviation in such a heatmap may indicate that there is a low probability that objects different from the eye region used during the model training process exist, or that they are eye regions.

[0291] It may be determined whether the calculated standard deviation passes the standard deviation critical criterion (1020).

[0292] For example, if the standard deviation is greater than the value that falls within the critical threshold, it can be defined as having failed to meet the critical standard deviation threshold. For example, if the standard deviation is less than or equal to the value that falls within the critical threshold, it can be defined as having met the critical standard deviation threshold.

[0293] For example, the larger the standard deviation value included in the analysis of the standard deviation of such a heatmap, the higher the probability of occlusion occurring within the image.

[0294] In one embodiment, if it is determined that the calculated standard deviation did not pass the standard deviation critical criterion, the computing device 100 may decide not to generate an eye-closed alarm corresponding to the image in question (1030).

[0295] In one embodiment, if it is determined that the calculated standard deviation has passed the standard deviation critical criterion, the computing device 100 may proceed with an additional process to determine whether or not to generate an eye-closed alarm corresponding to the image in question.

[0296] In one embodiment, if it is determined that the calculated standard deviation has passed the standard deviation critical criterion, the computing device 100 may determine that the image in question is an image of the eyes open.

[0297] In one embodiment, the computing device 100 may calculate a closed-eyes score and determine whether the closed-eyes score passes the closed-eyes critical criterion if it is determined that the standard deviation has passed the standard deviation critical criterion (1040).

[0298] The closed-eye score can be calculated based on the results of eye landmark detection 1010. For example, the closed-eye score can be calculated from the results of eye landmark detection 1010, based on the size of the eye and / or the distance between the top and bottom of the eye region. For example, the closed-eye score may be output by a model for eye landmark detection 1010. For example, the closed-eye score can be calculated based on the results of transforming the results of eye landmark detection 1010 into a normalized space (e.g., a transformed space). The closed-eye criticality criterion can be defined, for example, as a criterion value that quantitatively represents the degree of closed eyes. If the closed-eye criticality criterion is not passed (i.e., the quantitative intensity of closed eyes is lower than the criticality criterion), it can be determined that the eyes are open. If the closed-eye criticality criterion is passed (i.e., the quantitative intensity of closed eyes is equal to or greater than the criticality criterion), it can be determined that the eyes are closed.

[0299] In one embodiment, a smaller value for the measured distance between the eyes in a normalized space may indicate a higher probability of eye closure. Depending on the implementation, the computing device 100 may set a higher eye closure score for smaller values ​​for the measured distance between the eyes.

[0300] In one embodiment, when it is determined that the eyes-closed critical criterion has not been passed, the computing device 100 may determine not to generate an eyes-closed alarm (1050).

[0301] In one embodiment, when it is determined that the eyes-closed critical criterion has not been passed, the computing device 100 may determine that there is no eyes-closed in the corresponding image.

[0302] In one embodiment, when it is determined that the eyes-closed critical criterion has not been passed, the computing device 100 may determine that there is no abnormality corresponding to drowsiness.

[0303] In one embodiment, when it is determined that the eyes-closed critical criterion has been passed, the computing device 100 may determine to generate an eyes-closed alarm (1060).

[0304] In one embodiment, when it is determined that the eyes-closed critical criterion has been passed, the computing device 100 can determine that there is an abnormality corresponding to eyes-closed.

[0305] In one embodiment, when it is determined that the eyes-closed critical criterion has been passed, the computing device 100 can determine that there is an abnormality corresponding to drowsiness.

[0306] As described above, the technology according to an embodiment of the present disclosure can provide an accurate determination result for eyes-closed or drowsiness by using a heat map that is an output of a model for detecting eye landmarks. The technology according to an embodiment of the present disclosure may further utilize the heat map used to detect eye landmarks to determine whether there are abnormal objects such as an hourglass. The technology according to an embodiment of the present disclosure can increase the accuracy of the alarm by determining not to generate an alarm when the standard deviation criterion of the heat map is not satisfied.

[0307] Figure 11 illustrates a methodology for generating an abnormal alarm according to one embodiment of the present disclosure.

[0308] The example shown in Figure 11 will be replaced with the previous explanation to avoid repetition of the explanation.

[0309] As shown in Figure 11, a voting technology according to one embodiment of the present disclosure may utilize multiple queues to determine an anomaly alarm from the detection results of a detection model, or to determine whether or not to generate an anomaly alarm. The queues in the present disclosure can be used interchangeably with a voting queue.

[0310] In one embodiment, the boating queue may include a first queue 1110, a second queue 1130, a third queue 1150, and a fourth queue 1170. Such first queues 1110, 2130, 3150, and 4170 can be used to determine abnormal alarms according to one embodiment of the present disclosure. Although four queues are shown in Figure 11, it will be apparent to those skilled in the art that various numbers of queues are available depending on various implementations, such as adding new boatings or deleting existing boatings.

[0311] In one embodiment, the X-axis direction (i.e., the lateral direction) of the first queue 1110, second queue 1130, third queue 1150, and fourth queue 1170 is the direction for representing images acquired in time. The first queue 1110, second queue 1130, third queue 1150, and fourth queue 1170 are configured such that moving to the right approaches the current point in time, and moving to the left approaches a past point in time. The first queue 1110, second queue 1130, third queue 1150, and fourth queue 1170 contain multiple units. In Figure 11, the first queue 1110, second queue 1130, third queue 1150, and fourth queue 1170 are shown to contain five units corresponding to each other. It will be apparent to those skilled in the art that a single queue may contain a variety of units depending on the mode of implementation.

[0312] In one embodiment, each of the multiple units in the queue may correspond to each of the acquired images. In the first queue 1110, the second queue 1130, the third queue 1150, and the fourth queue 1170, the first units 1110a, 1130a, 1150a, and 1170a may include results corresponding to the first image from the first viewpoint; the second units 1110b, 1130b, 1150b, and 1170b may include results corresponding to the second image from the second viewpoint; the third units 1110c, 1130c, 1150c, and 1170c may include results corresponding to the third image from the third viewpoint; the fourth units 1110d, 1130d, 1150d, and 1170d may include results corresponding to the fourth image from the fourth viewpoint; and the fifth units 1110e, 1130e, 1150e, and 1170e may include results corresponding to the fifth image from the fifth viewpoint. Here, the first viewpoint is the viewpoint prior to the second viewpoint, the second viewpoint is the viewpoint prior to the third viewpoint, the third viewpoint is the viewpoint prior to the fourth viewpoint, and the fifth viewpoint is the viewpoint prior to the fourth viewpoint.

[0313] In one embodiment, the units corresponding to the first queue 1110, the second queue 1130, the third queue 1150, and the fourth queue 1170 may each contain the judgment result for the same image acquired at the same time. For example, the first image acquired at the first time may be processed in the following order: the first unit 1110a of the first queue 1110, the first unit 1130a of the second queue 1130, the first unit 1150a of the third queue 1150, and the first unit 1170a of the fourth queue 1170. As a result of such processing, the fourth queue 1170, which is the last queue, may contain the result of turning the abnormal alarm ON or OFF for each image.

[0314] In one embodiment, the first queue 1110 may include a first unit 1110a, a second unit 1110b, a third unit 1110c, a fourth unit 1110d, and a fifth unit 1110e. For example, the fifth unit 1110e may include a decision result corresponding to the currently (or most recently) retrieved image. In one embodiment, the first queue 1110 is a queue representing the output of a model or a queue showing the result of post-processing the output of a model. For example, the first queue 1110 may include model output information for each of a plurality of images. For example, the first queue 1110 is a queue showing the detection result of a detection model. The first unit 1110a includes first model output information corresponding to the first image of the first viewpoint, the second unit 1110b includes second model output information corresponding to the second image of the second viewpoint, the third unit 1110c includes third model output information corresponding to the third image of the third viewpoint, the fourth unit 1110d includes fourth model output information corresponding to the fourth image of the fourth viewpoint, and the fifth unit 1110e may include fifth model output information corresponding to the fifth image of the fifth viewpoint. Here, the first viewpoint is the earliest viewpoint, and viewpoints can be represented as you move from the first viewpoint to the fifth viewpoint. In addition, the sixth unit 1110f is assigned sixth model output information corresponding to the sixth image acquired at the sixth viewpoint, which is a future viewpoint. In one embodiment, the first queue 1110 may be set to include a predetermined number of units in a manner in which the oldest unit is removed when a new unit is drawn in. For example, if the sixth model output information is assigned to the sixth unit 1110f, the first unit 1110a is removed from the queue.

[0315] In one embodiment, the second queue 1130 may include a first unit 1130a, a second unit 1130b, a third unit 1130c, a fourth unit 1130d, and a fifth unit 1130e. The values ​​contained in each of the units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130 may also be determined based on a comparison between the information contained in the unit corresponding to the first queue 1110 and the threshold assigned to the corresponding unit of the first queue 1110. The values ​​contained in each of the units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130 may include a primary prediction result of an anomaly.

[0316] In one embodiment, the second queue 1130 may include a primary prediction result for an anomaly determined based on a comparison between model output information and a threshold. The first unit 1130a includes a primary prediction result for a first anomaly corresponding to a first image of a first viewpoint, the second unit 1130b includes a primary prediction result for a second anomaly corresponding to a second image of a second viewpoint, the third unit 1130c includes a primary prediction result for a third anomaly corresponding to a third image of a third viewpoint, the fourth unit 1130d includes a primary prediction result for a fourth anomaly corresponding to a fourth image of a fourth viewpoint, and the fifth unit 1130e includes a primary prediction result for a fifth anomaly corresponding to a fifth image of a fifth viewpoint.

[0317] The expression “the threshold is variable” as used herein may be used to encompass embodiments in which the threshold is set to one of several (e.g., two) threshold options, and embodiments in which the threshold is increased, decreased, or kept equal to a previous threshold.

[0318] In one embodiment, the threshold assigned to the current unit of the first queue 1110 may be variable based on values ​​included in previous units of other queues. For example, the threshold assigned to the current unit of the first queue 1110 may be determined based on the first predicted result of an anomaly of a previous unit of the first queue 1110. For example, in Figure 11, if the model output information exceeds the threshold, the first predicted result of the anomaly may have a value of 1, and if the model output information does not exceed the threshold, the first predicted result of the anomaly may have a value of 0. The first predicted result of an anomaly can be determined by comparing the first model output information of 0.6 with the threshold of 0.5, and may have a value of 1. The first predicted result of an anomaly can be determined by comparing the second model output information of 0.6 with the threshold of 0.5, and may have a value of 1. The first predicted result of an anomaly can be determined by comparing the third model output information of 0.5 with the threshold of 0.4, and may have a value of 1. The first predicted result of an anomaly can be determined by comparing the third model output information of 0.5 with the threshold of 0.4, and may have a value of 1. The first predicted result of an anomaly can be determined by comparing the fourth model output information of 0.3 with the threshold of 0.4, and may have a value of 0. The primary prediction result for the fifth anomaly can be determined by comparing the fifth model output information of 0.6 with a threshold of 0.4, and may have a value of 1. As shown in Figure 11, if the value of a specific unit in the first queue 1110 is greater than or equal to the threshold assigned to that specific unit in the first queue 1110, the value included in that specific unit in the second queue 1130 is determined to be 1. If the value of a specific unit in the first queue 1110 is less than the threshold assigned to that specific unit in the first queue 1110, the value included in that specific unit in the second queue 1130 is determined to be 0. For example, the threshold assigned to the current unit (e.g., the Nth unit) in the first queue 1110 may be variable depending on the value included in the unit immediately preceding the current unit in the second queue 1130 (e.g., the N-1th unit). In such an example, if the value of the second unit 1130b of the second queue 1130 is set to 1, the threshold assigned to the third unit 1110c of the first queue 1110 may be determined to be 0.4 by subtracting 0.1 from the threshold of the previous unit, the second unit 1110b, which is 0.5.In another embodiment, if the value of the second unit 1130b of the second queue 1130 is set to 0, the threshold assigned to the third unit 1110c of the first queue 1110 may be kept equal to the threshold of the previous unit, the second unit 1110b, which is 0.5, or 0.1 may be added to determine a threshold of 0.6. In this specification, N may mean a natural number.

[0319] In one embodiment, the threshold corresponding to each unit of the first queue 1110 may be set to one of a first value and a second value. In such an embodiment, the threshold corresponding to the current unit of the first queue 1110 may be set to one of a first value and a second value using the result value corresponding to the previous unit. In such an embodiment, if the result value corresponding to the previous unit indicates the presence of an anomaly, the threshold corresponding to the current unit of the first queue 1110 may be set to the lower of the first value and the second value. In such an embodiment, if the result value corresponding to the previous unit indicates the absence of an anomaly, the threshold corresponding to the current unit of the first queue 1110 may be set to the higher of the first value and the second value.

[0320] In one embodiment, the current threshold corresponding to the current unit may be maintained or modified relative to the previous threshold corresponding to the previous unit. In an example where the threshold changes, for example, if the predicted anomaly result corresponding to the previous unit is determined to indicate the presence of an anomaly, the current threshold may be set to have a downward trend. Here, a downward trend may include keeping the current threshold the same as the previous threshold or setting the current threshold lower than the previous threshold. Setting the current threshold lower may mean that the current image is more likely to be judged as an anomaly. In an example where the threshold changes, for example, if the predicted anomaly result corresponding to the previous unit is determined to indicate the absence of an anomaly, the current threshold may be set to have an upward trend. Here, an upward trend may include keeping the current threshold the same as the previous threshold or setting the current threshold higher than the previous threshold. Setting the current threshold higher may mean that the current image is less likely to be judged as an anomaly. As described above, the technology according to one embodiment of the present disclosure can improve the accuracy of anomaly detection by determining the predicted anomaly result for the current image to have a similar (corresponding) trend to the predicted anomaly result for the previous image.

[0321] In one embodiment, the threshold assigned to the current unit of the first queue 1110 may be determined based on values ​​included in a plurality of previous units of the second queue 1130. In one embodiment, the threshold assigned to a particular unit of the first queue 1110 may be variable depending on the values ​​of previous units of other queues other than the first queue 1110. For example, the threshold assigned to the current unit of the first queue 1110 may be determined based on a ratio value obtained from the values ​​of previous units of the second queue 1130, or on a comparison of the average value with a particular threshold. Here, the particular threshold is a threshold for determining the threshold assigned to the current unit, and may mean a threshold in the form of a ratio, such as 30%, 40%, 50%, and 80%. For example, the threshold of the second unit 1110b of the first queue 1110 may be determined based on the values ​​of the first unit 1130a and the previous units of the first unit 1130a of the second queue 1130. If it is determined that the ratio of values ​​of 1 (i.e., the ratio judged to be abnormal) among the values ​​of the first unit 1130a and the previous units of the first unit 1130a of the second queue 1130 does not exceed a certain threshold, then the threshold for the second unit 1110b of the first queue 1110 can be kept at 0.5. The threshold for the third unit 1110c of the first queue 1110 may be determined based on the values ​​of the first unit 1130a and the second unit 1130b of the second queue 1130. Alternatively, the threshold for the third unit 1110c of the first queue 1110 may be determined based on the values ​​of the first unit 1130a, the second unit 1130b, and the previous units of the first unit 1130a of the second queue 1130. For example, if it is determined that the ratio of values ​​of 1 (i.e., the ratio judged to be abnormal) among the values ​​of the first unit 1130a of the second queue 1130 and the previous units of the first unit 1130a exceeds a certain threshold, the threshold for the third unit 1110c of the first queue 1110 may be set to 0.4. As another example, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be determined based on a comparison between the values ​​of the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 that represent abnormalities and a predetermined threshold.For example, if it is determined that the proportion of values ​​indicating an anomaly (i.e., a value of 1) among the values ​​of the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 is 40% or more, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be set to 0.4. For example, if it is determined that the proportion of values ​​indicating an anomaly (i.e., a value of 1) among the values ​​of the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 is less than 40%, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be set to 0.5.

[0322] In one embodiment, the threshold assigned to the current unit (e.g., the Nth unit) of the first queue 1110 may be variable depending on the values ​​contained in the previous units (e.g., the N-1st, N-2nd, N-3rd, and N-4th units) of the current unit of the second queue 1130. In such an example, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be determined based on the dominant or representative values ​​of the values ​​in the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130. The dominant or representative value may be determined to be the value with the highest weight among the values ​​contained in a predetermined number of units. For example, the dominant or representative value may be determined from the values ​​of 1 or 0 contained in five sequential units. In such an example, since four of the units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 have a value of 1 and one unit has a value of 0, the mainstream or representative value of the five units of the second queue 1130 may be determined to be 1. If the mainstream or representative value of the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 is 1, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be set to 0.4. In such an example, if the mainstream or representative value of the previous units 1130a, 1130b, 1130c, and 1130d of the second queue 1130 is 0, the threshold assigned to the fifth unit 1110e of the first queue 1110 may be set to 0.5. As shown in Figure 11, the threshold assigned to the sixth unit 1110f of the first queue 1110 may be determined based on the mainstream, representative, or ratio values ​​of the previous units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130.

[0323] In this specification, for the sake of clarity, the relevant details are omitted, but it will be apparent to those skilled in the art that, depending on the characteristics of the queue, when a new unit is added to the queue, the oldest unit is removed, the value of the first unit 1110a of the first queue 1110 relative to the previous unit can also be considered.

[0324] In one embodiment, the threshold assigned to the current unit of the first queue 1110 may be determined based on a value included in at least one previous unit of the third queue 1150. The threshold assigned to the fifth unit 1110e of the first queue 1110 may be determined based on the value of a previous unit of the third queue 1150 (e.g., the fourth unit 1150d). For example, if the value of the fourth unit 1150d of the third queue 1150 is 1, the threshold for the fifth unit 1110e of the first queue 1110 may be determined to be 0.4. Alternatively, the threshold may be kept the same as the previous threshold, or the threshold may be decreased compared to the previous threshold. For example, if the value of the first unit 1150a of the third queue 1150 is 0, the threshold for the second unit 1110b of the first queue 1110 may be determined to be 0.5. Alternatively, the threshold may be kept the same as the previous threshold, or the threshold may be increased compared to the previous threshold.

[0325] In one embodiment, the values ​​of units 1150a, 1150b, 1150c, 1150d, and 1150e of the third queue 1150 may be determined based on the results of a first voting that utilizes the values ​​of units 1130a, 1130b, 1130c, 1130d, and 1130e contained in the second queue 1130. The values ​​contained in the units of the third queue 1150 may be determined based on the mainstream, representative, or ratio of the primary predicted results of anomalies. For example, the fifth unit 1150e of the third queue 1150 may be determined based on the representative or mainstream values ​​of units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130. In such an example, the fifth unit 1150e of the third queue 1150 may be set to 1, which is the mainstream or representative value of units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130. The number of units used to determine the mainstream or representative value may vary, not just five, but also three or four, for example. In another example, the fifth unit 1150e of the third queue 1150 may also be determined based on a comparison of the ratio of units with a value of 1 from the values ​​of units 1130a, 1130b, 1130c, 1130d, and 1130e of the second queue 1130 with a threshold ratio.

[0326] In one embodiment, the fourth queue 1170 may correspond to a queue for generating alarms. In one embodiment, the values ​​of units 1170a, 1170b, 1170c, 1170d, and 1170e of the fourth queue 1170 may be determined based on a second voting that utilizes the values ​​of the units of the third queue 1150. The fourth queue 1170 may decide whether or not to generate an alarm by considering the continuity of the unit values ​​of the third queue 1150 via the second voting. For example, assume a criterion of continuity of unit values ​​of 3. Under such assumption, since neither the first unit 1150a of the third queue 1150 nor the two preceding units have a value of 1, the first unit 1170a of the fourth queue 1170 may be set to 0 (i.e., alarm off). Since neither the second unit 1150b of the third queue 1150 nor the two preceding units (1150a and the previous unit) have a value of 1, the second unit 1170b of the fourth queue 1170 may be set to 0 (i.e., alarm off). Since neither the third unit 1150c of the third queue 1150 nor the two preceding units 1150a and 1150b have a value of 1, the third unit 1170c of the fourth queue 1170 may be set to 0 (i.e., alarm off). Since neither the fourth unit 1150d of the third queue 1150 nor the two preceding units 1150b and 1150c have a value of 1, the fourth unit 1170c of the fourth queue 1170 may be set to 0 (i.e., alarm off). Since the fifth unit 1150e of the third queue 1150 and the two preceding units 1150c and 1150d all have a value of 1, the fifth unit 1170e of the fourth queue 1170 may be set to 1 (i.e., alarm on).

[0327] As described above, the technology according to one embodiment of the present disclosure can utilize multiple queues to determine the presence or absence of an anomaly or to generate an alarm corresponding to such an anomaly. Values ​​contained in other queues (e.g., previous queues) among these multiple queues may be used to determine the values ​​contained in a particular queue or the threshold to be assigned to a particular queue.

[0328] Figure 12 illustrates a methodology for determining abnormal alarms according to one embodiment of the present disclosure.

[0329] In the examples shown in Figure 12, the previously mentioned examples will be replaced with the previously mentioned explanations to avoid repetition of explanations.

[0330] The first queue 1210 and the units 1210a, 1210b, 1210c, 1210d, 1210e, and 1210f included in the first queue in Figure 12 may correspond to the first queue 1110 and the units 1110a, 1110b, 1110c, 1110d, 1110e, and 1110f included in the first queue in Figure 11, respectively.

[0331] The second queue 1230 and the units 1230a, 1230b, 1230c, 1230d, and 1230e included in the second queue in Figure 12 may correspond to the second queue 1130 and the units 1130a, 1130b, 1130c, 1130d, and 1130e included in the second queue in Figure 11, respectively.

[0332] The third queue 1250 in Figure 12 may include a first unit 1250a, a second unit 1250b, a third unit 1250c, a fourth unit 1250d, and a fifth unit 1250e. In one embodiment, the values ​​of units 1250a, 1250b, 1250c, 1250d, and 1250e of the third queue 1250 may be determined based on the results of a first voting that utilizes the values ​​of units 1230a, 1230b, 1230c, 1230d, and 1230e contained in the second queue 1230. The values ​​included in the units of the third queue 1250 may be determined based on the mainstream value, representative value, or ratio of the primary predicted result of anomalies.

[0333] Figure 12 illustrates an embodiment in which, by performing a first voting, the predicted results of secondary anomalies are generated in a manner that utilizes a counter or counter value. The example in Figure 12 shows an embodiment in which, as a result of the first voting, the counter value increases as the mainstream value or representative value becomes 1, and maintains its value as a result of the first voting, the mainstream value or representative value becomes 0. However, embodiments in which the counter value decreases as the mainstream value or representative value becomes 0, depending on the mode of implementation, may also be included within the scope of the rights of this disclosure.

[0334] In one embodiment, by performing a first voting using multiple units, including the first unit 1230a of the second queue 1230, the second unit 1250b of the third queue 1250 can have a value of 0. As a result of the first voting using multiple units, including the first unit 1230a and the second unit 1230b of the second queue 1230, the third unit 1250c of the third queue 1250 can have a value of 1 as the counter value of 1 is added. As a result of the first voting using multiple units, including the first unit 1230a, the second unit 1230b, and the third unit 1230c of the second queue 1230, the fourth unit 1250d of the third queue 1250 can have a value of 2 as the counter value of 1 is added. As a result of the first voting using multiple units, including the first unit 1230a, second unit 1230b, third unit 1230c, and fourth unit 1230d of the second queue 1230, the fifth unit 1250e of the third queue 1250 can have a value of 3 as the counter value of 1 is added. In this way, when generating the secondary prediction result of anomalies, if the result of the first voting is 0, the counter value is not added, and if the result of the first voting is 1, the counter value can be added.

[0335] In one embodiment, the fourth queue 1270 may correspond to a queue for generating alarms. In one embodiment, the values ​​of units 1270a, 1270b, 1270c, 1270d, and 1270e of the fourth queue 1270 may be determined based on a second voting that utilizes the values ​​of the units of the third queue 1250.

[0336] In one embodiment, the second voting may include comparing the value of each unit in the third queue 1250 (e.g., a counter value) with one or more thresholds. For example, if the value of a unit in the third queue 1250 is greater than or equal to a threshold, the second voting may decide to generate an alarm (ON) for the corresponding unit in the fourth queue 1270. For example, if the value of a unit in the third queue 1250 is less than a threshold, the second voting may decide not to generate an alarm (OFF) for the corresponding unit in the fourth queue 1270. In the example in Figure 12, since the threshold is set to 3, the computing device 100 may decide to generate an abnormal alarm for the fifth unit 1270e of the fourth queue 1270, which corresponds to the fifth unit 1250e of the third queue 1250 having a value of 3.

[0337] In one embodiment, the counter value may have a predetermined range. When the counter value reaches a boundary value within the predetermined range, the computing device 100 can maintain or change the counter value in a different manner than when it has not reached the boundary value. For example, suppose the range of the counter value is 0 to 5, and the threshold is 3. Under such assumptions, the threshold of 3 and each unit of the third queue 1250 can be compared. If the first unit of the third queue 1250 has a value of 3, the computing device 100 can set the corresponding first unit of the fourth queue 1270 to ON. If the second unit, which is the next unit of the third queue 1250, has a value of 4, the computing device 100 can set the corresponding unit of the fourth queue 1270 to ON. If the third unit, which is the next unit of the third queue 1250, has a value of 5, the computing device 100 can set the corresponding third unit of the fourth queue 1270 to ON. In such a situation, if it is determined through the first voting that the fourth unit, which is the next unit in the third queue 1250, will have a value of 1 added to it, the computing device 100 may set the fourth unit of the third queue 1250 to 5 (i.e., not add 1). As another example, in the same situation, if it is determined through the first voting that the value of 0 is the dominant or representative value for the fourth unit, which is the next unit in the third queue 1250, the computing device 100 may set the fourth unit of the third queue 1250 to 4 (i.e., subtract 1). In this way, when the secondary prediction result of an anomaly reaches a critical range, the counter value will not increase any further, and the counter value may decrease or be maintained. In such a situation, if the counter value changes from 3 to 2 as the counter value decreases, the computing device 100 may decide to turn off the anomaly alarm. As described above, the technology according to one embodiment of the present disclosure can set a critical range for the counter value and maintain or increase the counter value when the counter value reaches the minimum value of the critical range.The technology according to one embodiment of the present disclosure can set a critical range for a counter value and maintain or decrease the counter value when it reaches the maximum value of the critical range. The technology according to one embodiment of the present disclosure can set a critical range for a counter value and replace the counter value with the minimum or maximum value when it falls outside the minimum or maximum value of the critical range.

[0338] In one embodiment, the type or intensity of an alarm may be set to differ depending on the difference between the counter value and the threshold. If the counter value is 4 and the threshold is 3, the alarm may be set to the first type or first intensity. If the counter value is 5 and the threshold is 3, the alarm may be set to the second type or second intensity. Here, the second intensity may be greater than the first intensity. Here, the second type of alarm may be a different alarm from the first type of alarm. In such an example, the second type of alarm can convey a more intuitive and powerful message to the user than the first type of alarm.

[0339] In one embodiment, the technology according to one embodiment of the present disclosure may utilize one counter or multiple counters. Furthermore, multiple thresholds to be compared with the counters may be used. Alternatively, multiple counters and multiple thresholds may be used.

[0340] An example of using a single counter can be described as follows: In a situation where there is one counter for N alarms, the counter may be incremented, decremented, or maintained by 1 according to the predicted result (predicted value) of a quadratic anomaly. A starting value and / or critical range (min and max values) of the counter may be set, and N thresholds between the min and max values ​​for triggering an alarm may be set, where N is a natural number. The counter is compared to each of the N thresholds, and an alarm is triggered if the counter is greater than or equal to a specific threshold, and in the case of a combined condition (when the counter is greater than or equal to multiple thresholds), an alarm corresponding to the higher threshold may be triggered. For example, suppose the critical range of the counter is min=0 and max=6, and the first threshold=3 and the second threshold=5. In such a situation, when the predicted result of an anomaly in a particular image is 1, the counter has a value of 1, and no alarm is triggered. When the predicted result of an anomaly in the next image is 1, the counter has a value of 2, and no alarm is triggered. When the secondary prediction result for an anomaly in the next image is 1, the counter has a value of 3, and the first alarm corresponding to the first threshold can be generated. When the secondary prediction result for an anomaly in the next image is 1, the counter has a value of 4, and the first alarm corresponding to the first threshold can be generated. When the secondary prediction result for an anomaly in the next image is 1, the counter has a value of 5, and the second alarm corresponding to the second threshold can be generated. When the secondary prediction result for an anomaly in the next image is 1, the counter remains at a value of 5, and the second alarm corresponding to the second threshold can be generated. When the secondary prediction result for an anomaly in the next image is 0, the counter has a value of 4, and the first alarm corresponding to the first threshold can be generated. When the secondary prediction result for an anomaly in the next image is 0, the counter has a value of 3, and the first alarm corresponding to the first threshold can be generated. When the secondary prediction result for an anomaly in the next image is 0, the counter has a value of 3, and the alarm can be turned OFF.

[0341] An example of using multiple counters can be described as follows: In a situation where there are M counters for N alarms, each counter can be increased, decreased, or maintained by 1 according to the predicted result (predicted value) of a quadratic anomaly, where N and M are natural numbers. A starting value and / or critical range (min and max values) may be set for each counter, and M thresholds between the min and max values ​​for triggering an alarm may be set for each counter. Each of the N counters is compared with each of the M thresholds, and an alarm can be triggered if a particular counter is greater than or equal to a specific threshold. For example, the critical range of the first counter may be min=0 and max=4, and the first counter may be assigned a first threshold value of 3. Alternatively, the critical range of the second counter may be min=0 and max=6, and the second counter may be assigned a second threshold value of 5. In such an exemplary situation, when the predicted result of an anomaly in a particular image is 0, both the first and second counters can have a value of 0, and no alarm is triggered. When the secondary prediction result for an anomaly in the next image is 1, both the first and second counters can have a value of 1, and no alarm is generated. When the secondary prediction result for an anomaly in the next image is 1, both the first and second counters can have a value of 2, and no alarm is generated. When the secondary prediction result for an anomaly in the image after that is 1, both the first and second counters can have a value of 3, and since the first counter is determined to be above the first threshold, the first alarm or primary alarm can be generated. When the secondary prediction result for an anomaly in the next image is 1, both the first and second counters can have a value of 4, and since the first counter is determined to be above the first threshold, the first alarm or primary alarm can persist. When the secondary prediction result for an anomaly in the image after that is 1, the first counter is maintained at its maximum value of 4, the second counter can have a value of 5, the first counter may be determined to be above the first threshold, and the second counter may be determined to be above the second threshold. In this case, the second alarm or secondary alarm can be generated.When the secondary prediction result for an anomaly in the next image is 0, the first counter can decrease to a value of 3 and the second counter can decrease to a value of 4. Since the first counter is determined to be above the first threshold, the first alarm or primary alarm can be generated. When the secondary prediction result for an anomaly in the next image is 0, the first counter decreases to a value of 2 and the second counter decreases to a value of 3, which can turn the alarm OFF.

[0342] The first queue 1310 and the units 1310a, 1310b, 1310c, 1310d, 1310e, and 1310f included in the first queue in Figure 13 may correspond to the first queue 1110 and the units 1110a, 1110b, 1110c, 1110d, 1110e, and 1110f included in the first queue in Figure 11, respectively.

[0343] The second queue 1330 and the units 1330a, 1330b, 1330c, 1330d, and 1330e included in the second queue in Figure 13 may correspond to the second queue 1130 and the units 1130a, 1130b, 1130c, 1130d, and 1130e included in the second queue in Figure 11, respectively.

[0344] In one embodiment, the values ​​of units 1350a, 1350b, 1350c, 1350d, and 1350e of the third queue 1350 may be determined based on the results of a first voting that utilizes the values ​​of units 1330a, 1330b, 1330c, 1330d, and 1330e contained in the second queue 1330. The values ​​contained in the units of the third queue 1350 may be determined based on the mainstream value, representative value, or ratio of the primary predicted result of anomalies.

[0345] In Figure 13, the third queue 1350 and the first unit 1350a, second unit 1350b, third unit 1350c, fourth unit 1350d, and fifth unit 1350e included in the third queue 1350 may correspond to the first unit 1150a, second unit 1150b, third unit 1150c, fourth unit 1150d, and fifth unit 1150e of the third queue 1150, respectively. That is, Figure 13 shows a boating queue to which a counter is applied. Figure 13 exemplifies a method in which the unit of increase or decrease of the counter value is not 1. Figure 13 shows an embodiment in which the difference value of the acquisition time between images is used as the unit of increase or decrease of the counter value.

[0346] In one embodiment, the difference between the acquisition time of the first image corresponding to the first unit 1350a of the third queue 1350 and the acquisition time of the second image corresponding to the second unit 1350b may be determined to be 500ms. This allows the unit of increase or decrease of the counter value between the first unit 1350a and the second unit 1350b to be set to 500ms. The difference between the acquisition time of the second image corresponding to the second unit 1350b of the third queue 1350 and the acquisition time of the third image corresponding to the third unit 1350c may be determined to be 700ms. This allows the unit of increase or decrease of the counter value between the second unit 1350b and the third unit 1350c to be set to 700ms. The difference between the acquisition time of the third image corresponding to the third unit 1350c of the third queue 1350 and the acquisition time of the fourth image corresponding to the fourth unit 1350d may be determined to be 300ms. This allows the unit of increase or decrease of the counter value between the third unit 1350c and the fourth unit 1350d to be set to 300ms. The difference between the acquisition time of the fourth image corresponding to the fourth unit 1350d of the third queue 1350 and the acquisition time of the fifth image corresponding to the fifth unit 1350e may be determined to be 500ms. This allows the unit of increase or decrease in the counter value between the fourth unit 1350d and the fifth unit 1350e to be set to 500ms.

[0347] In one embodiment, the fourth queue 1370 may correspond to a queue for generating alarms. In one embodiment, the values ​​of units 1370a, 1370b, 1370c, 1370d, and 1370e of the fourth queue 1370 may be determined based on a second voting that utilizes the values ​​of the units of the third queue 1350.

[0348] As described above, the technology according to one embodiment of the present disclosure may utilize multiple queues and multiple voting to generate an anomaly alarm and / or to determine the type and intensity of such anomaly alarm. Furthermore, the technology according to one embodiment of the present disclosure may utilize one or more counters and one or more thresholds to determine a user-friendly and more accurate alarm.

[0349] Figure 13 illustrates an embodiment where the threshold is 1.1. This may result in an abnormal alarm being turned ON for the fifth unit 1370e of the fourth queue 1370.

[0350] Figure 14 illustrates an example of a method for determining an eye-closing alarm according to one embodiment of the present disclosure.

[0351] In Figure 14, the previously mentioned example will be replaced with the previously mentioned content to avoid repetition of explanations.

[0352] In one embodiment, the X-axis direction (i.e., lateral direction) of the first cue 1410, second cue 1430, third cue 1450, fourth cue 1470, and fifth cue 1490 is the direction for representing images acquired in time. The images here may include target images for determining whether or not the eyes are closed or drowsy. Each of the first cue 1410, second cue 1430, third cue 1450, fourth cue 1470, and fifth cue 1490 contains multiple units. Each unit of the first cue 1410, second cue 1430, third cue 1450, fourth cue 1470, and fifth cue 1490 can represent subsequent time points from left to right. In Figure 14, the first cue 1410, second cue 1430, third cue 1450, fourth cue 1470, and fifth cue 1490 are shown to contain five units corresponding to each other.

[0353] In the first queue 1410, the second queue 1430, the third queue 1450, the fourth queue 1470, and the fifth queue 1490, the units corresponding to each other may contain processing results for the same image. In the example in Figure 14, the first units 1410a, 1430a, 1450a, 1470a, and 1490a may include results corresponding to the first image at the first time point; the second units 1410b, 1430b, 1450b, 1470b, and 1490b may include results corresponding to the second image at the second time point; the third units 1410c, 1430c, 1450c, 1470c, and 1490c may include results corresponding to the third image at the third time point; the fourth units 1410d, 1430d, 1450d, 1470d, and 1490d may include results corresponding to the fourth image from the fourth viewpoint; and the fifth units 1410e, 1430e, 1450e, 1470e, and 1490e may include results corresponding to the fifth image from the fifth viewpoint.

[0354] As mentioned above, the corresponding units in the first queue 1410, the second queue 1430, the third queue 1450, the fourth queue 1470, and the fifth queue 1490 may include the judgment result for the image acquired at the same time.

[0355] In one embodiment, the first queue 1410 may include multiple units, each containing a standard deviation value of a heatmap corresponding to an retrieved image. Such a heatmap may be obtained from a model (e.g., a detection model) for detecting an eye region or a target point within an eye region. For example, the larger the standard deviation of the heatmap corresponding to the retrieved image, the larger the value contained in the unit of the first queue 1410 may be. For example, the standard deviation value of a heatmap represents a quantitative value, and a large such value may indicate a high probability of occlusion in the image. For example, if the standard deviation of the heatmap in the first unit 1410a is 0.5 and the threshold assigned to the first unit 1410a is 0.6, then the first image contained in the first unit 1410a may be determined to have passed the standard deviation-related criterion. For example, the standard deviation of the heatmap of the fourth image corresponding to the fourth unit 1410d may be determined to be greater than the threshold corresponding to the fourth unit 1410d. As a result, the computing device 100 can set the value of the fourth unit 1450d in the third queue 1450 to 0 without performing any analysis or processing on the fourth image in the second queue 1430. The computing device 100 may also determine that the fourth image is in an "open eyes" state. In the example in Figure 14, if the standard deviation of the heatmap (e.g., a value of 0.7) exceeds a critical criterion (e.g., a value of 0.6), the computing device 100 may determine that there is occlusion in the fourth image corresponding to the fourth unit, that the eyes are open, or that the occlusion is above a threshold. Examples of such occlusion may include sunglasses. For example, if it is determined that there is occlusion in the fourth image, or that the occlusion is above a threshold, the computing device 100 may decide not to calculate an occlusion score for the fourth image.For example, if it is determined that a hilt exists in the fourth image, or that the hilt is greater than a threshold, the computing device 100 may decide not to detect a target point corresponding to the eye area for the fourth image. For example, if it is determined that a hilt exists in the fourth image, or that the hilt is greater than a threshold, the computing device 100 may decide not to generate a drowsiness alarm for the fourth image. For example, if it is determined that a hilt exists in the fourth image, or that the hilt is greater than a threshold, the computing device 100 may decide that the driver is not closing their eyes for the fourth image. If the standard deviation exceeds a threshold, the computing device 100 may determine the primary prediction result of eye closure in the image to be an open state, regardless of the eye closure score determined using the transformed position information.

[0356] In one embodiment, the thresholds (thresholds compared to the standard deviation of the heatmap) for each of the units 1410a, 1410b, 1410c, 1410d, and 1410e of the first queue 1410 can be set to predetermined fixed values. In one embodiment, the thresholds (thresholds compared to the standard deviation of the heatmap) for each of the units 1410a, 1410b, 1410c, 1410d, and 1410e of the first queue 1410 may be set to have the same or different values ​​for each unit.

[0357] In one embodiment, the second cue 1430 may include a plurality of units 1430a, 1430b, 1430c, 1430d, and 1430e, which include values ​​related to the closed-eyes score. The closed-eyes score may be determined, for example, based on the measurement of the distance between the eyes in the transformed space. As a non-restrictive example, the measurement of the distance between the eyes may represent a value based on the distance between the upper and lower boundaries of the eye region. As an example, the measurement of the distance between the eyes and the closed-eyes score may have a negative correlation with each other. In such an example, a larger value for the measurement of the distance between the eyes may result in a smaller closed-eyes score. A larger value for the closed-eyes score here may indicate a higher probability of closed eyes.

[0358] In one embodiment, thresholds may be set for each of the multiple units 1430a, 1430b, 1430c, 1430d, and 1430e of the second queue 1430. For example, the thresholds for each of the multiple units 1430a, 1430b, 1430c, 1430d, and 1430e may be variably determined based on the blind-closing scores of one or more previous units of the unit in question. For example, the thresholds for each of the multiple units 1430a, 1430b, 1430c, 1430d, and 1430e may be variable depending on the values ​​included in the previous units of the unit in question. For example, the thresholds for each of the multiple units 1430a, 1430b, 1430c, 1430d, and 1430e may be determined based on the primary blind-closing prediction results of one or more previous units of the unit in question. For example, the thresholds corresponding to each of the multiple units 1430a, 1430b, 1430c, 1430d, and 1430e may be determined based on the closed-eye secondary prediction results of one or more previous units of the unit in question. The method for determining the thresholds corresponding to each of the units 1430a, 1430b, 1430c, 1430d, and 1430e in the second queue 1430 may correspond to the method for determining the threshold for determining the primary prediction results of the anomaly described above.

[0359] In Figure 14, the closed-eyes score may correspond to the model output information, the primary closed-eyes prediction result may correspond to the primary anomaly prediction result, and the secondary closed-eyes prediction result may correspond to the secondary anomaly prediction result.

[0360] In one embodiment, the third queue 1450 may include a primary eye-close prediction result determined based on a comparison of the image's eye-close score with a threshold. The first unit 1450a includes a first primary eye-close prediction result corresponding to the first image of the first viewpoint, the second unit 1450b includes a second primary eye-close prediction result corresponding to the second image of the second viewpoint, the third unit 1450c includes a third primary eye-close prediction result corresponding to the third image of the third viewpoint, the fourth unit 1450d includes a fourth primary eye-close prediction result corresponding to the fourth image of the fourth viewpoint, and the fifth unit 1450e includes a fifth primary eye-close prediction result corresponding to the fifth image of the fifth viewpoint.

[0361] In Figure 14, the first closed-eyes primary prediction result corresponding to the first image may be determined by comparing a first closed-eyes score with a value of 110 with a threshold value of 100. The first closed-eyes primary prediction result can have a value of 1 because the first closed-eyes score exceeds the corresponding threshold. The second closed-eyes primary prediction result corresponding to the second image may be determined by comparing a second closed-eyes score with a value of 110 with a threshold value of 100. The second closed-eyes primary prediction result can have a value of 1 because the second closed-eyes score exceeds the corresponding threshold. The third closed-eyes primary prediction result corresponding to the third image may be determined by comparing a third closed-eyes score with a value of 100 with a threshold value of 90. The third closed-eyes primary prediction result can have a value of 1 because the third closed-eyes score exceeds the corresponding threshold. The fourth closed-eyes primary prediction result corresponding to the fourth image can be set to a value of 0 because the standard deviation of the heatmap in the fourth unit 1410d of the first queue 1410 exceeded the threshold. The fifth closed-eyes primary prediction result corresponding to the fifth image may be determined by comparing a fifth closed-eyes score with a value of 100 with a threshold with a value of 90. The fifth closed-eyes primary prediction result can have a value of 1 because the fifth closed-eyes score exceeds the corresponding threshold.

[0362] In one embodiment, the values ​​of units 1470a, 1470b, 1470c, 1470d, and 1470e of the fourth queue 1470 may be determined based on a first voting that utilizes the values ​​of units 1450a, 1450b, 1450c, 1450d, and 1450e contained in the third queue 1450. Units 1470a, 1470b, 1470c, 1470d, and 1470e of the fourth queue 1470 may include closed-eye secondary prediction results. Such closed-eye secondary prediction results may correspond to anomaly secondary prediction results. The values ​​contained in the units of the fourth queue 1470 may be determined based on mainstream values, representative values, or ratios of closed-eye primary prediction results. For example, the fifth unit 1470e of the fourth queue 1470 may be determined based on a processing result that includes the previous unit and the fifth image corresponding to the fifth unit 1470e. For example, the fifth unit 1470e of the fourth queue 1470 may be determined based on representative, ratio, or mainstream values ​​of the values ​​of units 1450a, 1450b, 1450c, 1450d, and 1450e of the third queue 1450. For example, the fifth unit 1470e of the fourth queue 1470 may be determined based on a comparison of the ratio of units having a value of 1 in the values ​​of units 1450a, 1450b, 1450c, 1450d, and 1450e of the third queue 1450 with a threshold ratio.

[0363] In one embodiment, the fifth queue 1490 may correspond to a queue for generating an alarm. In one embodiment, the fifth queue 1490 may determine the alarm using the closed-eyes secondary prediction result.

[0364] In one embodiment, the values ​​of units 1490a, 1490b, 1490c, 1490d, and 1490e of the fifth queue 1490 may be determined based on a second voting that utilizes the values ​​of the units of the fourth queue 1470. The fifth queue 1490 may include the result of a second voting that determines whether or not to generate an alarm, taking into account the continuity of the unit values ​​of the fourth queue 1470. For example, assume a criterion of 3 for the continuity of the unit values. Under such assumption, since neither the first unit 1470a of the fourth queue 1470 nor the two preceding units have a value of 1, the first unit 1490a of the fifth queue 1490 may be set to 0 (i.e., alarm off). Since neither the second unit 1470b of the fourth queue 1470 nor the two preceding units (1470a and earlier) have a value of 1, the second unit 1490b of the fifth queue 1290 may be set to 0 (i.e., alarm off). Since neither the third unit 1490c of the fourth queue 1470 nor the two preceding units 1470a and 1470b have a value of 1, the third unit 1490c of the fifth queue 1490 may be set to 0 (i.e., alarm off). Since neither the fourth unit 1470d of the fourth queue 1470 nor the two preceding units 1470b and 1470c have a value of 1, the fourth unit 1490c of the fifth queue 1490 may be set to 0 (i.e., alarm off). Since neither the fifth unit 1470e of the fourth queue 1470 nor the two preceding units 1470c and 1470d have a value of 1, the fifth unit 1490e of the fifth queue 1490 may be set to 1 (i.e., alarm on).

[0365] As described above, the technology according to one embodiment of the present disclosure can utilize multiple cues to determine the presence or absence of eye closure or drowsiness, or to generate an alarm corresponding to eye closure or drowsiness. The technology according to one embodiment of the present disclosure can provide an accurate alarm in response to eye closure by utilizing multiple cues and one or more voting. Values ​​included in other cues (e.g., previous cues) among such multiple cues may be used to determine the values ​​included in a particular cue or the threshold assigned to a particular cue. This ensures not only the accuracy of determining the presence or absence of eye closure or drowsiness, but also the accuracy of alarm generation.

[0366] Figure 15 illustrates an example of a method for generating a closed-eye alarm or drowsiness alarm according to one embodiment of the present disclosure.

[0367] In the examples shown in Figure 15, the previously mentioned examples will be replaced with the previously mentioned content to avoid repetition of explanations.

[0368] In one embodiment, the first queue 1510, the second queue 1530, the third queue 1550, the fourth queue 1570, and the fifth queue 1590 may correspond to the first queue 1410, the second queue 1430, the third queue 1450, the fourth queue 1470, and the fifth queue 1490 in Figure 14, respectively.

[0369] In the first queue 1510, the second queue 1530, the third queue 1550, the fourth queue 1570, and the fifth queue 1590, the units corresponding to each other may contain processing results for the same image.

[0370] In the example in Figure 15, the first units 1510a, 1530a, 1550a, 1570a, and 1590a may include results corresponding to the first image at the first time point; the second units 1510b, 1530b, 1550b, 1570b, and 1590b may include results corresponding to the second image at the second time point; the third units 1510c, 1530c, 1550c, 1570c, and 1590c may include results corresponding to the third image at the third time point; the fourth units 1510d, 1530d, 1550d, 1570d, and 1590d may include results corresponding to the fourth image from the fourth viewpoint; and the fifth units 1510e, 1530e, 1550e, 1570e, and 1590e may include results corresponding to the fifth image from the fifth viewpoint.

[0371] In one embodiment, the first queue 1510 may include multiple units, each containing a standard deviation value of a heatmap corresponding to an retrieved image. Such a heatmap may be obtained from a model (e.g., a detection model) for detecting eye regions or target points within eye regions. For example, the larger the standard deviation of the heatmap corresponding to the retrieved image, the larger the value contained in the unit of the first queue 1510. For example, the standard deviation value of a heatmap represents a quantitative value, and a large such value can indicate a high probability of the presence of hidden areas in the image. For example, the standard deviation of the heatmap in the first unit 1510a is 0.7, and the threshold assigned to the first unit 1510a is 0.6, so the first image contained in the first unit 1510a may be determined to be outside the standard deviation-related criteria. Thus, the computing device 100 may process the value of the first unit 1550a in the third queue 1550 as 0 for the first image without performing any analysis or processing in the second queue 1530. As a result, the first image may be determined to be an open-eyed state by the first unit 1550a of the third queue 1550. The computing device 100 can determine that there is an obstruction in the first image corresponding to the first unit 1510a, that the eyes are open, or that the obstruction is above a threshold. Examples of such obstructions may include sunglasses. As an example, the computing device 100 may determine that it does not detect a target point corresponding to the eye area in the first image. As another example, the computing device 100 may determine that it does not generate a drowsiness alarm for the first image. As yet another example, the computing device 100 may determine that the driver is not closing their eyes in the first image. If the standard deviation exceeds a threshold, the computing device 100 may determine the primary predicted eye-closed result in the image to be an open-eyed state, regardless of the eye-closed score determined using the transformed positional information.If the standard deviation exceeds a threshold, the computing device 100 may determine the primary predicted eye-closed state for the image to be an open-eyed state, regardless of the eye-distance measurement result determined using the converted position information.

[0372] In one embodiment, since the other units 1550b, 1550c, 1550d, and 1550e of the first queue 1510 have standard deviation values ​​that are less than or equal to a threshold, data processing in the second queue 1530, which shows the measurement results of the distance between the eyes, can be performed on the images corresponding to the other units 1550b, 1550c, 1550d, and 1550e.

[0373] In one embodiment, the thresholds (thresholds compared to the standard deviation of the heatmap) for each of the units 1510a, 1510b, 1510c, 1510d, and 1510e of the first queue 1510 may be set to predetermined fixed values. In one embodiment, the thresholds (thresholds compared to the standard deviation of the heatmap) for the units 1510a, 1510b, 1510c, 1510d, and 1510e of the first queue 1510 may also be set to have the same or different values ​​for each of the units.

[0374] In one embodiment, the second cue 1530 may include a plurality of units 1530a, 1530b, 1530c, 1530d, and 1530e, each containing a value related to the measurement result of the eye distance. The measurement result of the eye distance can, for example, represent the measurement result of the eye distance in a transformed space (e.g., a normalized space). As an unrestrictive example, the measurement result of the eye distance can represent a value based on the distance between the upper and lower boundaries of the eye region. As an example, the measurement result of the eye distance and the closed-eye score can have a negative correlation with each other. In such an example, a larger value for the measurement result of the eye distance can result in a smaller closed-eye score. A smaller measurement result of the eye distance here indicates a higher probability of closed eyes. A larger value for the closed-eye score here indicates a higher probability of closed eyes.

[0375] In one embodiment, thresholds may be set for each of the multiple units 1530a, 1530b, 1530c, 1530d, and 1530e of the second queue 1530. For example, the thresholds for each of the multiple units 1530a, 1530b, 1530c, 1530d, and 1530e may be variably determined based on the measurement results of the eye distance of one or more previous units of the unit in question. For example, the thresholds for each of the multiple units 1530a, 1530b, 1530c, 1530d, and 1530e may be variable depending on the values ​​included in the previous units of the unit in question. For example, the thresholds for each of the multiple units 1530a, 1530b, 1530c, 1530d, and 1530e may be determined based on the eye-closed primary prediction results of one or more previous units of the unit in question. For example, the thresholds corresponding to each of the multiple units 1530a, 1530b, 1530c, 1530d, and 1530e may be determined based on the closed-eye secondary prediction results of one or more previous units of the unit in question. The method for determining the thresholds corresponding to each of the units 1530a, 1530b, 1530c, 1530d, and 1530e in the second queue 1530 may correspond to the method for determining the threshold for determining the primary prediction result of the anomaly described above.

[0376] The eye distance measurement results in Figure 15 may conceptually correspond to the model output information. However, a difference from the model output information may be that if the eye distance measurement result is smaller than the threshold, the primary eye-closed prediction result is assigned as "eyes closed" (i.e., a value of 1). The primary eye-closed prediction result in Figure 15 may correspond to the primary anomaly prediction result, and the secondary eye-closed prediction result may correspond to the secondary anomaly prediction result.

[0377] The units 1550a, 1550b, 1550c, 1550d, and 1550e of the third queue 1550 in Figure 15 may include a primary eye-closed prediction result indicating whether or not the eyes are closed. If the primary eye-closed prediction result is 1, the eyes are considered to be open; if the primary eye-closed prediction result is 0, the eyes are considered to be closed. If the value of the second queue 1530 is less than the threshold, the primary eye-closed prediction result may have a value of 1. If the value of the second queue 1530 is greater than or equal to the threshold, the primary eye-closed prediction result may have a value of 0.

[0378] The values ​​of units 1570a, 1570b, 1570c, 1570d, and 1570e of the fourth queue 1570 in Figure 15 may be determined by a first voting that utilizes the closed-eye primary prediction results. The first voting may generate the closed-eye secondary prediction results using the mainstream, representative, or proportion values ​​of the closed-eye primary prediction results.

[0379] Units 1570a, 1570b, 1570c, 1570d, and 1570e of the fourth queue 1570 may include a closed-eyes secondary prediction result used as a parameter for determining a closed-eyes alarm or a drowsiness alarm. If the closed-eyes secondary prediction result is 1, it can be considered an open-eyes state, and if the closed-eyes secondary prediction result is 0, it can be considered a closed-eyes state. The concept of a counter in which the value of the closed-eyes secondary prediction result is maintained or increased if the primary voting result is 1, and maintained or decreased if the primary voting result is 1, can be applied to one embodiment of the present disclosure. If the value of the closed-eyes secondary prediction result falls outside a preset maximum or minimum value, the closed-eyes secondary prediction result can be replaced with the maximum or minimum value. In the example of Figure 15, the maximum value of the closed-eyes secondary prediction result may be set to 3 and the minimum value to 0.

[0380] The units 1590a, 1590b, 1590c, 1590d, and 1590e of the fifth queue 1590 in Figure 15 may have values ​​to determine whether or not to generate an alarm, or the type or intensity of the alarm. The values ​​of the units 1590a, 1590b, 1590c, 1590d, and 1590e of the fifth queue 1590 may be determined based on a second voting that performs a comparison between the corresponding units of the fourth queue 1590 and one or more thresholds. As an example, the second voting may perform a comparison between the closed-eyes quadratic prediction result to a threshold to which the concept of a counter is applied. As another example, the second voting may include determining whether or not there is continuity in the closed-eyes quadratic prediction result.

[0381] In the example in Figure 15, the threshold used in the second voting can be exemplified as 2. Therefore, since the value (2) of the fourth unit 1570d of the fourth queue 1570 is greater than or equal to the threshold (2), the value of the fourth unit 1590e of the fifth queue 1590 may be determined to be ON. For example, the alarm can be turned OFF when the values ​​of the units from the fifth unit 1570e onward become 0 and the counter value becomes 2. As another example, the alarm can be kept ON when the values ​​of the units from the fifth unit 1570e onward become 0 and the counter value becomes 2, and then turned OFF when the values ​​of subsequent units become 0 and the counter value becomes 1.

[0382] As mentioned above, multiple counters and / or multiple alarm thresholds can also be applied in Figure 15, thereby enabling the setting of alarms under various conditions.

[0383] Figure 16 is a schematic diagram of the computing environment of a computing device 100 according to one embodiment of the present disclosure.

[0384] In this disclosure, computing devices, computing apparatus, computers, systems, components, modules, or units include routines, procedures, programs, components, data structures, etc., that perform a particular task or realize a particular type of abstract data. Furthermore, a person skilled in the art will readily recognize that the methods presented in this disclosure can be implemented in other computer system configurations, including single-processor or multi-processor computing devices, minicomputers, mainframe computers, as well as personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, and others (each of which may operate in conjunction with one or more related devices).

[0385] The embodiments described herein can also be implemented in a distributed computing environment in which a task is performed by remote processing units connected via a communication network. In a distributed computing environment, program modules can reside in both local and remote memory storage devices.

[0386] Computing devices typically include various computer-readable media. Any media accessible by a computer can be computer-readable, and such computer-readable media include volatile and non-volatile media, transient and non-transitory media, and mobile and non-mobile media. As an unrestricted example, computer-readable media may include computer-readable storage media and computer-readable transmission media.

[0387] Computer-readable storage media include volatile and non-volatile media, temporary and non-temporary media, mobile and non-mobile media, implemented in any way or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD (digital video disk) or other optical disc storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other media that can be accessed by a computer and used to store desired information.

[0388] Computer-readable transmission media typically include all information transmission media that implement computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism. The term modulated data signal means a signal whose one or more characteristics have been set or modified in order to encode information within the signal. As an unrestricted example, computer-readable transmission media include wired media such as wired networks or direct-wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the aforementioned media is also included within the scope of computer-readable transmission media.

[0389] An exemplary environment 2000 that realizes various aspects of the present invention, including a computer 2002, is shown, the computer 2002 including a processing unit 2004, system memory 2006, and a system bus 2008. The computer 200 in this specification may be used in a manner that is interoperable with computing device 100. The system bus 2008 connects system components, including (but not limited to) system memory 2006, to the processing unit 2004. The processing unit 2004 may be any processor from a variety of commonly used processors. Dual-processor and other multi-processor architectures can also be used as the processing unit 2004.

[0390] The system bus 2008 may be any of several types of bus structures that can be further interconnected to a memory bus, a peripheral bus, and a local bus using any of the various common bus architectures. System memory 2006 includes read-only memory (ROM) 2010 and random access memory (RAM) 2012. The basic input / output system (BIOS) is stored in non-volatile memory 2010, such as ROM, EPROM, or EEPROM, and this BIOS includes basic routines that help transfer information between components within the computer 2002, such as during startup. RAM 2012 may also include high-speed RAM, such as static RAM, for caching data.

[0391] Computer 2002 also includes an internal hard disk drive (HDD) 2014 (e.g., EIDE, SATA), a magnetic floppy disk drive (FDD) 2016 (e.g., for reading from or writing to a portable diskette 2018), an SSD, and an optical disk drive 2020 (e.g., for reading CD-ROM disks 2022, or for reading from or writing to other high-capacity optical media such as DVDs). The hard disk drive 2014, magnetic disk drive 2016, and optical disk drive 2020 may be coupled to the system bus 2008 by a hard disk drive interface 2024, a magnetic disk drive interface 2026, and an optical drive interface 2028, respectively. Interfaces 2024 for implementing external drives include, for example, at least one or both of USB (Universal Serial Bus) and IEEE 1394 interface technologies.

[0392] These drives and associated computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, and the like. In the case of Computer 2002, the drives and media correspond to storing any data in a suitable digital format. While the above description of computer-readable storage media refers to HDDs, portable magnetic disks, and portable optical media such as CDs or DVDs, those skilled in the art will understand that other types of computer-readable storage media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, and the like, can also be used in the exemplary operating environment, and that any such media may contain computer-executable instructions for performing the methods of the present invention.

[0393] Multiple program modules, including an operating system 2030, one or more application programs 2032, other program modules 2034, and program data 2036, can be stored in the drive and RAM 2012. All or part of the operating system, applications, modules, and / or data can also be cached in RAM 2012. It is understood that the present invention can be implemented with various commercially available operating systems or combinations of operating systems.

[0394] The user can input commands and information to the computer 2002 via one or more wired / wireless input devices, such as a keyboard 2038 and a pointing device such as a mouse 2040. Other input devices (not shown) include microphones, IR remote controls, joysticks, gamepads, stylus pens, touchscreens, and others. These and other input devices are often connected to the processing unit 2004 via an input device interface 2042 connected to the system bus 2008, but may also be connected via other interfaces such as parallel ports, IEEE 1394 serial ports, game ports, USB ports, IR interfaces, and others.

[0395] Monitor 2044 or other types of display devices are also connected to system bus 2008 via interfaces such as video adapter 2046. In addition to monitor 2044, the computer generally includes other peripheral output devices (not shown) such as speakers, printers, and others.

[0396] Computer 2002 may operate in a networked environment using logical connections to one or more remote computers, such as remote computer 2048, via wired and / or wireless communication. Remote computer 2048 may be a workstation, server computer, router, personal computer, portable computer, microprocessor-based entertainment device, peer device, or other ordinary network node, and generally include many or all of the components described for computer 2002, except for memory storage device 205, which is shown for simplicity. The illustrated logical connections include wired / wireless connections to a near-field communication network (LAN) 2052 and / or a larger network, such as a far-field communication network (WAN) 2054. Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks such as intranets, all of which may be connected to a global computer network, such as the Internet.

[0397] When used in a LAN networking environment, computer 2002 is connected to the local network 2052 via a wired and / or wireless network interface or adapter 2056. Adapter 2056 facilitates wired or wireless communication to LAN2052, which also includes a wireless access point installed therein for communication with the wireless adapter 2056. When used in a WAN networking environment, computer 2002 is connected to a communication server on WAN2054, including a modem 2058, or has other means of establishing communication over WAN2054, such as via the Internet. The modem 2058, which may be internal or external, and wired or wireless, is connected to the system bus 2008 via a serial port interface 2042. In a networked environment, program modules or parts thereof described for computer 2002 may be stored in remote memory / storage device 2050. The illustrated network connections are illustrative, and it should be understood that other means may be used to establish communication links between computers.

[0398] Computer 2002 operates to communicate with any wireless device or individual that is located and operates wirelessly, such as a printer, scanner, desktop and / or portable computer, PDA (portable data assistant), communication satellite, any device or location associated with a wirelessly discoverable tag, and a telephone. This includes at least Wi-Fi and Bluetooth wireless technologies. Thus, the communication may be a predefined structure, like a conventional network, or simply ad hoc communication between at least two devices.

[0399] It is understood that the specific order or hierarchical structure of the process stages presented is an example of an exemplary approach. It is understood that the specific order or hierarchical structure of the stages within the process can be rearranged within the scope of this disclosure based on design priorities. The claims of the methods in this disclosure provide elements of various stages in sample order, but are not limited to the specific order or hierarchical structure presented.

[0400] These and other modifications may be added to the embodiments in light of the detailed description above. In general, the terms used in the following claims should not be construed as limiting the claims to any specific embodiment disclosed in the specification and claims, but rather as encompassing all possible embodiments, along with the entire scope of the rightsable equivalents that such claims have. Accordingly, the claims are not limited by this disclosure. (Embodiments of the Invention)

[0401] As mentioned above, the relevant information is described in the best mode for carrying out the invention. [Industrial applicability]

[0402] This disclosure can be applied to a method, computer program, and / or computing device for providing alarms based on driver behavior in a driver monitoring system.

Claims

1. A method for providing alarms based on driver behavior in a Driver Monitoring System (DMS), which is performed by a computing device, wherein the method is: Receiving a first image including the driver inside the vehicle, Using an artificial intelligence model, first model output information is generated from the first image indicating the possibility that an anomaly that impairs the driver's safety exists in the first image. The first model output information is compared with a first threshold to generate a primary prediction result of the first anomaly indicating whether or not the anomaly exists in the first image. By performing a first voting using the primary prediction result of the first anomaly, a secondary prediction result of the first anomaly is generated; and By performing a second voting using the secondary prediction result of the first anomaly, an anomaly alarm corresponding to the first image is determined. Includes, The artificial intelligence model, The system takes an image of the driver as input and outputs the possibility that a predetermined abnormality that would impair the driver's safety exists in the driver's image, corresponding to a pre-trained deep learning-based model. method.

2. The aforementioned predetermined abnormalities are This includes a first abnormality corresponding to not wearing a seatbelt, a second abnormality corresponding to distraction while driving, and a third abnormality corresponding to drowsiness while driving. The first model output information includes a quantitative value for comparison with the first threshold, and if the first model output information is greater than or equal to the first threshold, the first prediction result for the first anomaly indicates that the anomaly exists in the first image, and if the first model output information is less than the first threshold, the first prediction result for the first anomaly indicates that the anomaly does not exist in the first image. The secondary prediction result of the first anomaly represents a quantitative value used as a parameter for determining the anomaly alarm corresponding to the first image. The method according to claim 1.

3. The first threshold is, This is determined based on the primary prediction result of at least one previous anomaly corresponding to at least one previously acquired image of the first image, In response to a previously acquired second image of the first image, it is determined based on the second model output information generated by the model, or, The second threshold is determined based on the primary prediction result of the second anomaly, which is obtained by comparing the second threshold, which is determined based on the primary prediction result of the third anomaly corresponding to the third image previously acquired from the second image, with the second model output information. Here, the second threshold is a threshold for determining whether or not the abnormality exists in the second image. The method according to claim 1.

4. The first threshold is determined based on the second predicted result of the second anomaly corresponding to the second image previously acquired for the first image, and the second predicted result of the second anomaly is obtained by performing the first voting using the first predicted result of the second anomaly corresponding to the second image. The method according to claim 1.

5. The first threshold is, The first image is determined based on the ratio of result values ​​indicating the presence of the anomaly, from the primary prediction results of previous anomalies corresponding to a predetermined first number of previously acquired images of the first image. In the aforementioned initial prediction results for anomalies, if the ratio of result values ​​indicating the presence of the anomaly is equal to or greater than the first ratio, the first threshold is set to the first value. If, in the aforementioned primary prediction results for an anomaly, the ratio of result values ​​indicating the presence of the anomaly is less than the first ratio, the first threshold is set to a second value that is higher than the first value. The method according to claim 1.

6. Generating the secondary prediction result for the first anomaly is, Determining the majority value of the primary prediction result of the anomaly corresponding to the first image and a predetermined second number of previously acquired images of the first image, and Using the determined mainstream value, generate the secondary prediction result for the first anomaly. Includes, The aforementioned mainstream value is, Of the result values ​​indicating the presence of the anomaly and the result values ​​indicating the absence of the anomaly, the result value that shows a higher proportion than the primary prediction result of the anomaly is determined. The method according to claim 1.

7. Generating the secondary prediction result for the first anomaly is, Based on the results of the first voting, the first current counter value and the second current counter value corresponding to the first image are determined by changing each of the multiple previous counter values ​​corresponding to the second image previously acquired for the first image, and To generate a secondary prediction result of the first anomaly, including the first current counter value and the second current counter value, It includes, and, Determining the abnormal alarm corresponding to the first image is: By comparing the first current counter value with the first counter threshold, it is determined whether to turn the first abnormal alarm corresponding to the first image ON or OFF, and By comparing the second current counter value with the second counter threshold, it is determined whether to turn the second abnormal alarm corresponding to the first image ON or OFF. including, The method according to claim 1.

8. Generating the secondary prediction result for the first anomaly is, If the result of the first voting, which represents the mainstream value of the primary prediction result of an anomaly or the prediction result of a cluster anomaly in a predetermined set of images, indicates the presence of an anomaly, then it is decided to increment at least one previous counter value corresponding to a second image previously acquired for the first image, or If the result of the first voting, which represents the mainstream value of the primary prediction result of anomalies or the prediction result of aggregate anomalies in a predetermined set of images, indicates the absence of anomalies, then it is decided to decrease the at least one previous counter value. including, The method according to claim 1.

9. Generating the secondary prediction result for the first anomaly is, Based on the results of the first voting, the current counter value corresponding to the first image is determined by changing the previous counter value corresponding to the second image previously acquired for the first image, and To generate a secondary prediction result of the first anomaly, including the current counter value, It includes, and also, Determining the abnormal alarm corresponding to the first image is: By comparing the current counter value with the third counter threshold, it is determined whether to turn the third abnormal alarm corresponding to the first image ON or OFF, and By comparing the current counter value with the fourth counter threshold, it is determined whether to turn the fourth abnormal alarm corresponding to the first image ON or OFF. including, The method according to claim 1.

10. Generating the secondary prediction result for the first anomaly is, Based on the results of the first voting, determine at least one current counter value corresponding to the first image by changing at least one previous counter value corresponding to a second image previously acquired for the first image. Includes, Here, if the current counter value falls outside the range defined by a predetermined minimum and maximum value, the current counter value is determined by the minimum or maximum value, and the unit of change of the at least one previous counter value is determined based on the time difference between the acquisition time of the second image and the acquisition time of the first image. The method according to claim 1.

11. The aforementioned second boating is, In order to ensure accuracy in predicting the occurrence of the alarm, a set of images consisting of a predetermined third number of sequential images including the first image is determined, and it is determined whether or not there is continuity in the secondary prediction results of anomalies corresponding to the images constituting the set of images, where the third number of sequential images includes the first image and images acquired prior to the first image. Determining the abnormal alarm corresponding to the first image is: If all secondary prediction results for anomalies within an image set composed of a predetermined third number of sequential images including the first image indicate the presence of the anomaly, it is decided to generate the anomaly alarm corresponding to the first image. It includes, and also, The third sequential image includes the first image and an image previously acquired of the first image. The method according to claim 1.

12. Determining the abnormal alarm corresponding to the first image is: Based on the secondary prediction result of the first anomaly corresponding to the first image and the secondary prediction results of previous anomalies corresponding to each of the predetermined fourth images prior to the first image, it is determined whether or not to generate the anomaly alarm corresponding to the first image. including, The method according to claim 1.

13. The first voting utilizes the primary prediction result of a previous anomaly corresponding to at least one previously received previous image of the first image and the primary prediction result of the first anomaly, The second voting utilizes a combination of the secondary prediction result of a previous anomaly corresponding to the at least one previous image and the secondary prediction result of the first anomaly, or compares the secondary prediction result of the first anomaly with a counter threshold. The method according to claim 1.

14. A computer program stored on a computer-readable storage medium, wherein, when the computer program is executed by at least one processor, the at least one processor allows the Driver Monitoring Systems (DMS) to perform an operation to provide an alarm in response to the driver's actions, and the operation is Receiving a first image including the driver inside the vehicle, Using an artificial intelligence model, first model output information is generated from the first image indicating the possibility that an abnormality that could impair the driver's safety exists in the first image. The first model output information is compared with a first threshold to generate a first anomaly prediction result indicating whether or not the anomaly exists in the first image. By performing a first voting using the primary prediction result of the first anomaly, a secondary prediction result of the first anomaly is generated, and By performing a second voting using the secondary prediction result of the first anomaly, an anomaly alarm corresponding to the first image is determined. The system takes an image of the driver as input and outputs the possibility that a predetermined abnormality that would impair the driver's safety exists in the driver's image, corresponding to a pre-trained deep learning-based model. including, A computer program stored on a computer-readable storage medium.

15. A computing device, at least one processor; and Memory; Includes, The aforementioned at least one processor is Receiving a first image including the driver inside the vehicle, Using an artificial intelligence model, first model output information is generated from the first image indicating the possibility that an abnormality that could impair the driver's safety exists in the first image. The first model output information is compared with a first threshold to generate a first anomaly prediction result indicating whether or not the anomaly exists in the first image. By performing a first voting using the primary prediction result of the first anomaly, a secondary prediction result of the first anomaly is generated, and By performing a second voting using the secondary prediction result of the first anomaly, an anomaly alarm corresponding to the first image is determined. Execute The system takes an image of the driver as input and outputs the possibility that a predetermined abnormality that would impair the driver's safety exists in the driver's image, corresponding to a pre-trained deep learning-based model. Computing device.