Artificial intelligence-based behavior monitoring method, program, and device
The AI-based method uses a deep learning model to analyze multiple sensing items and a rule set for accurate behavior monitoring, addressing the limitations of existing technologies in online environments by improving detection of actions like cheating.
Patent Information
- Application Number
- JP2025550406
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2022-12-19
- Publication Date
- 2025-11-26
AI Technical Summary
Existing technologies struggle to accurately monitor human behavior in online environments, particularly in online exams, due to limited information analysis, leading to potential misjudgment of cheating or failure to detect it.
An artificial intelligence-based method using a deep learning model to analyze various sensing items, including body parts, objects, and sounds, combined with a predetermined rule set to predict behavior, enabling comprehensive and accurate monitoring.
The method provides detailed and precise monitoring of human actions by integrating multiple sensing data sources, enhancing the accuracy of detecting behaviors like cheating in online exams.
Smart Images

Figure 2025538276000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to data analysis technology, and more particularly to a method and apparatus for inferring and monitoring human behavior based on complex judgment results based on artificial intelligence. [Background technology]
[0002] In environments built for specific purposes, there are situations where it is necessary to monitor the actions taken by people and analyze the consequences of those actions. For example, in an educational environment where exams are taken, it is necessary to monitor the actions taken by test takers during the exam. In particular, unlike offline exams, it is difficult to effectively monitor the actions of test takers and their surrounding environment in online exams. Therefore, in online exam environments, it is even more important for administrators to accurately analyze the actions taken by test takers in real time to determine whether cheating has occurred.
[0003] As can be seen from the above examples, it is not easy to effectively monitor a person's behavior and the surrounding environment in an online environment. While there are conventional technologies that analyze a person's specific behavior using a sensing device such as a camera, most of them analyze a person's specific behavior based only on fragmentary information acquired in a specific situation. However, such an analysis based only on fragmentary information makes it difficult to accurately interpret whether a person is engaging in behavior that requires judgment in a specific environment. For example, if cheating is detected by analyzing only a frontal image of a person in an online exam environment, the limited information that can be acquired from the frontal image increases the likelihood of failing to determine cheating even when cheating is suspected, or of misjudging cheating even when it is not. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure has been devised in response to the above-mentioned background art, and aims to provide a method and apparatus that can comprehensively determine and accurately monitor what actions a person will take in a specific environment based on various sensing results.
[0005] However, the problems to be solved by the present disclosure are not limited to those mentioned above, and other problems not mentioned will be clearly understood from the following description. [Means for solving the problem]
[0006] To achieve the above object, one embodiment of the present disclosure discloses an artificial intelligence-based behavior monitoring method performed by a computing device, the method including: generating an analysis result for the sensing item based on observation data of a person being a subject of behavior monitoring by using a deep learning model that matches at least one sensing item included in each of a plurality of sensing targets, the deep learning model being executed by the computing device including at least one processor; and predicting the person's behavior based on the generated analysis result by using a predetermined rule set.
[0007] Alternatively, the sensing item may be status information identified based on a subclass of the sensing object, and the status information may be changeable depending on the person's actions.
[0008] Alternatively, the plurality of sensing objects may include, in addition to the body parts of the person, at least one of an object other than the person, a sound of an object related to the person's actions, or a time of an object related to the person's actions.
[0009] Alternatively, the deep learning model may include at least one of a first model that estimates a person's pose based on an image, a second model that estimates a person's facial shape and direction based on an image, a third model that tracks a person's gaze based on an image, a fourth model that recognizes objects other than people based on an image, or a fifth model that detects sound elements of objects related to human actions based on at least one of an image or audio.
[0010] Alternatively, the step of using a deep learning model that matches at least one sensing item included in each of the plurality of sensing targets and generating an analysis result for the sensing item based on observation data of a person who is the subject of behavior monitoring may include the steps of acquiring the observation data at a predetermined period, and inputting the acquired observation data into at least one of the first model, the second model, the third model, the fourth model, or the fifth model, and generating an analysis result for the sensing item that reflects the behavioral results of the person performed during the predetermined period.
[0011] Alternatively, the predetermined period may be determined according to environmental conditions set via a client in charge of behavior monitoring.
[0012] Alternatively, the step of using a predetermined rule set to predict the person's behavior based on the generated analysis results may include the steps of: identifying, from the generated analysis results, analysis results that match judgment conditions for each behavior class included in the predetermined rule set; estimating accuracy of the identified analysis results and a correlation between the behavior classes included in the predetermined rule set and the identified analysis results; and combining the identified analysis results based on the estimated accuracy and correlation.
[0013] Alternatively, the step of combining the identified analysis results based on the estimated accuracy and correlation to infer the person's behavior may include the steps of: assigning a first weight to each of the identified analysis results based on the estimated accuracy; assigning a second weight to each of the identified analysis results based on the estimated correlation; and determining whether the person has performed at least one of the behavior classes included in the predetermined rule set based on a numerical value derived by combining the first weight and the second weight.
[0014] Alternatively, the behavior classes included in the predetermined rule set may include a first behavior class, set via a client in charge of behavior monitoring, corresponding to exam cheating, and a second behavior class, set via the client, corresponding to abnormal behavior that is unnecessary for taking the exam.
[0015] According to one embodiment of the present disclosure, there is provided a computer program stored on a computer-readable storage medium. When executed by one or more processors, the computer program performs operations for monitoring behavior based on artificial intelligence. Here, the operations include: generating an analysis result for the sensing item based on observation data of a person who is a target of behavior monitoring, using a deep learning model that matches at least one sensing item included in each of a plurality of sensing objects; and estimating the person's behavior based on the generated analysis result using a predetermined rule set.
[0016] To achieve the above-described object, one embodiment of the present disclosure discloses a computing device for monitoring behavior based on artificial intelligence. The device may include a processor including at least one core, a memory including program code executable by the processor, and a network unit for acquiring observation data of a person who is a target of behavior monitoring. Here, the processor may use a deep learning model that matches at least one sensing item included in each of a plurality of sensing objects to generate an analysis result for the sensing item based on the observation data, and may use a predetermined rule set to estimate the person's behavior based on the generated analysis result. [Effects of the Invention]
[0017] The present disclosure can provide a method and apparatus that can comprehensively determine and accurately monitor what actions a person will take in a specific environment based on various sensing results. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a block diagram of a computing device according to one embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating a process for performing behavioral monitoring of a computing device according to one embodiment of the present disclosure. [Figure 3] FIG. 10 is a block diagram illustrating a process for performing behavioral monitoring of a computing device according to an alternative embodiment of the present disclosure. [Figure 4] 4a is a table summarizing an analysis method and analysis results for each sensing item according to an embodiment of the present disclosure, and FIG. 4b is a table summarizing a rule set for behavior estimation and behavior estimation results according to an embodiment of the present disclosure. [Figure 5][Figure 5a] A conceptual diagram of a detailed estimation process for each activity of a computing device according to an embodiment of the present disclosure, [Figure 5b] A conceptual diagram of a detailed estimation process for each activity of a computing device according to an embodiment of the present disclosure, and [Figure 5c] A conceptual diagram of a detailed estimation process for each activity of a computing device according to an embodiment of the present disclosure. [Figure 6] 1 is a flowchart illustrating an artificial intelligence-based activity monitoring method according to one embodiment of the present disclosure. [Figure 7] 1 is a flowchart illustrating a method for monitoring behavior in an online testing environment according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0019] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. The embodiments presented in this disclosure are provided to enable those skilled in the art to use or practice the contents of the present disclosure. Therefore, various modifications to the embodiments of the present disclosure will be apparent to those skilled in the art. That is, the present disclosure may be embodied in various different forms and is not limited to the following embodiments.
[0020] Throughout the specification of the present disclosure, the same or similar reference numerals refer to the same or similar components. In addition, in order to clearly explain the present disclosure, reference numerals of parts that are not relevant to the explanation of the present disclosure may be omitted from the drawings.
[0021] The term "or" as used in this disclosure is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or otherwise clear from the context in this disclosure, "X uses A or B" should be understood to mean one of the natural inclusive permutations. For example, unless otherwise specified or otherwise clear from the context in this disclosure, "X uses A or B" can be interpreted as either X uses A, X uses B, or X uses both A and B.
[0022] The term "and / or" as used in this disclosure must be understood to indicate and include all possible combinations of one or more of the associated listed concepts.
[0023] The terms "comprises" and / or "comprising" as used in this disclosure should be understood to mean that the specified features and / or components are present. However, the terms "comprises" and / or "comprising" should not be understood to exclude the presence or addition of one or more other features, other components and / or combinations thereof.
[0024] In this disclosure, unless otherwise specified or clear from the context as referring to the singular form, the singular should generally be construed as including "one or more."
[0025] The term "nth (n is a natural number)" used in this disclosure can be understood as an expression used to distinguish components of the present disclosure from one another based on a predetermined criterion, such as functional, structural, or convenience of description. For example, in this disclosure, components that perform different functional roles can be classified as a first component or a second component. However, components that are substantially identical within the technical concept of the present disclosure but must be distinguished for convenience of description can also be classified as a first component or a second component.
[0026] The term "acquire" as used in this disclosure may be understood to refer to generating or receiving data in an on-device form, as well as receiving data from an external device or system via a wireless communication network.
[0027] Meanwhile, the terms "module" or "unit" used in this disclosure may be understood to refer to an independent functional unit that processes computing resources, such as a computer-related entity, firmware, software or a portion thereof, hardware or a portion thereof, or a combination of software and hardware. Here, a "module" or "unit" may refer to a unit composed of a single element or a unit expressed as a combination or collection of multiple elements. For example, as a concept of a "module," a "unit" may refer to a hardware element or a collection of hardware elements of a computing device, an application program that achieves a specific software function, a processing procedure implemented by the execution of software, or a collection of instructions for executing a program. Furthermore, as a broad concept, a "module" or "unit" may refer to a computing device itself that constitutes a system, or an application executed on a computing device. However, the above concepts are merely examples, and the concepts of a "module" or "unit" may be defined in various ways within the scope of understanding of those skilled in the art based on the contents of this disclosure.
[0028] The term "model" as used in this disclosure may be understood as a system implemented using mathematical concepts and language to solve a specific problem, a collection of software units for solving a specific problem, or an abstract model of a processing process for solving a specific problem. For example, a deep learning "model" may refer to a system generally implemented as a neural network that has problem-solving capabilities through learning. Here, a neural network may have problem-solving capabilities by optimizing parameters connecting nodes or neurons through learning. A deep learning "model" may include a single neural network or a neural network ensemble in which multiple neural networks are combined.
[0029] As used in this disclosure, the term "image" refers to multidimensional data composed of discrete image elements. In other words, "image" can be understood as a term referring to a digital representation of an object that can be viewed by the human eye. For example, "image" can refer to multidimensional data composed of elements that correspond to pixels in a two-dimensional image. "Image" can refer to multidimensional data composed of elements that correspond to voxels in a three-dimensional image.
[0030] The explanations of the above terms are intended to facilitate understanding of the present disclosure. Therefore, unless the above terms are explicitly stated as matters limiting the contents of the present disclosure, care should be taken not to use them in a way that limits the technical ideas of the contents of the present disclosure.
[0031] FIG. 1 is a block diagram of a computing device according to one embodiment of the present disclosure.
[0032] The computing device 100 according to an embodiment of the present disclosure may be a hardware device or part of a hardware device that performs comprehensive data processing and calculations, or may be a software-based computing environment connected via a communication network. For example, the computing device 100 may be a server that performs intensive data processing functions and shares resources, or a client that shares resources by interacting with the server. The computing device 100 may also be a cloud system that enables multiple servers and clients to interact with each other to comprehensively process data. The above description is merely an example of a type of computing device 100, and various types of computing device 100 may be configured within the scope that can be understood by those skilled in the art based on the contents of the present disclosure.
[0033] 1, a computing device 100 according to an embodiment of the present disclosure may include a processor 110, a memory 120, and a network unit 130. However, since FIG. 1 is merely an example, the computing device 100 may include other components for implementing a computer environment. Also, the computing device 100 may include only some of the disclosed components.
[0034] The processor 110 according to an embodiment of the present disclosure may be understood as a component including hardware and / or software for performing computing operations. For example, the processor 110 may read a computer program to perform data processing for machine learning. The processor 110 may process operations such as input data processing for machine learning, feature extraction for machine learning, and error calculation based on backpropagation. The processor 110 for performing such data processing may include a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc. The types of processor 110 described above are merely examples, and various types of processor 110 may be configured within the scope of what one skilled in the art would understand based on the present disclosure.
[0035] The processor 110 may use a pre-trained deep learning model to generate an analysis result for each of a plurality of sensing objects based on observation data of a person who is the target of activity monitoring. Here, the sensing object may be understood as a component of the observation data that serves as a reference for estimating a person's activity. The analysis result for the sensing object may be information indicating what behavior a person will perform based on the sensing object present in the observation data. Specifically, the sensing object may be any one of a body part of a person, an object other than a person, the sound of an object related to a person's activity, or the time of an object related to a person's activity. The object related to a person's activity may be a body part of a person or an object that can change depending on a person's activity. The analysis result for the sensing object may be information about the person's activity based on the body part of a person, an object, the sound of an object, or the time of an object present in the observation data. That is, the processor 110 may input the observation data into a pre-trained deep learning model and generate an analysis result for each sensing object present in the observation data as basic data for detecting a specific human behavior performed in a specific environment for activity monitoring.
[0036] The processor 110 may infer specific human behavior from analysis results generated through a deep learning model using a predetermined rule set. Here, the rule set may be a set of behavior classes that are candidates for detection in a specific environment for behavior monitoring and judgment conditions for each behavior class. The rule set may be created, changed, or modified by an administrator who has built the specific environment for behavior monitoring. That is, the processor 110 may comprehensively judge the analysis results for each detection target present in the observation data based on a rule set that can be customized for the specific environment for behavior monitoring, and infer human behavior present in the observation data. For example, assuming that the environment for behavior monitoring is an environment for an online exam, the rule set generated by the client of the exam proctor may be a set of judgment conditions for exam cheating and / or abnormal behavior that may be suspected of cheating, and for each of the cheating and / or abnormal behavior. The processor 110 may identify analysis results of the deep learning model that match the judgment conditions included in the rule set. The processor 110 may then combine the analysis results of the deep learning model that match the judgment criteria. Here, the combination of the analysis results that match the judgment criteria may be understood as an operation of performing a mathematical calculation based on the accuracy of each analysis result and a correlation with the fraudulent behavior or abnormal behavior. The processor 110 may determine whether the test taker has engaged in fraudulent behavior or abnormal behavior included in the rule set in a situation confirmed by the observation data, based on a numerical value derived by combining the analysis results of the deep learning model that match the judgment criteria.
[0037] In this way, the processor 110 can derive individual information indicating what actions a person will take based on various sensing targets using artificial intelligence, and can monitor specific human actions by comprehensively determining the individual information derived for each sensing target based on a rule set generated for a specific environment. In other words, the processor 110 performs monitoring by comprehensively considering all information obtainable from observation data, thereby enabling more detailed and accurate estimation of human actions to be detected in a specific environment, thereby providing an effective monitoring environment.
[0038] The memory 120 according to an embodiment of the present disclosure may be understood as a component including hardware and / or software for storing and managing data processed by the computing device 100. That is, the memory 120 may store any type of data generated or determined by the processor 110 and any type of data received by the network unit 130. For example, the memory 120 may include at least one type of storage medium selected from the group consisting of flash memory, hard disk, multimedia card micro, card-type memory, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. The memory 120 may also include a database system that manages data in a predetermined manner. The types of memory 130 described above are merely examples, and various configurations of the memory 120 are possible within the scope of what would be understood by one skilled in the art based on the present disclosure.
[0039] The memory 120 may structure and organize and manage data, a combination of data, and program code executable by the processor 110 required for the processor 110 to perform calculations. For example, the memory 120 may store observation data acquired via the network unit 130 (described later). The memory 120 may store program code that causes the processor 110 to learn a deep learning model, program code that causes the processor 110 to estimate human behavior using the learned deep learning model, and various data calculated by executing the program code.
[0040] The network unit 130 according to an embodiment of the present disclosure may be understood as a component that transmits and receives data via any type of known wired or wireless communication system. For example, the network unit 130 may transmit and receive data using a wired or wireless communication system such as a local area network (LAN), wideband code division multiple access (WCDMA), long term evolution (LTE), wireless broadband internet (WiBro), 5th generation mobile communication (5G), ultra wideband, ZigBee, radio frequency (RF) communication, wireless LAN, wireless fidelity (WiFi), near field communication (NFC), or Bluetooth. The above-described communication systems are merely examples, and various other wired or wireless communication systems for transmitting and receiving data by the network unit 130 may be used.
[0041] The network unit 130 may receive data necessary for the processor 110 to perform calculations via wired or wireless communication with any system, server, client, etc. The network unit 130 may also transmit data generated by calculations of the processor 110 via wired or wireless communication with any system, server, client, etc. The network unit 130 may also transmit data generated by calculations of the processor 110 via wired or wireless communication with any system, server, client, etc. For example, the network unit 130 may receive observation data of a person who is a target of behavior monitoring via wired or wireless communication with a sensing device such as a camera or a client equipped with a sensing device. The network unit 130 may also receive user input via a user interface implemented in the sensing device or the client equipped with the sensing device. The network unit 130 may transmit various data generated by calculations of the processor 110 based on the observation data via wired or wireless communication with the sensing device or the client equipped with the sensing device.
[0042] FIG. 2 is a block diagram illustrating a process for performing behavioral monitoring on a computing device according to one embodiment of the present disclosure.
[0043] 2, a computing device 100 according to an embodiment of the present disclosure can input observation data 11 about a person who is the subject of behavioral monitoring into a pre-trained deep learning model 200. Here, the observation data 11 can be understood as data that can be acquired by behavioral monitoring a person performing a behavior in a specific environment.
[0044] For example, the observation data 11 may be at least one of images or video captured by a camera installed in a space for the online exam or audio collected by a microphone installed in the space for the online exam. Here, a sensor for the observation data 11, such as a camera or a microphone, may be a component of a client held by the test taker. When the online exam begins, the sensor included in the test taker's client may generate at least one of images, video, or audio of the test taker and the exam space. The computing device 100 may acquire the observation data 11 generated by the test taker's client through wired or wireless communication with the test taker's client. The computing device 100 may then input the acquired observation data 11 into the pre-trained deep learning model 200.
[0045] The computing device 100 may generate an analysis result for at least one sensing item included in each of a plurality of sensing objects through the deep learning model 200 to which the observation data 11 is input. Here, the sensing item may be state information identified based on a subclass of the sensing object. The sensing item may indicate a state that can change depending on a human action.
[0046] For example, if the sensing object is a human body part, the subclass of the sensing object can be divided into a face, an arm, etc. The face can be further divided into eyes, a nose, a mouth, and ears. The arm can be further divided into a hand, a palm, a finger, etc. The sensing item is state information that can be sensed based on each of the subclasses of the sensing object, and can be a gaze direction, whether or not there is speech, a hand position, a palm direction, etc. In other words, the sensing item can indicate a specific state or appearance that can appear when the subclass of the sensing object moves or changes due to a human action.
[0047] In other words, the computing device 100 can input the observation data 11 into the pre-trained deep learning model 200 to generate a plurality of analysis results 13, 15, and 17 for various sensing items. Here, each of the first analysis result 13, second analysis result 15, and third analysis result 17 generated through the deep learning model 200 can be matched with each sensing item such as gaze direction, presence or absence of speech, hand position, palm direction, etc.
[0048] For example, assuming that the behavior monitoring environment is an environment for an online exam, the first analysis result 13 matching the gaze direction may indicate whether the examinee is looking at a display where they are checking a test paper. The second analysis result 15 matching the presence or absence of speech may indicate whether the examinee's mouth shape has changed. The third analysis result 17 matching the hand position may indicate whether the examinee's left or right hand is moving within a reference space determined by the examinee's body and the desk arrangement. In this manner, the computing device 100 may individually generate an analysis result for at least one sensing item included in each of a plurality of sensing targets using the pre-trained deep learning model 200. Through this calculation process, the computing device 100 may acquire various information that can be used to infer a behavior from the observation data 11 and use it in the calculation process for behavior estimation, which will be described later.
[0049] Meanwhile, the deep learning model 200 may be a model based on a neural network capable of processing single data, or a model based on a neural network capable of processing sequential data. For example, the deep learning model 200 may include a convolutional neural network that receives an image corresponding to single data, extracts image features, and recognizes an object. The deep learning model 200 may also include a recurrent neural network that receives sequential data, such as audio, and extracts and interprets features of the sequential data. In addition to the above examples, neural networks capable of processing single data or sequential data may be included in the deep learning model 200 of the present disclosure.
[0050] The deep learning model 200 may be pre-trained using labels, which are pre-verified analysis results for sensing items of sensing objects present in observation data, as ground truth (GT). Specifically, during the training process, the deep learning model 200 may receive observation data and generate analysis results for each sensing item of the sensing objects present in the observation data. The deep learning model 200 may then perform training by repeatedly comparing the generated analysis results with the labels and updating neural network parameters based on the comparison results. Here, the comparison operation may be performed in a direction that minimizes a loss calculated using a loss function such as cross entropy. While the above example illustrates a training process based on supervised learning, the deep learning model 200 may also be trained based on semi-supervised learning, unsupervised learning, self-supervised learning, and the like, in addition to supervised learning.
[0051] Referring to FIG. 2, the computing device 100 may combine multiple analysis results generated through the deep learning model 200 to generate a behavior estimation result 19 of a person being monitored. Here, the computing device 100 may use a rule set predetermined for a specific environment for behavior monitoring. Specifically, the computing device 100 may identify analysis results that match behavior class-specific judgment conditions included in the predetermined rule set from among the analysis results generated through the deep learning model 200. The computing device 100 may estimate the accuracy of the identified analysis results and the correlation between the identified analysis results and the behavior classes included in the predetermined rule set. The computing device 100 may then combine the analysis results based on the estimated accuracy and correlation to estimate a person's behavior that should be detected for behavior monitoring in a specific environment. Through this calculation process, the computing device 100 may precisely and accurately determine and detect a specific behavior based on various information.
[0052] For example, assuming that the environment for behavior monitoring is an environment for an online test, the computing device 100 may identify an analysis result corresponding to a first behavioral class determination condition for cheating and / or a second behavioral class determination condition for abnormal behavior from among a first analysis result 13 matching a gaze direction, a second analysis result 15 matching the presence or absence of speech, and a third analysis result 17 matching a hand position by screening a predetermined rule set. The computing device 100 may estimate the accuracy of the analysis result corresponding to the first behavioral class and / or the second behavioral class from among the first analysis result 13, the second analysis result 15, and the third analysis result 17. Here, the accuracy may be understood as a quantitative indicator indicating how accurately the deep learning model 200 performed the analysis. The computing device 100 may then estimate the correlation between the analysis result corresponding to the first behavioral class and / or the second behavioral class and the first behavioral class and / or the second behavioral class using the predetermined rule set. Here, correlation may be understood as a quantitative indicator indicating the degree to which a specific analysis result influences the determination of a specific behavioral class. The computing device 100 may assign weights to the analysis results corresponding to the first behavioral class and / or the second behavioral class based on the estimated accuracy and correlation. The computing device 100 may combine the weights assigned based on the accuracy and correlation to derive a numerical value for finally determining one of the behavioral classes included in the predetermined rule set as the behavioral inference result 19. Here, the numerical value may be a value that matches the grade of the behavioral class included in the predetermined rule set. That is, the computing device 100 may select one of the behavioral classes included in the predetermined rule set based on the numerical value derived by combining the weights and derive it as the behavioral inference result 19.
[0053] If there is no analysis result that does not match the judgment criteria for the first behavior class or the judgment criteria for the second behavior class among the first analysis result 13, the second analysis result 15, and the third analysis result 19, the computing device 100 may determine that no cheating or abnormal behavior has occurred by the examinee based on the currently input observation data 11. Then, the computing device 100 may input the observation data of the next time point into the deep learning model 200 and perform the above-described analysis and behavior estimation process again.
[0054] FIG. 3 is a block diagram illustrating a process for performing behavioral monitoring of a computing device according to an alternative embodiment of the present disclosure.
[0055] A computing device 100 according to an alternative embodiment of the present disclosure may use a deep learning model 200 including at least one sub-model to generate an analysis result for at least one sensing item included in a plurality of sensing objects based on observation data 21 of a person who is the subject of behavior monitoring. For example, referring to FIG. 3, the deep learning model 200 may include a first model 210 that estimates a person's pose based on an image, a second model 220 that estimates a person's facial shape and direction based on an image, a third model 230 that tracks a person's gaze based on an image, a fourth model 240 that recognizes objects other than a person based on an image, or a fifth model 250 that detects sound elements of an object related to a person's behavior based on at least one of an image or audio. While FIG. 3 illustrates the deep learning model 200 as including all of the first model 210 to the fifth model 250, the present invention is not limited to this. That is, the deep learning model 200 may include at least one of a first model 210, a second model 220, a third model 230, a fourth model 240, or a fifth model 250.
[0056] According to an alternative embodiment of the present disclosure, each of the first model 210 to the fifth model 250 may be matched with at least one sensing item included in each of a plurality of sensing targets. That is, each of the first model 210 to the fifth model 250 may be pre-trained to derive analysis results optimized for sensing items according to a specific environment for activity monitoring. Each of the first model 210 to the fifth model 250 may receive observation data 21 and derive analysis results for the learned sensing items. Here, the first model 210 to the fifth model 250 may be individually matched with two or more different sensing items according to a specific environment for activity monitoring to derive two or more analysis results.
[0057] For example, the first model 210, which receives an image and estimates a person's pose, can be trained to derive analysis results for two different sensing items corresponding to the position of the hand and the orientation of the palm, which are included in the sensing target, i.e., a body part. Thus, the first model 210 can receive observation data 21 and output a first-first analysis result 22 indicating the result of sensing whether the examinee's left or right hand is moving within a reference space determined by the examinee's body and the desk layout. The first model 210 can also receive observation data 21 and output a first-second analysis result 23 indicating the result of determining whether the orientation of the examinee's palm matches the examinee's gaze direction. Although not shown in FIG. 3 , each of the second model 220 through the fifth model 250 can also generate analysis results for two or more different sensing items, similar to the first model 210 described above.
[0058] In this manner, by utilizing the first model 210 to the fifth model 250 individually optimized for one or more sensing items, it is possible to efficiently and quickly process the calculation process for deriving analysis results for various sensing items from the observation data 21. Furthermore, it is possible to realize a system that can perform behavior monitoring in real time through efficient and fast processing by the first model 210 to the fifth model 250.
[0059] Meanwhile, the first model 210 can receive an image of a person and detect the pose of the person in the image. For example, the first model 210 can receive the image, classify body parts and background based on a plurality of feature points for identifying the person's pose, and generate a mask for the body parts. The first model 210 can then analyze the mask for the body parts to estimate what pose the person is in. For this pose estimation, the first model 210 can include a neural network optimized for image processing. The first model 210 can be trained based on supervised learning, semi-supervised learning, unsupervised learning, self-supervised learning, etc.
[0060] The second model 220 receives an image of a person and can detect the shape and orientation of the person's face present in the image. For example, the second model 220 can extract a human face region from an image of the person and generate a cropped image. The second model 220 can generate a feature map based on the cropped image. The second model 220 can then perform an attention operation based on an affine matrix based on the generated feature map and the cropped image to generate 3D landmarks in a mesh form for the human face. The second model 220 can estimate the face shape based on the 3D landmarks and estimate the face direction based on changes in feature points included in the 3D landmarks. To estimate the face shape and orientation, the second model 220 can include a neural network optimized for image processing. The second model 220 can be trained based on supervised learning, semi-supervised learning, unsupervised learning, self-supervised learning, etc.
[0061] The third model 230 can receive an image of a person and track the gaze of the person present in the image. For example, the third model 230 can extract a facial region of the person from the image of the person and generate a cropped image. The third model 230 can extract features based on the cropped image to recognize the person's eyes. The third model 230 can then track the person's gaze by analyzing the movement and changes of the pupils included in the recognized eyes. For such gaze tracking, the third model 230 can include a neural network optimized for image processing. The third model 230 can be trained based on supervised learning, semi-supervised learning, unsupervised learning, self-supervised learning, etc.
[0062] The fourth model 240 receives an image of an object and can detect objects present in the image, excluding people. For example, the fourth model 240 can receive an image and classify the objects present in the image into people and objects. The fourth model 240 can estimate the type and location of the object present in the image by performing semantic segmentation based on the classified object. Here, the semantic segmentation can be performed using a pixel-based method, an edge-based method, or a region-based method, without limitation. To estimate the type and location of the object, the fourth model 240 can include a neural network optimized for image processing. The fourth model 240 can be trained using supervised learning, semi-supervised learning, unsupervised learning, self-supervised learning, etc.
[0063] The fifth model 250 receives an image or audio representing sounds generated in a space where a person who is a target of activity monitoring is present, and can detect sound elements related to the person's activities. For example, the fifth model 250 can extract features of sound elements from an image representing a sound waveform or audio representing a sound waveform signal. The fifth model 250 can estimate the volume, source, and type of sound generated in a space where a person who is a target of activity monitoring is present, based on the features of sound elements extracted from the image or audio. For this sound estimation, the fifth model 250 can include a neural network optimized for processing sequential data. The fifth model 250 can be trained based on supervised learning, semi-supervised learning, unsupervised learning, self-supervised learning, and the like.
[0064] Meanwhile, the calculation process for deriving the behavior estimation result 28 based on the analysis results 22, 23, 24, 25, 26, and 27 output by each of the first model 210 to the fifth model 250 corresponds to the calculation process for deriving the behavior estimation result 19 in Fig. 2 described above, and therefore a detailed description thereof will be omitted below. Also, the specific types of sensing targets, sensing items, and analysis results detailed in Figs. 2 and 3 are merely examples, and the types of sensing targets, sensing items, and analysis results may be configured in various ways within the scope that can be understood by a person skilled in the art based on the contents of the present disclosure.
[0065] 4a is a table summarizing an analysis method and analysis results for each sensing item according to an embodiment of the present disclosure, and FIG. 4b is a table summarizing a rule set for behavior estimation and behavior estimation results according to an embodiment of the present disclosure.
[0066] Referring to Table 30 in Figure 4a, sensing objects according to an embodiment of the present disclosure may be classified into subclasses, which are further subdivided into major, middle, and minor classes. Sensing items may be classified based on the subclasses of the sensing objects. Sensing items correspond to status information measured based on the subclasses of the sensing objects, and may be change information that appears due to human behavior.
[0067] For example, based on a sensing target such as a body part, the subclasses of the sensing target may be divided into a major category including a face, an arm, etc., a middle category including eyes, nose, mouth, ears, hands, palms, fingers, etc., and a minor category that distinguishes the middle category into left and right sides depending on the direction. Sensing items may be classified by matching with the middle and minor categories of the sensing target. Specifically, sensing items may be classified into gaze direction measured based on the eyes, face measured based on the eyes, nose, ears, and mouth, speech measured based on the mouth, hand position measured based on the hand, palm orientation measured based on the palm, hand actions measured based on the fingers, etc.
[0068] According to an embodiment of the present disclosure, multiple deep learning models can be individually matched with sensing items. To efficiently perform analysis according to the analysis purpose for each sensing item, each of the multiple deep learning models can be classified by sensing item. Here, two or more deep learning models can be used to analyze one sensing item through classification, and one deep learning model can be used to analyze two or more sensing items.
[0069] For example, a detection item such as gaze direction can be matched with a second model that detects the shape and direction of a person's face and a third model that tracks a person's gaze. The second model that matches gaze direction can detect the shape and direction of a person's face and detect whether the person's gaze direction deviates from a screen that outputs specific information. The third model that matches gaze direction can detect a person's gaze and detect whether the person's gaze direction deviates from the screen. As shown in Table 30 of FIG. 4a, the outputs of the second model and the third model can each be used as individual analysis results for the gaze direction. Although not shown in Table 30 of FIG. 4a, the outputs of the second model and the third model can be combined into a single analysis result based on priority or accuracy, and used to predict a person's behavior.
[0070] The third model, which matches the sensing item of gaze direction, can also match other sensing items such as face recognition and speech. The third model, which tracks a person's gaze, can derive analysis results for gaze direction, face recognition, and speech based on input data. That is, when observation data of a person who is the target of behavior monitoring is input, multiple deep learning models according to an embodiment of the present disclosure can analyze all of the sensing items that match based on the input data. By matching models for each sensing item in this way, the computing device 100 according to an embodiment of the present disclosure can simultaneously analyze all conditions defined to determine the behavior to be monitored from a single observation data in real time.
[0071] Referring to Table 40 of Fig. 4b, a rule set according to one embodiment of the present disclosure may include behavior classes that are candidates for monitoring and judgment conditions for the behavior classes, where the behavior classes may correspond to the overall judgments and judgment classifications shown in Table 40 of Fig. 4b, and the judgment conditions may correspond to the sensing items shown in Table 40 of Fig. 4b.
[0072] For example, in an online testing environment, behavior classes included in a rule set may be classified into cheating and abnormal behavior. The rule set may define specific behaviors or situations that each of the cheating and abnormal behaviors included in the rule set represents. Furthermore, a judgment condition indicating a combination of behaviors that determines each of the cheating and abnormal behaviors included in the rule set may be defined by matching each class with the rule set. Specifically, cheating corresponding to a situation in which a mobile phone is detected for 5 seconds or more may be detected by combining a judgment condition for detecting a mobile phone and a judgment condition that the mobile phone is exposed for 5 seconds or more. Thus, the rule set may define a behavior class in which cheating corresponding to a situation in which a mobile phone is detected for 5 seconds or more, and define a judgment condition for detecting a mobile phone and a judgment condition that the mobile phone is exposed for 5 seconds or more, matching the class. Here, each judgment condition may be assigned an identification code corresponding to the detection code shown in Table 40 of FIG. 4b. The identification code assigned to the judgment condition can be used to identify which detection object, detection item, detection device, and detection result each judgment condition is derived from.
[0073] The grades displayed in table 40 of Fig. 4b may correspond to values calculated based on weights according to the detection accuracy for a combination of judgment criteria. The relevance levels displayed in table 40 of Fig. 4b may correspond to values calculated based on weights according to the correlation for a combination of judgment criteria. And, the overall grades displayed in table 40 of Fig. 4b may correspond to values calculated based on the grades and relevance levels of the analysis results of the deep learning model matching the judgment criteria.
[0074] For example, assuming that four analysis results corresponding to the bold boxes shown in Table 30 of FIG. 4a are derived, the judgment criteria matching the four analysis results in the rule set can be identified as the bold boxes shown in Table 40 of FIG. 4b. Each of the four analysis results matching the judgment criteria can be assigned a first weight based on the output accuracy of each model and classified into grades such as A1, A2, B1, and C1. Furthermore, each of the four analysis results matching the judgment criteria can be assigned a second weight based on the matching behavioral class and correlation and classified into relevance levels such as very high, high, average, low, and very low. Once each of the four analysis results matching the judgment criteria is classified by grade and relevance, an overall grade for final judgment can be calculated by combining the four values using the following formula: 1, and the final behavioral class can be determined based on the calculated overall grade.
[0075]
number
[0076] Referring to Table 30 in Figure 4a, it appears that all eight judgment conditions must match the analysis results to completely determine the fraudulent behavior of using a mobile phone with the left hand for eight seconds. However, as in the above-described example of the present disclosure, when a final judgment is made based on the grade and relevance, even if only four judgment conditions, rather than all eight, are confirmed as the analysis results, the fraudulent behavior of using a mobile phone with the left hand for eight seconds can be detected with a high probability. In other words, as in the present disclosure, when a calculation is performed to combine the underlying actions based on the detection accuracy of the deep learning model and the correlation between the analysis results and a specific behavior class, a specific behavior can be accurately inferred based only on the analysis results derived through the deep learning model, even if not all judgment conditions defined in the rule set are satisfied.
[0077] In addition to the above examples, various fraudulent or abnormal behaviors may be defined by an administrator who creates an online testing environment and generated as a rule set. Furthermore, the computing device 100 according to an embodiment of the present disclosure may be applied to various monitoring environments other than the online testing environment.
[0078] 5a to 5c are conceptual diagrams illustrating a detailed process of inferring behaviors of a computing device according to an embodiment of the present disclosure.
[0079] Comparing Figures 5a and 5b, it can be seen that even though the same two sensing objects are analyzed, subtle differences in the analysis results lead to different results. Specifically, as in Figure 5a, if the deep learning model outputs an analysis result of 5 seconds based on the sensing item "time," the person's behavior can be inferred as fraudulent behavior, i.e., the mobile phone was detected for 5 seconds or more. On the other hand, as in Figure 5b, if the deep learning model outputs an analysis result of 1 second based on the sensing item "time," the person's behavior can be inferred as anomalous behavior, i.e., the mobile phone was detected for 1 second or more. Here, anomalous behavior may indicate behavior that is suspicious of fraudulent behavior, but does not constitute fraud. A computing device 100 according to an embodiment of the present disclosure can precisely interpret the above-mentioned differences by deriving analysis results for each sensing item through a deep learning model and combining the derived analysis results based on accuracy and correlation, thereby accurately distinguishing between fraudulent behavior and anomalous behavior.
[0080] Referring to FIG. 5c, it can be seen that the computing device 100 according to an embodiment of the present disclosure detects the fraudulent behavior of using a mobile phone with the left hand for 8 seconds by comprehensively considering the results of analyzing the sensing results of various sensing items included in various sensing objects. The computing device 100 can derive analysis results for different sensing items, such as the left hand action, the left hand position, the position of the left arm, and the horizontal / horizontal angle, for a single sensing object, such as a body part, through a deep learning model that individually matches the analysis results. Furthermore, the computing device 100 can accurately determine fraudulent behavior by combining the analysis results for each sensing item derived using the deep learning model based on a ranking based on sensing accuracy and a correlation based on correlation. In this way, by deriving the analysis results for each sensing item through a deep learning model that matches the sensing items and combining all the individual results to make a final judgment on the behavior, a highly reliable behavior inference result can be obtained.
[0081] FIG. 6 is a flowchart illustrating an artificial intelligence-based activity monitoring method according to one embodiment of the present disclosure.
[0082] Referring to FIG. 6, a computing device 100 according to an embodiment of the present disclosure may generate an analysis result for a sensing item based on observation data of a person who is a target of activity monitoring by using a deep learning model that matches at least one sensing item included in each of a plurality of sensing objects (S110). Specifically, the computing device 100 may acquire observation data at a predetermined period. Here, the predetermined period may be determined according to environmental conditions set through a client that manages activity monitoring. When a manager of a specific environment for activity monitoring sets environmental conditions through the client, the client may transmit observation data to the computing device 100 at a period determined according to the set conditions. The observation data may be at least one of images or videos captured of a space built around a person according to environmental conditions, or audio captured in the space. The computing device 100 may receive observation data transmitted from the client at a predetermined period. The computing device 100 can then input the acquired observation data into at least one of a first model for pose estimation, a second model for facial shape and orientation estimation, a third model for gaze tracking, a fourth model for object recognition, and a fifth model for sensing sound elements, and generate analysis results for the sensing items that reflect the results of the person's actions performed during a predetermined period.
[0083] The computing device 100 may use the rule set to infer a person's behavior based on the analysis results generated in step S110 (S120). Here, the rule set may be predetermined according to environmental conditions set through a client managing behavior monitoring. Specifically, the computing device 100 may identify analysis results that match judgment conditions for each behavior class included in the predetermined rule set from among the analysis results in step S110. The computing device 100 may estimate the accuracy of the identified analysis results and a correlation between the identified analysis results and the behavior classes included in the rule set. The computing device 100 may then infer a person's behavior by combining the identified analysis results based on the estimated accuracy and correlation. For example, the computing device 100 may assign a first weight to each identified analysis result based on the estimated accuracy, assign a second weight to each identified analysis result based on the estimated correlation, and determine whether the person has performed at least one of the behavior classes included in the predetermined rule set based on a value derived by combining the first weight and the second weight. Here, the behavior classes included in the predetermined rule set may include a first behavior class set via a client in charge of behavior monitoring, which corresponds to test cheating, and a second behavior class set via the client, which corresponds to behavior that is not cheating but is suspected of cheating or abnormal behavior that is unnecessary for taking a test. For example, checking a cell phone for about one second to check the time is not cheating but is suspected of cheating, so it may be pre-set to the second behavior class rather than the first behavior class. The types of behavior classes described above are merely examples, and various types of behavior classes may be configured within a scope that would be understandable to one skilled in the art based on the contents of this disclosure.
[0084] FIG. 7 is a flow chart illustrating a method for monitoring behavior in an online testing environment according to one embodiment of the present disclosure.
[0085] 7, a computing device 100 according to an embodiment of the present disclosure may generate an online exam based on a user request input via a promoter client for the online exam (S210). Here, environmental conditions for the online exam, a rule set for monitoring the behavior of examinees, etc. may be determined based on the user request input via the promoter client. For example, the computing device 100 may determine a rule set including a period for acquiring observation data, definitions and judgment conditions for cheating 61 or abnormal behavior 62, etc., based on the user request input via the promoter client. After the rule set is generated based on the user request, it may be dynamically updated as the computing device 100 repeatedly performs behavior estimation.
[0086] Once the online exam is generated (S210), the computing device 100 may acquire observation data at a predetermined interval (S220). For example, the computing device 100 may acquire observation data at intervals of 100 ms to 1 s via wired or wireless communication with a sensing device installed in the exam space. Here, the sensing device may be a component installed in the examinee's client or may be a component of the computing device 100. The observation data acquisition interval may be predetermined in step S210 depending on the environmental conditions of the online exam.
[0087] The computing device 100 may perform an analysis for each sensing item to infer misconduct 61 or abnormal behavior 62 included in a rule set based on observation data acquired at a predetermined period (S230). Here, the computing device 100 may use a plurality of deep learning models 210, 220, 230, 240, and 250, each matching at least one sensing item. The plurality of deep learning models 210, 220, 230, 240, and 250 may generate an analysis result for each sensing item based on a sensing object present in the observation data by matching at least one sensing item. Here, the analysis result for each sensing item is status information measured based on the sensing item and may be information that can change depending on a person's behavior. For example, the first model 210 may receive observation data and analyze whether the left hand is adjacent to the desk based on the sensing item, the position of the left hand. The first model 210 may also perform an analysis for other sensing items included in the sensing object, the body part, in addition to the position of the left hand. The second model 220 receives the observation data and can analyze whether the shape of the examinee's mouth changes based on a detection item called "utterance." The third model 230 receives the observation data and can analyze whether the examinee's gaze deviates from the display area where test questions are displayed based on a detection item called "gaze direction." The fourth model 240 receives the observation data and can detect whether a mobile phone is present within a predetermined radius around the examinee based on a detection item called "mobile phone" among objects. If a mobile phone is present, the fourth model 240 can measure the amount of time the mobile phone is exposed within the predetermined radius based on a detection item called "time." The fifth model 250 receives the observation data and can analyze the entity that generated the sound within the examination space based on a detection item called "sound generating entity." The above examples are intended to facilitate understanding of the contents of the present disclosure, and the analysis results by model of the present disclosure are not limited to the above examples.
[0088] The computing device 100 may identify an analysis result that matches a determination condition included in a predetermined rule set from among the analysis results for each sensing item derived in operation S230. The computing device 100 may compare the determination conditions for each of the fraudulent behavior 61 and abnormal behavior 62 defined in the predetermined rule set with the analysis results for each sensing item derived in operation S230 to identify matching analysis results. Here, if no matching results exist, the computing device 100 may perform the process again from operation S220.
[0089] The computing device 100 may estimate accuracy and correlation for the analysis results identified in step S240 (S250). The computing device 100 may estimate accuracy and correlation to determine the weight to which analysis results matching the judgment conditions included in the rule set should be weighed in final judgment of fraudulent activity 61 or anomalous behavior 62. Here, accuracy may be estimated based on the detection accuracy of each of the multiple deep learning models 210, 220, 230, 240, and 250. Correlation may be estimated based on how much the analysis results identified in step S240 affect the judgment of a specific fraudulent activity or a specific anomalous behavior.
[0090] The computing device 100 may assign weights to the analysis results identified in step S240 according to the accuracy and correlation estimated in step S250. The computing device 100 may then generate a basis for determining fraudulent behavior 61 or abnormal behavior 62 by combining the weighted analysis results. For example, the computing device 100 may assign a higher weight to the analysis results identified in step S240 according to the accuracy estimated in step S250, and may assign a higher weight to the analysis results identified in step S240 according to the correlation. The computing device 100 may derive a value for determining fraudulent behavior 61 or abnormal behavior 62 by performing a mathematical operation that combines a weight assigned according to the accuracy and a weight assigned according to the correlation for all analysis results identified in step S240.
[0091] The computing device 100 can infer any one of the types of behavior included in the rule set based on the numerical value derived by the combination of step S260. The computing device 100 can infer that, among the types of behavior included in the rule set, a specific fraudulent behavior or a specific abnormal behavior corresponding to the numerical value derived by the combination of step S260 is the behavior of the person being observed.
[0092] The various embodiments of the present disclosure described above can be combined with additional embodiments and can be modified within the scope that can be understood by those skilled in the art based on the above detailed description. It should be understood that the embodiments of the present disclosure are illustrative in all respects and are not limiting. For example, each component described as a single type can also be implemented in a distributed form, and similarly, components described as distributed can be implemented in a combined form. Therefore, all modifications and variations derived from the meaning, scope, and equivalent concepts of the claims of the present disclosure should be construed as being within the scope of the present disclosure.
Claims
1. 1. An artificial intelligence based behavioral monitoring method executed by a computing device including at least one processor, comprising: generating an analysis result for the sensing item based on observation data of the person being the subject of behavior monitoring by using a deep learning model that matches at least one sensing item included in each of the plurality of sensing objects; Inferring the person's behavior according to the generated analysis results using a predetermined rule set; A method comprising:
2. The sensing item is status information identified based on a subclass of the sensing object, The method of claim 1 , wherein the state information is changeable by an action of the person.
3. The method of claim 1, wherein the plurality of sensing objects include, in addition to the body parts of the person, at least one of an object other than the person, a sound of an object related to the person's actions, or a time of an object related to the person's actions.
4. The deep learning model is a first model for estimating a person's pose based on an image; a second model that estimates the shape and orientation of a person's face based on the image; The third model tracks people's gaze based on images. A fourth model that recognizes objects other than people based on images, or a fifth model for detecting sound elements of an object related to human activity based on at least one of an image or an audio; The method of claim 1 , comprising at least one of:
5. generating an analysis result for the sensing item based on observation data of a person who is a target of behavior monitoring by using a deep learning model that matches at least one sensing item included in each of the plurality of sensing objects; acquiring the observation data at a predetermined interval; inputting the acquired observation data into at least one of the first model, the second model, the third model, the fourth model, or the fifth model, and generating an analysis result for the sensing item that reflects the result of the person's behavior performed during the predetermined period; The method of claim 4, comprising:
6. The method according to claim 5 , wherein the predetermined period is determined in accordance with an environmental condition set via a client in charge of behavior monitoring.
7. Inferring the person's behavior according to the generated analysis results using a predetermined rule set includes: identifying an analysis result that matches a judgment condition for each behavioral class included in the predetermined rule set from among the generated analysis results; estimating the accuracy of the identified analysis results and the correlation of the identified analysis results with the behavioral classes included in the predetermined rule set; combining the identified analysis results based on the estimated accuracy and correlation to infer the person's behavior; The method of claim 1 , comprising:
8. Inferring the human behavior by combining the identified analysis results based on the estimated accuracy and correlation, assigning a first weight to each of the identified analysis results according to the estimated accuracy; assigning a second weight to each of the identified analysis results according to the estimated correlation; determining whether the person has performed at least one of the behavior classes included in the predetermined rule set based on a value derived by combining the first weighted value and the second weighted value; The method of claim 7, comprising:
9. The behavior classes included in the predetermined rule set are: a first behavior class configured via a client administering behavior monitoring to address exam cheating; a second behavior class configured via the client corresponding to abnormal behavior that is not necessary for taking the test; The method of claim 8, comprising:
10. A computer program stored on a computer-readable storage medium, the computer program configured, when executed by one or more processors, to perform operations for monitoring behavior based on artificial intelligence; The operation is an operation of generating an analysis result for the sensing item based on observation data of a person who is a target of behavior monitoring by using a deep learning model that matches at least one sensing item included in each of a plurality of sensing targets; Inferring the person's behavior according to the generated analysis results using a predetermined ruleset; a computer program comprising:
11. 1. A computing device for monitoring behavior based on artificial intelligence, comprising: a processor including at least one core; a memory containing program code executable by the processor; a network unit for acquiring observation data of a person who is a target of behavior monitoring; Including, The processor: Using a deep learning model that matches at least one sensing item included in each of a plurality of sensing targets, and generating an analysis result for the sensing item based on the observation data; A device that uses a predetermined rule set to infer the person's behavior according to the generated analysis results.
Citation Information
Patent Citations
Crime prevention
JP1985262299A
Vehicle security remote monitoring system
JP2009161093A
Information processing apparatus and method, computer program, and monitoring system
JP2019176423A
Monitoring camera and detection method
JP2021043666A
Analysis device, analysis method, and analysis program
JP2021101318A