Abnormal behavior detection system

The abnormal behavior detection system uses a comprehensive visual language model to analyze object size and joint displacements, ensuring accurate detection of small objects and managing system load, addressing the challenge of recognizing fine details in natural language and vision models.

JP7894613B1Active Publication Date: 2026-07-24ASILLA INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ASILLA INC
Filing Date
2025-12-25
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Natural language and vision language models struggle to detect fine details of human actions, particularly when humans appear small in videos, leading to inaccurate recognition of abnormal behaviors.

Method used

An abnormal behavior detection system utilizing a comprehensive visual language model to determine object size and analyze joint or key point displacements, with a first abnormal behavior determination unit for high accuracy in detecting small objects and a control mechanism to distribute load between models.

Benefits of technology

Efficiently determines abnormal behaviors based on object size, reducing overall system load while maintaining high accuracy for small objects, and allowing for prioritization of high-urgency behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894613000001_ABST
    Figure 0007894613000001_ABST
Patent Text Reader

Abstract

This invention provides an abnormal behavior detection system, method, and program that efficiently determines abnormal behavior of an object based on its size as captured in video. [Solution] The abnormal behavior determination system 1 comprises a comprehensive visual language model 3 that determines whether or not an abnormal behavior has occurred by a target behavioral object Z shown in a target video Z based on feature quantities in the target video Y, and a first abnormal behavior determination unit 5 that determines whether or not an abnormal behavior has occurred by a target behavioral object Z in the target video Y based on the displacement of joints or key points. The first abnormal behavior determination unit 5 is capable of detecting joints or key points with a higher accuracy than the comprehensive visual language model 3, and if the target behavioral object on the target video is of a size greater than or equal to the size that can be determined, the comprehensive visual language model 3 determines that an abnormal behavior has occurred, and if it is smaller than the size that can be determined, the first abnormal behavior determination unit 5 determines that an abnormal behavior has occurred.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an abnormal behavior determination system capable of efficiently determining abnormal behavior of a moving object according to the size of the moving object shown in an image.

Background Art

[0002] Conventionally, in order to facilitate understanding of the basis of image recognition results, a learned image recognition model is applied to a processing target image to generate a feature map including at least one attention area, and an image recognition result output unit that outputs the feature map and the image recognition result of the processing target image, a map generation unit that generates an importance map from the feature map based on the importance of each of a plurality of pixels constituting the processing target image, a template acquisition unit that acquires a prompt template including an instruction for a natural language model, and an image recognition result to be output, coordinate values and pixel values of a predetermined number of pixels in order from pixels with high importance in the importance map, and the number of pixels constituting the processing target image are fitted to the prompt template to generate a prompt, and a prompt output unit that outputs the prompt to a natural language model. An information processing apparatus including these is known (for example, see Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, in a natural language model or a vision language model, for example, when recognizing that a human has performed a predetermined action, a set of feature amounts in the case where the human has performed the predetermined action is learned in advance, and when a set of feature amounts close to the learned set of feature amounts is detected, it is recognized that the human has performed the predetermined action.

[0005] However, if the performance of natural language models or the resolution of the video is low, it may not be possible to detect the fine details of the feature set. If these details are important elements for recognizing a particular action, this can lead to a problem where the model fails to recognize that a human performed that action. This problem is particularly pronounced when humans appear small in the video.

[0006] Therefore, the present invention aims to provide an abnormal behavior detection system that can efficiently determine abnormal behavior of an object based on its size as captured in a video. [Means for solving the problem]

[0007] The present invention comprises: an acquisition unit that acquires target video; a comprehensive visual language model that has learned the displacement of feature points of an object when an object in the video performs abnormal behavior, and is capable of determining whether or not an abnormal behavior occurred by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of detecting the size of the object in the target video; a storage unit that stores the size of the object in the video that the comprehensive visual language model can determine whether or not an abnormal behavior occurred based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when an object in the video performs abnormal behavior; a detection unit that can detect the joints or key points based on the feature quantities of the object in the target video; and the detected displacement of the joints or key points. An abnormal behavior determination system is provided, comprising: a first abnormal behavior determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video based on the displacement of a joint or key point, wherein the first abnormal behavior determination unit is capable of detecting the same joint or key point of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model, and when the target behavioral body in the target video is larger than the size that can be determined, the integrated visual language model determines whether or not the abnormal behavior was performed by the target behavioral body or analyzes the abnormal behavior, and when the target behavioral body in the target video is smaller than the size that can be determined, the integrated visual language model controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior was performed by the target behavioral body or analyzes the abnormal behavior.

[0008] With this configuration, under normal circumstances, the overall load (GPU or CPU usage) of the abnormal behavior detection system is kept down by using a comprehensive visual language model to easily determine abnormal behavior. Furthermore, when the target behavioral entity Z is smaller than the detectable size, the first abnormal behavior detection unit can perform abnormal behavior detection with high accuracy. Thus, it becomes possible to efficiently determine abnormal behavior of the target behavioral entity according to its size as seen in the target video. [Effects of the Invention]

[0009] According to the abnormal behavior detection system of the present invention, it is possible to efficiently determine whether an object is exhibiting abnormal behavior based on the size of the object captured in the video. [Brief explanation of the drawing]

[0010] [Figure 1] Diagram illustrating the abnormal behavior of a target organism according to an embodiment of the present invention. [Figure 2] Block diagram of an abnormal behavior detection system according to an embodiment of the present invention. [Figure 3] Diagram illustrating the neural network of a visual language model according to an embodiment of the present invention. [Figure 4] Flowchart for determining abnormal behavior according to an embodiment of the present invention [Figure 5] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 6] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 7] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 8] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 9] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 10] Block diagram of an abnormal behavior detection system according to a modified version of the present invention. [Figure 11] Block diagram of an abnormal behavior detection system according to a modified version of the present invention.

Embodiments for Carrying out the Invention

[0011] Hereinafter, the abnormal behavior determination system 1 according to the embodiment of the present invention will be described with reference to FIGS. 1 to 4.

[0012] As shown in FIG. 1, the abnormal behavior determination system 1 is for checking whether an abnormal behavior has been performed by the target moving body Z in the target video Y (in FIG. 1, the frames constituting the video) captured by the imaging means X or analyzing the abnormal behavior. In the present embodiment, a human will be described as an example of the moving body, and in FIG. 1, for easy understanding, the target moving body Z is simply displayed only by its skeleton.

[0013] As shown in FIG. 2, the abnormal behavior determination system 1 includes an acquisition unit 2, a general visual language model 3, a storage unit 4, and a first abnormal behavior determination unit 5.

[0014] The acquisition unit 2 acquires the target video Y. In the present embodiment, the acquisition unit 2 acquires the target video Y by sequentially acquiring a plurality of consecutive frames captured by the imaging means X. Also, in the present embodiment, the imaging means X performs continuous shooting, and the acquisition unit 2 always acquires the video captured by the imaging means X.

[0015] The general visual language model 3 has learned the displacement of the feature amount of the moving body when the moving body shown in the video performs an abnormal behavior, and can determine whether an abnormal behavior has been performed by the target moving body Z shown in the target video Y based on the feature amount in the target video Y or analyze the abnormal behavior, and can detect the size of the target moving body Z on the target video Y in the target video Y. In the present embodiment, the determination is started when the acquisition unit 2 acquires the target video Y.

[0016] ] / > The general visual language model 3 is an AI model that can simultaneously understand and process visual information such as images and videos and text information, and is configured by a deep neural network such as a CNN, for example.

[0017] As shown in FIG. 3, the integrated visual language model 3 includes an input layer 31, an intermediate layer 32, and an output layer 33. Although shown simply in FIG. 3, the input layer 31 is provided in a number corresponding to each of the plurality of pixels of the frames constituting the video, and the intermediate layer 32 is preferably composed of a large number of layers. Note that the number of layers of the intermediate layer 32 greatly affects the performance of the integrated visual language model 3, and as the number of layers of the intermediate layer 32 increases, the usage rate of the GPU and CPU also increases.

[0018] In the integrated visual language model 3, when the video of the actor in the case where the actor performs an abnormal action is input to the input layer 31, the displacement of the feature amount of the actor in the case where the actor performs an abnormal action is learned in the intermediate layer 32 so that a determination result indicating that "an abnormal action has been performed" is output from the output layer 33. Therefore, in the present embodiment, when the displacement of the feature amount of the video input to the input layer 31 matches the learned displacement of the feature amount by a predetermined amount or more, it is determined that "the abnormal action has been performed" (actually, it is output from the output layer 33 as "probability of abnormal action: ○○%"). In the example of FIG. 1, based on the displacement of the feature points in the continuous target video Y, it is determined as "falling down".

[0019] In the present embodiment, it is assumed that the integrated visual language model 3 has learned the displacements of the feature amounts of a plurality of actors corresponding to each of a plurality of types of abnormal actions.

[0020] Examples of abnormal actions include actions such as falling down, hitting, and kicking. Further, the abnormal actions may include "precursor actions" such as preparatory actions for shoplifting. Furthermore, in order to estimate actions other than abnormal actions (such as walking and stopping), the integrated visual language model 3 may learn the displacements of the feature amounts when those actions occur. Also, the integrated visual language model 3 may be made to learn normal actions and regard actions other than normal actions as abnormal actions.

[0021] Naturally, the integrated visual-language model 3 is also able to recognize the existence of any behavioral entities in which it has determined whether or not abnormal behavior has occurred. In particular, in this embodiment, the integrated visual-language model 3 is capable of detecting the size of behavioral entities in the video.

[0022] The memory unit 4 stores the size of an object in the target video Y that the integrated visual language model 3 can determine whether or not abnormal behavior has occurred based on the features within the target video Y. The size of the object could, for example, be the size that the integrated visual language model 3 was able to accurately determine with a predetermined probability or higher through experiments (e.g., the number of pixels containing the object). Alternatively, a predetermined size could be stored depending on the performance of the integrated visual language model 3 and the resolution of the target video Y.

[0023] Furthermore, the memory unit 4 stores the characteristic quantities of the joints (neck, right elbow, left elbow, waist, right knee, left knee, etc.) or key points (between the eyes, etc.) of the moving body shown in the video, as well as the displacement of the joints or key points of the moving body when the moving body performs abnormal behavior.

[0024] In this embodiment, as will be described later, in order to detect "characteristic points of the behavioring body" and "displacement of characteristic points of the behavioring body when abnormal behavior occurs," "joint identification criteria," "behavioral identification criteria," and "behavioral identification criteria" are stored.

[0025] The "joint identification criteria" are used to identify multiple joints and key points in a moving body, and for each joint and key point, they specify characteristic quantities such as shape, direction, and size to identify it.

[0026] The "behavioral body identification criteria" indicate the "basic posture," "range of motion of each joint," and "distance between each joint (key point)" for various variations of a behavioral body ("walking," "standing," etc.).

[0027] The "behavioral identification criteria" indicate the movement of each joint and key point when an organism exhibits abnormal behavior.

[0028] Examples of abnormal behaviors include falling, punching, and kicking, and it is possible to store not just one but multiple behaviors. In this embodiment, the memory unit 4 stores the displacement of the characteristic points of the corresponding behavioral entities for each of the multiple types of abnormal behaviors.

[0029] Furthermore, abnormal behavior may include "premonitory behaviors" such as preparatory actions for shoplifting. The memory unit 4 may also store the displacement of characteristic point information when behaviors other than abnormal behavior (such as walking or stopping) occur, in order to estimate such behaviors. Alternatively, the memory unit 4 may store normal behaviors, and any behavior other than normal behaviors may be considered abnormal behavior.

[0030] The first abnormal behavior determination unit 5 comprises a detection unit 51 and a determination unit 52.

[0031] The detection unit 51 can detect joints or key points based on the characteristic quantities of the target moving object Z captured in the target video Y.

[0032] In this embodiment, when detecting the characteristic points of the target behavioral entity Z, first, the target behavioral entity Z captured in the target video Y is identified.

[0033] In detail, the system detects feature quantities corresponding to the "joint identification criteria" stored in memory unit 4, and then refers to the "behavioral body identification criteria" to identify multiple joints and key points included in a single target behavioral body Z. In the example in Figure 1, all joints and key points are identified as being included in a single target behavioral body Z, and thus the existence of a single target behavioral body Z is confirmed.

[0034] The determination unit 52 can determine whether or not abnormal behavior occurred by the target behavioral body Z within the target video Y, or analyze the abnormal behavior, based on the detected joint or key point displacement and the stored joint or key point displacement.

[0035] In this embodiment, if the detected displacement of a joint or key point matches the displacement of a joint or key point in the “behavior identification criteria” by a predetermined amount or more, it is determined that “the abnormal behavior has occurred.”

[0036] In the example in Figure 1, the "fall" is determined based on the displacement of feature points in the continuous target image Y.

[0037] Furthermore, the first abnormal behavior determination unit 5 (detection unit 51) can detect the joints or key points of the same target behavioral body Z within the same target video Y with a higher accuracy than the integrated visual language model 3. On the other hand, in this embodiment, when the first abnormal behavior determination unit 5 determines whether the same abnormal behavior has occurred within the same target video Y, the GPU or CPU usage will be higher than that of the integrated visual language model 3. Whether or not the accuracy rate is high can be determined, for example, by comparing the detection results for the same predetermined number of targets, and if the overall accuracy rate is high, then it can be determined that the accuracy rate is high.

[0038] In natural language models and visual language models, for example, when recognizing that a human has performed a specific action, a set of features representing the characteristics of that action is learned in advance. When a set of features similar to the learned set of features is detected, the model recognizes that the human has performed the specific action.

[0039] However, if the performance of natural language models or the resolution of the video is low, it may not be possible to detect the fine details of the feature set. If these details are important elements for recognizing a particular action, this can lead to a problem where the model fails to recognize that a human performed that action. This problem is particularly pronounced when humans appear small in the video.

[0040] Therefore, in this embodiment, the first abnormal behavior determination unit 5 utilizes a configuration in which it individually identifies joints or key points. If the target behavioral object Z on the target image Y is larger than or equal to the size that can be determined, the integrated visual language model 3 determines whether or not an abnormal behavior has occurred by the target behavioral object Z or analyzes the abnormal behavior. If the target behavioral object Z on the target image Y is smaller than the size that can be determined, the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not an abnormal behavior has occurred by the target behavioral object Z or analyzes the abnormal behavior. In other words, in the latter case, the integrated visual language model 3 also functions as a control unit for the first abnormal behavior determination unit 5.

[0041] For example, the integrated visual language model 3 could be given a prompt in advance stating, "If the target behavioral object Z shown in the target video Y is smaller than the size that can be determined, control the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior occurred in the target video Y or to analyze the abnormal behavior." The integrated visual language model 3 would then, if the target behavioral object Z shown in the target video Y is smaller than the size that can be determined, input the target video Y from the output unit (not shown) of the integrated visual language model 3 to the input unit (not shown) of the first abnormal behavior determination unit 5. The first abnormal behavior determination unit 5 would then be configured to automatically perform abnormal behavior determination, etc., once the target video Y is input, thereby initiating the abnormal behavior determination, etc.

[0042] As a result, under normal circumstances, the overall load on the abnormal behavior detection system 1 is reduced (GPU or CPU usage is reduced) by using the integrated visual language model 3 to easily determine whether or not abnormal behavior has occurred. However, if the target behavioral entity Z is smaller than the size that can be determined, the first abnormal behavior detection unit 5 can determine with high accuracy whether or not abnormal behavior has occurred.

[0043] Next, the flow of abnormal behavior detection according to this embodiment will be explained using the flowchart in Figure 4.

[0044] First, it is determined whether the target object Z, as seen in the target video Y, is smaller than the size that can be judged (S1).

[0045] If the size is larger than the size that can be judged (S1: NO), the comprehensive visual language model 3 determines whether or not abnormal behavior occurred based on the features in the target video Y (S2).

[0046] On the other hand, if the size is smaller than the size that can be judged (S1: YES), the first abnormal behavior determination unit 5 determines whether or not an abnormal behavior occurred in the target video Y or analyzes the abnormal behavior (S3).

[0047] As described above, in the abnormal behavior determination system 1 according to this embodiment, the first abnormal behavior determination unit 5 utilizes a configuration that individually identifies joints or key points, and the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not an abnormal behavior occurred in the target image Y or to analyze the abnormal behavior if the target behavior object Z on the target image Y is smaller than the size that can be determined.

[0048] With this configuration, under normal circumstances, the overall load (GPU or CPU usage) of the abnormal behavior detection system 1 is kept low by using the integrated visual language model 3 to easily determine abnormal behavior. However, when the target behavioral object Z is smaller than the detectable size, the first abnormal behavior detection unit 5 can perform abnormal behavior detection with high accuracy. Therefore, it is possible to efficiently determine abnormal behavior of the target behavioral object Z according to its size as seen in the target video Y.

[0049] Furthermore, the abnormal behavior detection system of the present invention is not limited to the embodiments described above, and various modifications and improvements are possible within the scope described in the claims.

[0050] For example, the integrated visual language model 3 may have a configuration that includes an integrated function that further controls various devices other than the first abnormal behavior determination unit 5 in accordance with the circumstances of the abnormal behavior.

[0051] For example, if the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, a configuration can be considered in which the detailed visual language model 6, which has an even higher accuracy rate than the first abnormal behavior determination unit 5, performs a re-determination or detailed analysis, as shown in Figure 5.

[0052] More specifically, the detailed visual language model 6 has learned the displacement of feature points of an object when it performs an abnormal action in a video, and has a higher accuracy rate than the first abnormal action determination unit 5 when determining whether the same abnormal action occurred within the same target video Y or when analyzing the abnormal action.

[0053] The detailed visual language model 6, like the integrated visual language model 3, is an AI model (large-scale language model) that can simultaneously understand and process visual information such as images and videos, as well as text information. For example, it is composed of a deep neural network such as a CNN. However, the total number of intermediate layers is greater than that of the integrated visual language model 3. This allows for more accurate judgments than the integrated visual language model 3 and the first abnormal behavior judgment unit 5, but it also results in higher GPU or CPU usage.

[0054] The GPU or CPU usage will vary depending on the combination of factors such as memory capacity, number of cores, clock speed, number of threads, and training amount (or training parameters), as well as the actual processing content. For example, under the condition that the training parameters of the detailed visual language model 6 are greater than those of the integrated visual language model 3, and other factors such as memory capacity are the same, when determining whether the same abnormal behavior occurred within the same target video or analyzing abnormal behavior, the detailed visual language model 6 will have higher GPU or CPU usage than the integrated visual language model 3.

[0055] Furthermore, memory unit 4 also stores highly urgent abnormal behaviors. Highly urgent abnormal behaviors include things like "violence" and "failure to get up within a specified time after falling," and these can be set according to the user's requests.

[0056] Then, if the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, the integrated visual language model 3 controls the detailed visual language model 6 to determine whether or not the abnormal behavior occurred in the target video Y or to analyze the abnormal behavior. In other words, in this case, the integrated visual language model 3 also functions as the control unit for the detailed visual language model 6.

[0057] For example, one possible configuration is to pre-provide the integrated visual language model 3 with the prompt, "If the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, control the detailed visual language model 6 to determine whether or not the abnormal behavior occurred in the target video Y or to analyze the abnormal behavior," and then, if the abnormal behavior is of high urgency, the integrated visual language model 3 will prompt the detailed visual language model 6 with the prompt, "Determine whether or not the abnormal behavior occurred in the target video Y or to analyze the abnormal behavior."

[0058] Furthermore, when the detailed visual language model 6 is to perform an analysis of abnormal behavior (a more detailed understanding of the situation), it is preferable for the integrated visual language model 3 to also provide the detailed visual language model 6 with the details of the abnormal behavior (such as falling) determined by the first abnormal behavior determination unit 5. For example, if "falling" is provided as the details of the abnormal behavior, the detailed visual language model 6 will perform analyses such as "Is it really a fall?", "(If it is a fall) Is it really a high-priority emergency?", and "Which pixel movement corresponds to 'falling'?".

[0059] As explained above, according to the modified abnormal behavior determination system 1, if the abnormal behavior determined by the overall visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, the detailed visual language model 6 is controlled to determine whether or not the abnormal behavior occurred in the target video Y or to analyze the abnormal behavior.

[0060] With this configuration, when a low-priority abnormal behavior occurs, the overall load on the abnormal behavior detection system 1 (GPU or CPU usage) is kept low, while when a high-priority abnormal behavior occurs, the detailed visual language model 6 can perform abnormal behavior detection with even greater accuracy.

[0061] Furthermore, if the abnormal behavior determined by the Integrated Visual Language Model 3 or the First Abnormal Behavior Determination Unit 5 is deemed to be of high urgency, the Integrated Visual Language Model 3 may control other devices to identify and track the target behavioral entity Z that performed the abnormal behavior.

[0062] For example, if we want to identify a target behavioral entity Z that has exhibited abnormal behavior, as shown in Figure 5, we can consider a configuration in which "the system further includes an appearance analysis unit 7 capable of analyzing the appearance of the behavioral entity as seen in the video, the appearance analysis unit 7 can analyze the appearance of the same target behavioral entity Z within the same target video Y with a higher accuracy than the integrated visual language model 3, the memory unit 4 further stores highly urgent abnormal behaviors, and the integrated visual language model 3 controls the appearance analysis unit 7 to analyze the appearance of the target behavioral entity Z if the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is highly urgent." In other words, in this case, the integrated visual language model 3 also functions as the control unit for the appearance analysis unit 7. For appearance analysis, it is conceivable to use facial recognition, color recognition, body size recognition, etc., to depict the target behavioral entity Z.

[0063] For example, the integrated visual language model 3 could be pre-configured with a prompt stating, "If the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, control the appearance analysis unit 7 to analyze the appearance of the target behavioral entity Z." If the abnormal behavior is of high urgency, the integrated visual language model 3 would then input the target video Y from its output unit (not shown) to the input unit (not shown) of the appearance analysis unit 7. The appearance analysis unit 7 would then be configured to automatically perform an appearance analysis once the target video Y is input, thus initiating the appearance analysis.

[0064] Furthermore, the Integrated Visual Language Model 3 does not necessarily need to be able to analyze the appearance of the behavioral object shown in the video. Even if it does not have the function to analyze the appearance of the behavioral object shown in the video, it is still included in the statement that "the appearance analysis unit 7 can analyze the appearance of the same target behavioral object Z within the same target video Y with a higher accuracy than the Integrated Visual Language Model 3."

[0065] Furthermore, if it is desired to track a target behavioral entity Z that has performed abnormal behavior, a configuration can be considered as shown in Figure 5, in which "an appearance identification unit 8 capable of identifying the appearance of a behavioral entity captured in the video is provided, and the appearance identification unit 8 can identify the appearance of the same target behavioral entity Z within the same target video Y with a higher accuracy than the integrated visual language model 3, the memory unit 4 further stores highly urgent abnormal behaviors and sequentially stores acquired past videos, and the integrated visual language model 3 controls the appearance identification unit 8 to identify the appearance of the target behavioral entity Z if the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is highly urgent, and sequentially extracts videos from the stored videos in which the appearance of a behavioral entity matches the appearance of the target behavioral entity Z determined to have performed abnormal behavior to a predetermined extent or more." In other words, in this case, the integrated visual language model 3 also functions as a control unit for the appearance identification unit 8. For appearance identification, facial recognition, color recognition, body size recognition, etc., can be used to extract the same person.

[0066] For example, the integrated visual language model 3 may be given a prompt in advance stating, "If the abnormal behavior determined by the integrated visual language model 3 or the first abnormal behavior determination unit 5 is of high urgency, control the appearance identification unit 8 to identify the appearance of the target behavioral entity Z." If the abnormal behavior is of high urgency, the integrated visual language model 3 will input the target video Y from the output unit (not shown) of the integrated visual language model 3 to the input unit (not shown) of the appearance identification unit 8. In this case, it is preferable to identify the target behavioral entity Z that performed the abnormal behavior before inputting the target video Y. The appearance identification unit 8 is then set to automatically identify the appearance of the target behavioral entity Z once the target video Y is input, thereby initiating appearance identification. Once appearance (feature) identification is complete, the integrated visual language model 3 will sequentially extract videos from the stored videos in which the appearance of the behavioral entity Z that was determined to have performed the abnormal behavior matches a predetermined number of images.

[0067] Furthermore, the integrated visual language model 3 does not necessarily need to be able to identify the appearance of the behavioral object shown in the video. Even if it does not have the function to identify the appearance of the behavioral object shown in the video, it is still included in the statement that "the appearance identification unit 8 can identify the appearance of the same target behavioral object Z within the same target video Y with a higher accuracy than the integrated visual language model 3."

[0068] Furthermore, if object detection is desired, as shown in Figure 5, a configuration is also possible in which "an object detection unit 9 capable of detecting or analyzing multiple objects in the video using object detection or difference detection, and a detailed visual language model 6 that has already learned the feature points of objects in the video and is capable of detecting or analyzing objects in the video, the memory unit 4 stores objects of high urgency, and the integrated visual language model 3 controls the detailed visual language model 6 to detect or analyze the object in the target video Y if the object detected by the object detection unit 9 is of high danger." In this case, object detection or analysis will be performed by the object detection unit 9 or the integrated visual language model 3 separately from the detection of the target behavioral body Z by the integrated visual language model 3. In this case, the integrated visual language model 3 will also function as a control unit for the detailed visual language model 6. For known object detection and background difference, known methods can be used. For example, in object detection, one possible method is to "memorize the shape of an object (for example, the shape of a person or a car), and then, using segmentation, detect something similar in shape and determine it to be 'the object in question'."

[0069] Furthermore, in the above embodiment, "the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior occurred by the target behavior Z or to analyze the abnormal behavior when the target behavior Z in the target image Y is smaller than the size that can be determined." However, when the behavior object is small, it is conceivable that the integrated visual language model 3 may not be able to clearly determine whether or not what is shown in the target image Y is a behavior object in the first place.

[0070] Therefore, the configuration may be such that "the integrated visual language model 3 has already learned the feature quantities of behavioral objects captured in the video, and when it determines that a set of feature quantities in the target video Y is a behavioral object with a probability within a predetermined range, it controls the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior occurred in the target video Y or to analyze the abnormal behavior." The predetermined probability range could be set to, for example, 20-80%.

[0071] With this configuration, even if the integrated visual-language model 3 cannot clearly determine whether what is shown in the target video Y is a moving object, the failure to detect abnormal behavior is suppressed.

[0072] Furthermore, in the above embodiment, the integrated visual language model 3 detected the size of the target behavioral body Z, but other devices may also perform the size detection.

[0073] [Size detection using a simplified visual language model → determination by a comprehensive visual language model → if small, determination by the first abnormal behavior determination unit]

[0074] For example, as shown in Figure 6, the simplified visual language model 10 may detect the size of the target behavioral entity Z.

[0075] In detail, the system includes: an acquisition unit 2 that acquires the target video Y; a simple visual language model 10 capable of detecting the size of the target behavioral body Z shown in the target video Y; a comprehensive visual language model 3 that has already learned the displacement of the feature points of the behavioral body when the behavioral body shown in the video performs abnormal behavior, and is capable of determining whether or not abnormal behavior has occurred by the target behavioral body Z shown in the target video Y based on the feature quantities in the target video Y, and is capable of acquiring the size of the target behavioral body Z on the target video Y detected by the simple visual language model 10; a storage unit 4 that stores the size of the behavioral body on the video that the comprehensive visual language model 3 can determine whether or not abnormal behavior has occurred based on the feature quantities in the target video Y, the feature quantities of the joints or key points of the behavioral body shown in the video, and the displacement of the joints or key points of the behavioral body when the behavioral body shown in the video performs abnormal behavior; and the feature quantities of the target behavioral body Z shown in the target video Y. A possible configuration is described as having a first abnormal behavior determination unit 5 which includes a detection unit 51 capable of detecting joints or keypoints based on the displacement of the detected joints or keypoints and a determination unit 52 capable of determining whether or not abnormal behavior has occurred by the target behavioral body Z in the target video Y or analyzing the abnormal behavior, wherein the first abnormal behavior determination unit 5 is capable of detecting the joints or keypoints of the same target behavioral body Z in the same target video Y with a higher accuracy than the integrated visual language model 3, and if the target behavioral body Z on the target video Y is larger than the size that can be determined, the integrated visual language model 3 determines whether or not abnormal behavior has occurred by the target behavioral body Z or analyzes the abnormal behavior, and if the target behavioral body Z on the target video Y is smaller than the size that can be determined, the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior has occurred by the target behavioral body Z or analyzes the abnormal behavior.

[0076] In this case, for example, the simplified visual language model 10 is given a prompt in advance that says, "If the target behavioral object Z shown in the target video Y is smaller than the size that can be judged, report that fact to the integrated visual language model 3." Also, the integrated visual language model 3 is given a prompt in advance that says, "If the above report is received, control the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior occurred in the target video Y or to analyze the abnormal behavior." The integrated visual language model 3 is configured to input the target video Y from its output unit (not shown) to the input unit (not shown) of the first abnormal behavior determination unit 5 when the above report is received. The first abnormal behavior determination unit 5 is then configured to automatically perform abnormal behavior determination etc. when the target video Y is input, thereby initiating the abnormal behavior determination etc.

[0077] With this configuration, the simplified visual language model 10 is entrusted with size detection, and if the target behavioral object Z on the target video Y is smaller than the size that can be determined, the integrated visual language model 3 can entrust the determination of abnormal behavior to the first abnormal behavior determination unit 5 without performing the determination of abnormal behavior itself. This makes it possible to further distribute the load to each device while reducing unnecessary operations in the integrated visual language model 3.

[0078] [Size detection and determination using a simplified visual language model → Integrated visual language model → If small, determination is made by the first abnormal behavior determination unit]

[0079] Furthermore, as shown in Figure 7, the decision may also be made using the simplified visual-language model 10 instead of the comprehensive visual-language model 3.

[0080] In detail, the system comprises: an acquisition unit 2 that acquires target video Y; a simplified visual language model 10 that has learned the displacement of feature points of an object when it performs abnormal behavior in the video, and is capable of determining whether or not an abnormal behavior has occurred by a target object Z in the target video Y based on the feature quantities in the target video Y, and is capable of detecting the size of the target object Z on the target video Y; a storage unit 4 that stores the size of the object on the video that the simplified visual language model 10 can determine whether or not an abnormal behavior has occurred based on the feature quantities in the target video Y, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video; a comprehensive visual language model 3 that can acquire the size of the target object Z detected by the simplified visual language model 10; and a storage unit 4 that stores the size of the joints or key points of the object when it performs abnormal behavior in the video. A possible configuration is described as comprising a first abnormal behavior determination unit 5 having a detectable detection unit 51, a determination unit 52 capable of determining whether or not abnormal behavior has occurred by a target behavioral body Z in a target video Y based on the detected joint or key point displacement and the stored joint or key point displacement, wherein the first abnormal behavior determination unit 5 can detect the same joint or key point of the same target behavioral body Z in the same target video Y with a higher accuracy than the simplified visual language model 10, and if the target behavioral body Z on the target video Y is larger than the detectable size, the integrated visual language model 3 controls the simplified visual language model 10 to determine whether or not abnormal behavior has occurred by the target behavioral body Z or to analyze abnormal behavior, and if the target behavioral body Z on the target video Y is smaller than the detectable size, the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior has occurred by the target behavioral body Z or to analyze abnormal behavior.

[0081] With this configuration, the integrated visual language model 3 becomes more specialized in integration (it does not need to have the function of determining or analyzing abnormal behavior), and the workload can be distributed among the devices.

[0082] [Size detection by object detection unit → determination by integrated visual language model → if small, determination by first abnormal behavior determination unit]

[0083] Furthermore, as shown in Figure 8, the object detection unit 9 may also detect the size of the target moving object Z.

[0084] In detail, the system includes: an acquisition unit 2 that acquires the target video Y; an object detection unit 9 that can detect the size of the target behavioral body Z shown in the target video Y using object detection or difference detection; a comprehensive visual language model 3 that has learned the displacement of feature points of a behavioral body when it performs abnormal behavior, and is capable of determining whether or not abnormal behavior has occurred by the target behavioral body Z shown in the target video Y based on the feature quantities in the target video Y, and is capable of acquiring the size of the target behavioral body Z detected by the object detection unit 9 in the target video Y; a storage unit 4 that stores the size of the behavioral body on the video that the comprehensive visual language model 3 can determine whether or not abnormal behavior has occurred based on the feature quantities in the target video Y, the feature quantities of the joints or key points of the behavioral body shown in the video, and the displacement of the joints or key points of the behavioral body when it performs abnormal behavior; and the feature quantities of the target behavioral body Z shown in the target video Y. A possible configuration is described as having a first abnormal behavior determination unit 5 which includes a detection unit 51 capable of detecting joints or key points based on characteristics, and a determination unit 52 capable of determining whether or not abnormal behavior has occurred by a target behavioral body Z in a target video Y, or analyzing abnormal behavior, based on the displacement of the detected joints or key points and the stored displacement of the joints or key points. The first abnormal behavior determination unit 5 is capable of detecting the joints or key points of the same target behavioral body Z in the same target video Y with a higher accuracy than the integrated visual language model 3. If the target behavioral body Z on the target video Y is larger than or equal to the size that can be determined, the integrated visual language model 3 determines whether or not abnormal behavior has occurred by the target behavioral body Z or analyzes abnormal behavior. If the target behavioral body Z on the target video Y is smaller than the size that can be determined, the integrated visual language model 3 controls the first abnormal behavior determination unit 5 to determine whether or not abnormal behavior has occurred by the target behavioral body Z or analyzes abnormal behavior.

[0085] With this configuration, the integrated visual language model 3 can distribute the load to each device by having the object detection unit 9 perform size detection.

[0086] [Judgment by the Integrated Visual Language Model → If small, judgment is made by the second abnormal behavior judgment unit]

[0087] Furthermore, in the above embodiment, the first abnormal behavior determination unit 5 determined whether or not an abnormal behavior occurred or analyzed the abnormal behavior based on joints or key points. However, as shown in Figure 9, the second abnormal behavior determination unit 5A may determine whether or not an abnormal behavior occurred or analyze the abnormal behavior based on the shape of the behaving body detected using object detection or difference detection.

[0088] In detail, the system includes: an acquisition unit 2 that acquires target video Y; a comprehensive visual language model 3 that has learned the displacement of feature points of an object when it performs abnormal behavior in the video, and is capable of determining whether or not an abnormal behavior has occurred by a target object Z in the target video Y based on the feature quantities in the target video Y, and is capable of detecting the size of the target object Z in the target video Y on the target video Y; a storage unit 4 that stores the size of the object on the video that the comprehensive visual language model 3 can determine whether or not an abnormal behavior has occurred based on the feature quantities in the target video Y, the feature quantities of the shape of the object in the video, and the displacement of the shape of the object when it performs abnormal behavior in the video; and an object detection or difference detection that uses the target video A possible configuration is described as having a second abnormal behavior determination unit 5A which includes a detection unit 51A capable of detecting the shape of the target behavioral body Z reflected in Y, and a determination unit 52A capable of determining whether or not abnormal behavior was performed by the target behavioral body Z in the target image Y or analyzing the abnormal behavior based on the displacement of the shape of the detected target behavioral body Z and the displacement of the shape of the behavioral body, wherein the integrated visual language model 3 controls the second abnormal behavior determination unit 5A to determine whether or not abnormal behavior was performed by the target behavioral body Z or analyze the abnormal behavior when the target behavioral body Z on the target image Y is larger than the size that can be determined, and when the target behavioral body Z on the target image Y is smaller than the size that can be determined, it determines whether or not abnormal behavior was performed by the target behavioral body Z or analyzes the abnormal behavior.

[0089] With this configuration, taking advantage of the fact that object detection or difference detection can generally detect even small moving objects, if the target moving object Z is smaller than the detectable size, the second abnormal behavior determination unit 5A, which uses object detection or difference detection, performs abnormal behavior determination, etc., making it possible to perform abnormal behavior determination, etc., efficiently.

[0090] Furthermore, when using object detection or difference detection, the accuracy rate is often lower than when using joints or keypoints. Therefore, depending on the judgment / analysis results of the second abnormal behavior determination unit 5A (for example, when abnormal behavior is judged to be an action with a probability within a predetermined range (20-80%)), the detailed visual language model 6 may perform further judgments.

[0091] [Size detection using a simplified visual-language model → Judgment by a comprehensive visual-language model → If small, judgment by a second abnormal behavior judgment unit]

[0092] Furthermore, when the second abnormal behavior determination unit 5A makes a determination, the simplified visual language model 10 may also detect the size of the target behavioral entity Z, as shown in Figure 10.

[0093] In detail, the system includes: an acquisition unit 2 that acquires target video; a simple visual language model 10 capable of detecting the size of target behavioral body Z on target video Y as seen in target video Y; a comprehensive visual language model 3 that has already learned the displacement of feature points of behavioral bodies when they perform abnormal behavior, and is capable of determining whether or not abnormal behavior has occurred by target behavioral body Z seen in target video Y based on the feature quantities in target video Y, and is capable of acquiring the size of target behavioral body Z on target video Y detected by the simple visual language model 10; and the size of the behavioral body on the video that the comprehensive visual language model 3 can determine whether or not abnormal behavior has occurred based on the feature quantities in target video Y, the feature quantities of the shape of the behavioral body seen in the video, and the displacement of the shape of the behavioral body when it performs abnormal behavior. A possible configuration is described as having a second abnormal behavior determination unit 5A which includes a stored memory unit 4, a detection unit 51A capable of detecting the shape of a target behavioral body Z reflected in the target video Y using object detection or difference detection, and a determination unit 52A capable of determining whether or not abnormal behavior was performed by the target behavioral body Z in the target video Y or analyzing abnormal behavior based on the displacement of the detected shape of the target behavioral body Z and the displacement of the stored shape of the behavioral body. The second abnormal behavior determination unit 5A is controlled so that if the target behavioral body Z on the target video Y is larger than or equal to the size that can be determined, the integrated visual language model 3 determines whether or not abnormal behavior was performed by the target behavioral body or analyzes abnormal behavior, and if the target behavioral body Z on the target video Y is smaller than the size that can be determined, the integrated visual language model 3 determines whether or not abnormal behavior was performed by the target behavioral body Z or analyzes abnormal behavior.

[0094] [Size detection and determination using a simplified visual language model → Comprehensive visual language model → If small, determination is made by the second abnormal behavior determination unit]

[0095] Similarly, as shown in Figure 11, the simplified visual language model 10 may be used to perform further determinations instead of the comprehensive visual language model 3.

[0096] In detail, the system comprises: an acquisition unit 2 that acquires the target video Y; a simplified visual language model 10 that has already learned the displacement of feature points of an object when it performs abnormal behavior in the video, and is capable of determining whether or not an abnormal behavior has occurred by the target object Z shown in the target video Y based on the feature quantities in the target video Y, and is capable of detecting the size of the target object Z on the target video Y; a storage unit 4 that stores the size of the object on the video that the simplified visual language model 10 can determine whether or not an abnormal behavior has occurred based on the feature quantities in the target video Y, the feature quantities of the shape of the object shown in the video, and the displacement of the shape of the object when it performs abnormal behavior in the video; and a comprehensive visual language model 3 that can acquire the size of the target object Z on the target video Y detected by the simplified visual language model 10, and an object A possible configuration is described as having a second abnormal behavior determination unit 5A which includes a detection unit 51A capable of detecting the shape of a target behavioral body Z reflected in a target video Y using detection or difference detection, and a determination unit 52A capable of determining whether or not abnormal behavior was performed by the target behavioral body Z in the target video Y or analyzing abnormal behavior based on the displacement of the shape of the detected target behavioral body Z and the displacement of the shape of the behavioral body, wherein if the target behavioral body Z on the target video Y is larger than the size that can be determined, the integrated visual language model 3 controls the simplified visual language model 10 to determine whether or not abnormal behavior was performed by the target behavioral body Z or analyze abnormal behavior, and if the target behavioral body Z on the target video Y is smaller than the size that can be determined, the integrated visual language model 3 controls the second abnormal behavior determination unit to determine whether or not abnormal behavior was performed by the target behavioral body Z or analyze abnormal behavior.

[0097] Furthermore, in the above embodiment, the appearance analysis unit 7 and the appearance identification unit 8 were separate from the detailed visual language model 6, but the detailed visual language model 6 may also function as the appearance analysis unit 7 and the appearance identification unit 8, and this case is also included in the present invention.

[0098] Furthermore, in the above embodiment, the characteristic quantities of the joints (neck, right elbow, left elbow, waist, right knee, left knee, etc.) or key points (between the eyes, etc.) of the moving body shown in the video, and the displacement of the joints or key points of the moving body when the moving body shown in the video performs abnormal behavior, were stored in the memory unit 4, but they may also be stored in the memory unit (not shown) within the first abnormal behavior determination unit 5. Similarly, the characteristic quantities of the shape of the moving body shown in the video, and the displacement of the shape of the moving body when the moving body shown in the video performs abnormal behavior, may also be stored in the memory unit (not shown) within the second abnormal behavior determination unit 5A.

[0099] Furthermore, the present invention can also be applied to programs and methods corresponding to the processing performed by each component acting as a controller, as well as to recording media that store such programs. In the case of recording media, the program will be installed on a computer or the like. Here, the recording media that stores the program may be a non-transient recording media. Examples of non-transient recording media include CD-ROMs, but the invention is not limited to them. [Explanation of Symbols]

[0100] 1. Abnormal Behavior Detection System 2 Acquisition part 3. Comprehensive Visual Language Model 4 Storage section 5. First abnormal behavior determination unit 5A Second abnormal behavior detection unit 6. Detailed Visual Language Model 7. Visual Analysis Department 8. External Identification Section 9. Object detection unit 10. Simplified Visual Language Model 31 Input Layer 32 Middle Class 33 Output Layer 51 Detection unit 51A Detection Unit 52 Judgment section 52A Judgment section X Photography Method Y Target video Z Target Behavior Z Target video

Claims

1. An acquisition unit that acquires the target video, A comprehensive visual language model that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not the abnormal behavior was performed by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of detecting the size of the object in the target video. A storage unit that stores the following, which allows the integrated visual language model to determine whether or not the abnormal behavior occurred based on the features in the target video, or to analyze the abnormal behavior: the size of the behavioral body in the video that can be determined from the features in the target video; the features of the joints or key points of the behavioral body shown in the video; and the displacement of the joints or key points of the behavioral body when the behavioral body shown in the video performs abnormal behavior. A first abnormal behavior determination unit comprising: a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body captured in the target video; and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points; An abnormal behavior detection system equipped with, The first abnormal behavior detection unit is capable of detecting the same joint or key point of the target behavioral body within the same target video with a higher accuracy than the integrated visual language model. An abnormal behavior determination system characterized in that, when the target behavior object on the target video is larger than the size that can be determined, the integrated visual language model determines whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior, and when the target behavior object on the target video is smaller than the size that can be determined, the integrated visual language model controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior.

2. A detailed visual language model has been trained to learn the displacement of characteristic points of an object when it performs abnormal behavior in a video, and has a higher accuracy rate than the first abnormal behavior determination unit when determining whether the same abnormal behavior occurred in the same target video or when analyzing the abnormal behavior. Furthermore, The aforementioned memory unit further stores highly urgent abnormal behaviors. The abnormal behavior determination system according to claim 1, characterized in that the integrated visual language model controls the detailed visual language model to determine whether or not the abnormal behavior was performed by the target behavioral entity or to analyze the abnormal behavior when the abnormal behavior determined by the integrated visual language model or the first abnormal behavior determination unit is of high urgency.

3. It is further equipped with an appearance analysis unit capable of analyzing the appearance of the moving object captured in the video. The aforementioned appearance analysis unit is capable of analyzing the appearance of the same target behavioral object within the same target video with a higher accuracy rate than the comprehensive visual language model. The aforementioned memory unit further stores highly urgent abnormal behaviors. The abnormal behavior determination system according to claim 1, characterized in that the integrated visual language model controls the appearance analysis unit to analyze the appearance of the target behavioral body if the abnormal behavior determined by the integrated visual language model or the first abnormal behavior determination unit is of high urgency.

4. It further includes an appearance identification unit capable of identifying the appearance of the moving object captured in the video. The appearance identification unit is capable of identifying the appearance of the same target behavioral body within the same target video with a higher accuracy than the integrated visual language model. The memory unit further stores highly urgent abnormal behaviors and sequentially stores the acquired video footage. The unified visual language model controls the appearance identification unit to identify the appearance of the target behavioral body if the unified visual language model or the first unnatural behavior determination unit determines that the unnatural behavior is of high urgency, and extracts from the sequentially stored images images in which the behavioral body matches the appearance of the target behavioral body determined to have performed the unnatural behavior to a predetermined extent or more. This is the unnatural behavior determination system according to 1.

5. An object detection unit capable of detecting or analyzing multiple objects in a video using object detection or difference detection, A detailed visual language model that has already learned the feature points of objects shown in the video and is capable of detecting or analyzing objects shown in the video, Furthermore, The aforementioned memory unit stores items of high urgency, Abnormal behavior determination system according to 1, characterized in that the overall visual language model controls the detailed visual language model to detect or analyze the object in the target video if the object detected by the object detection unit is of high urgency.

6. The abnormal behavior determination system according to claim 1, characterized in that the integrated visual language model has already learned the feature quantities of behavioral objects captured in the video, and when it determines that a set of feature quantities in the target video is a behavioral object with a probability within a predetermined range, it controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior occurred in the target video or to analyze the abnormal behavior.

7. An acquisition unit that acquires the target video, A simple visual language model capable of detecting the size of the target object in the target video, A comprehensive visual language model that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not the abnormal behavior was performed by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of obtaining the size of the object in the target video detected by the simplified visual language model, A storage unit that stores the following, which allows the integrated visual language model to determine whether or not the abnormal behavior occurred based on the features in the target video, or to analyze the abnormal behavior: the size of the behavioral body in the video that can be determined from the features in the target video; the features of the joints or key points of the behavioral body shown in the video; and the displacement of the joints or key points of the behavioral body when the behavioral body shown in the video performs abnormal behavior. A first abnormal behavior determination unit comprising: a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body captured in the target video; and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points; An abnormal behavior detection system equipped with, The first abnormal behavior detection unit is capable of detecting the same joint or key point of the target behavioral body within the same target video with a higher accuracy than the integrated visual language model. An abnormal behavior determination system characterized in that, when the target behavior object on the target video is larger than the size that can be determined, the integrated visual language model determines whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior, and when the target behavior object on the target video is smaller than the size that can be determined, the integrated visual language model controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior.

8. An acquisition unit that acquires the target video, A simplified visual language model has been trained to learn the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not the abnormal behavior was performed by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of detecting the size of the object in the target video. The simplified visual language model can determine whether or not the abnormal behavior occurred based on the features in the target video, or analyze the abnormal behavior. The memory unit stores the following: the determinable size of the behavioring body in the video, the features of the joints or key points of the behavioring body shown in the video, and the displacement of the joints or key points of the behavioring body when the behavioring body shown in the video performs an abnormal behavior. A comprehensive visual language model capable of obtaining the size of the target behavioral body detected by the simplified visual language model, A first abnormal behavior determination unit comprising: a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body captured in the target video; and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points; An abnormal behavior detection system equipped with, The first abnormal behavior detection unit is capable of detecting the same joint or key point of the target behavioral body within the same target video with a higher accuracy than the simplified visual language model. An abnormal behavior determination system characterized in that, when the target behavior object on the target video is larger than the size that can be determined, the integrated visual language model controls the simplified visual language model to determine whether or not the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior, and when the target behavior object on the target video is smaller than the size that can be determined, the integrated visual language model controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior.

9. An acquisition unit that acquires the target video, An object detection unit capable of detecting the size of a target moving object captured in the target video using object detection or difference detection, A comprehensive visual language model that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not the abnormal behavior was performed by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of obtaining the size of the object detected by the object detection unit in the target video, A storage unit that stores the following, which allows the integrated visual language model to determine whether or not the abnormal behavior occurred based on the features in the target video, or to analyze the abnormal behavior: the size of the behavioral body in the video that can be determined from the features in the target video; the features of the joints or key points of the behavioral body shown in the video; and the displacement of the joints or key points of the behavioral body when the behavioral body shown in the video performs abnormal behavior. A first abnormal behavior determination unit comprising: a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body captured in the target video; and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points; An abnormal behavior detection system equipped with, The first abnormal behavior detection unit is capable of detecting the same joint or key point of the target behavioral body within the same target video with a higher accuracy than the integrated visual language model. An abnormal behavior determination system characterized in that, when the target behavior object on the target video is larger than the size that can be determined, the integrated visual language model determines whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior, and when the target behavior object on the target video is smaller than the size that can be determined, the integrated visual language model controls the first abnormal behavior determination unit to determine whether or not the abnormal behavior occurred by the target behavior object or analyzes the abnormal behavior.

10. A program executed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not an abnormal behavior occurred by an object in a target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and which can detect the size of an object in a target video on the target video, has stored the following: the determinable size on the video of an object that can determine whether or not an abnormal behavior occurred based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video, The steps include acquiring the aforementioned target video, The first abnormal behavior determination unit is controlled to determine whether the abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.

11. A method executed on a computer that has learned the displacement of feature points of an object when an object in a video performs abnormal behavior, and is capable of determining whether or not an abnormal behavior occurred by an object in a target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and is capable of detecting the size of an object in a target video on the target video, the method being executed on a computer that stores the following: the size of an object in a video that can determine whether or not an abnormal behavior occurred based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object in a video when the object in the video performs abnormal behavior, The steps include acquiring the aforementioned target video, The first abnormal behavior determination unit is controlled to determine whether an abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether an abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined, and to determine whether an abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.

12. A program executed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and which is capable of determining whether or not an abnormal behavior occurred or analyzing the abnormal behavior by an object in a target video based on the feature quantities in the target video, has stored the following: the determinable size of an object in the video that can determine whether or not an abnormal behavior occurred or analyze the abnormal behavior based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video. The steps include acquiring the aforementioned target video, The steps include: obtaining the size of the target behavior object in the target video, which is detected by a simplified visual language model capable of detecting the size of the target behavior object in the target video; The first abnormal behavior determination unit is controlled to determine whether the abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.

13. A method is performed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and which is capable of determining whether or not an abnormal behavior occurred or analyzing the abnormal behavior by an object in a target video based on the feature quantities in the target video, and which has stored the following information: the determinable size of an object in the video that can determine whether or not an abnormal behavior occurred or analyze the abnormal behavior based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video. The steps include acquiring the aforementioned target video, The steps include: obtaining the size of the target behavior object in the target video, which is detected by a simplified visual language model capable of detecting the size of the target behavior object in the target video; The first abnormal behavior determination unit is controlled to determine whether the abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.

14. A program executed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and is capable of determining whether or not the abnormal behavior was performed by an object in the target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and which has stored the following: the determinable size of the object in the video that can determine whether or not the abnormal behavior was performed based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video, The steps include acquiring the aforementioned target video, The steps include obtaining the size of the target behavioral body detected by the simplified visual language model on the target image, The steps include: controlling the simplified visual language model to determine whether the abnormal behavior occurred by the target behavior object in the target video if the target behavior object is larger than or equal to the size that can be determined, or to analyze the abnormal behavior; and controlling the first abnormal behavior determination unit to determine whether the abnormal behavior occurred by the target behavior object if the target behavior object is smaller than the size that can be determined, or to analyze the abnormal behavior. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the simplified visual language model.

15. A method executed on a computer that has learned the displacement of feature points of an object when an object in a video performs abnormal behavior, and is capable of determining whether or not an abnormal behavior occurred by an object in a target video based on the feature quantities in the target video, or analyzing the abnormal behavior, and which has stored the following: the determinable size on the video of an object that can determine whether or not an abnormal behavior occurred based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when the object in the video performs abnormal behavior; The steps include acquiring the aforementioned target video, The steps include obtaining the size of the target behavioral body detected by the simplified visual language model on the target image, The steps include: controlling the simplified visual language model to determine whether the abnormal behavior occurred by the target behavior object in the target video if the target behavior object is larger than or equal to the size that can be determined, or to analyze the abnormal behavior; and controlling the first abnormal behavior determination unit to determine whether the abnormal behavior occurred by the target behavior object if the target behavior object is smaller than the size that can be determined, or to analyze the abnormal behavior. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the simplified visual language model.

16. A program executed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and which is capable of determining whether or not an abnormal behavior occurred or analyzing the abnormal behavior by an object in a target video based on the feature quantities in the target video, has stored the following: the determinable size of an object in the video that can determine whether or not an abnormal behavior occurred or analyze the abnormal behavior based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video. The steps include acquiring the aforementioned target video, A step of obtaining the size of the target object in the target video detected by an object detection unit capable of detecting the size of the target object in the target video using object detection or difference detection, The first abnormal behavior determination unit is controlled to determine whether the abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.

17. A method is performed on a computer that has learned the displacement of feature points of an object when it performs abnormal behavior in a video, and which is capable of determining whether or not an abnormal behavior occurred or analyzing the abnormal behavior by an object in a target video based on the feature quantities in the target video, and which has stored the following information: the determinable size of an object in the video that can determine whether or not an abnormal behavior occurred or analyze the abnormal behavior based on the feature quantities in the target video, the feature quantities of the joints or key points of the object in the video, and the displacement of the joints or key points of the object when it performs abnormal behavior in the video. The steps include acquiring the aforementioned target video, A step of obtaining the size of the target object in the target video detected by an object detection unit capable of detecting the size of the target object in the target video using object detection or difference detection, The first abnormal behavior determination unit is controlled to determine whether the abnormal behavior occurred by the target behavior object in the target video is larger than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object in the target video is smaller than the size that can be determined, and to determine whether the abnormal behavior occurred by the target behavior object or to analyze the abnormal behavior if the target behavior object is smaller than the size that can be determined. Equipped with, The first abnormal behavior determination unit comprises a detection unit capable of detecting the joints or key points based on the characteristic quantities of the target behavioral body shown in the target video, and a determination unit capable of determining whether or not the abnormal behavior was performed by the target behavioral body in the target video or analyzing the abnormal behavior based on the displacement of the detected joints or key points and the stored displacement of the joints or key points, and is characterized in that it can detect the same joints or key points of the same target behavioral body in the same target video with a higher accuracy than the integrated visual language model.