Method for generating video contextual information using vision-language models and non-volatile computer-readable storage medium storing the same

KR103025470B1Active Publication Date: 2026-09-29JD1 CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
KR1020260043939
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-03-11
Publication Date
2026-09-29
Estimated Expiration
2046-03-11

Smart Images

  • Figure 112026029698163-PAT00001_ABST
    Figure 112026029698163-PAT00001_ABST
Patent Text Reader

Abstract

The present specification discloses a method for generating video situation information and a non-volatile computer-readable storage medium storing the same. The method for generating video situation information according to the present specification may include: (a) inputting a plurality of videos received from a plurality of video capturing devices installed at a preset location into an artificial neural network model to acquire situation analysis information for each video according to a preset time interval; (b) selecting a recommended action information corresponding to each situation analysis information from among a plurality of recommended action information stored in the storage medium, and setting a risk level for each situation analysis information according to a preset criterion to generate a plurality of integrated information including situation analysis information, recommended action information, and risk level; and (c) displaying the plurality of integrated information on a display device according to a preset priority. The artificial neural network model may be trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to the situation information for the input video.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for generating image situation information, and more specifically, to a method for generating image situation information using a visual language model and a non-volatile computer-readable storage medium storing the same. Background Technology

[0002] The content described in this section merely provides background information regarding the embodiments described in this specification and does not necessarily constitute prior art.

[0003] Intelligent video analysis technology utilizes a deep learning-based artificial neural network model to perform object detection and simple action recognition within a video, and generates information about the situation within the video. The intelligent video analysis technology is effective for generating identification information about objects such as people and vehicles in a video, or for generating information about fragmentary situations such as intrusion or falling. The intelligent video analysis technology can be utilized in control systems, security systems, etc.

[0004] As mentioned above, conventional intelligent image analysis technology can generate identification information about objects or information about fragmentary situations. However, conventional intelligent image analysis technology has limitations in generating information about the specific 'context' in which a particular situation occurred or the 'meaning' of a particular situation.

[0005] Furthermore, conventional intelligent video analysis technology frequently generates false positives, which produce incorrect information due to environmental factors such as lighting changes and weather. Consequently, control personnel operating the system face situations where they must distinguish between accurate information and erroneous information resulting from these false positives. This increases the workload of control personnel and undermines the reliability of the control system.

[0006] Furthermore, conventional intelligent video analysis technology generates information only regarding the situation within multiple videos and does not generate importance information for the information produced in each video. As a result, control personnel must individually identify which situations are critical in order to respond. In this case, control personnel may be unable to respond immediately to urgent situations, potentially leading to a worsening of the situation.

[0007] Therefore, technology is required to automatically generate specific situational information and response measure information so that control personnel can respond immediately in the event of an emergency. Prior art literature

[0008] Registered Patent Publication No. 10-1321444, Oct. 16, 2013. The problem to be solved

[0009] The present specification aims to provide a method for generating image situation information using a visual language model and a non-volatile computer-readable storage medium storing the same.

[0010] This specification is not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0011] A method for generating video situation information according to the present specification for solving the above-described problem may be implemented by a non-volatile computer-readable storage medium storing recommended action information according to a plurality of situation types, and computer-executable instructions written so that when executed by one or more processors, one or more processors perform each step of the method for generating video situation information. The computer-executable instructions may include: (a) a step of inputting a plurality of videos received from a plurality of video capturing devices installed at a preset location into an artificial neural network model to obtain situation analysis information for each video according to a preset time interval; (b) a step of selecting recommended action information corresponding to each situation analysis information among the plurality of recommended action information, setting a risk level for each situation analysis information according to a preset criterion, and generating a plurality of integrated information including situation analysis information, recommended action information, and risk level; and (c) a step of displaying the plurality of integrated information on a display device according to a preset priority. The artificial neural network model described above can be trained to generate situation analysis information for an input video, including at least one of a situation type, situation information, and an object related to the situation information, using training data that includes at least one of a plurality of videos, a situation type for each video, situation information, and an object related to each situation information.

[0012] According to one embodiment of the present specification, the artificial neural network model is trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to each situation in an input video using training data that includes at least one of a situation type, situation information, and an object related to each situation, when at least two of a plurality of situation types are included in at least one of a plurality of video types that are pre-set in the input video. The step (a) may be a step of obtaining situation analysis information for each situation type in a video when at least two of a plurality of situation types are pre-set in at least one of a plurality of video types that are input to the artificial neural network model.

[0013] According to one embodiment of the present specification, step (c) may be a step of preferentially displaying integrated information with a relatively high risk level.

[0014] According to one embodiment of the present specification, step (c) may be a step of displaying the plurality of integrated information according to a situation type that is pre-set to display them preferentially when there are a plurality of integrated information with the same risk level.

[0015] According to one embodiment of the present specification, step (a) is a step of further generating time information at which each situation analysis information was obtained, and step (c) may be a step of prioritizing the display of the most recently generated integrated information when there are multiple integrated informations with the same risk level.

[0016] According to one embodiment of the present specification, step (a) is a step of further generating time information at which each situation analysis information was acquired, and step (c) may be a step of prioritizing the display of the most recently generated integrated analysis information when there are multiple integrated informations having the same risk level and situation type.

[0017] According to one embodiment of the present specification, step (c) may be a step of further displaying location information where a corresponding image capturing device is installed for each of the integrated information.

[0018] According to one embodiment of the present specification, step (c) may be a step of further outputting an alarm output signal to a speaker device according to a preset risk level condition.

[0019] According to one embodiment of the present specification, step (c) may be a step of displaying at least one integrated information among the plurality of integrated information that is not classified as a false positive according to a preset criterion on the display device.

[0020] According to one embodiment of the present specification, step (c) may include classifying the first integrated information as a false detection when the similarity between the first integrated information generated from any one video and the second integrated information generated earlier than the first integrated information is less than a preset threshold value.

[0021] According to one embodiment of the present specification, the computer-executable instruction may further include a risk level adjustment step of gradually increasing the risk level of said integrated information when integrated information having a preset risk level is repeatedly generated for a preset time.

[0022] According to one embodiment of the present specification, the risk level adjustment step may be a step of further changing the situation type for the integrated information when the risk level for the integrated information increases.

[0023] According to one embodiment of the present specification, the risk level adjustment step may be a step of further displaying the integrated information on the display device when the risk level of the integrated information becomes greater than or equal to a preset risk level.

[0024] According to one embodiment of the present specification, the risk level adjustment step may be a step of further outputting an alarm output signal to a speaker device when the risk level of the integrated information becomes greater than or equal to a preset risk level.

[0025] A non-volatile computer-readable storage medium according to the present specification may be a component of a screening control system comprising a plurality of image capturing devices installed at preset locations and a display device that displays the plurality of integrated information.

[0026] Other specific details of the present invention are included in the detailed description and drawings. Effects of the invention

[0027] According to one aspect of the present specification, a method for generating video situation information can be used to construct a selective control system that can reduce the control personnel's judgment area compared to conventional methods by transmitting situation analysis information to control personnel without video screens.

[0028] According to another aspect of the present specification, the video situation information generation method can help automatically generate and provide initial briefing information to relevant agencies in order to improve initial response capabilities compared to conventional methods.

[0029] According to another aspect of the present specification, the video situation information generation method displays situation analysis information on a display device according to priority, thereby helping to operate a stable system with a minimum number of personnel regardless of the increase in CCTV video during the operation of a selective control system, and can increase the work efficiency of control personnel compared to conventional methods.

[0030] According to another aspect of the present specification, the video situation information generation method can reduce false detection compared to conventional methods by analyzing the situation prior to the analysis point of the CCTV video together.

[0031] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0032] FIG. 1 is a flowchart of a method for generating image situation information according to one embodiment of the present specification. Figure 2 is an example diagram showing integrated information displayed on a display device. Figure 3 is an example of training data for an artificial neural network model. Figure 4 is an example of situation analysis information generated by an artificial neural network model. Figure 5 is another example of integrated information displayed on a display device. FIG. 6 is a flowchart of a method for generating image situation information according to another embodiment of the present specification. Specific details for implementing the invention

[0033] The advantages and features of the invention disclosed herein, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, this specification is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of this specification is complete and to fully inform those skilled in the art (hereinafter referred to as "skilled in the art") of the scope of this specification, and the scope of rights of this specification is defined only by the scope of the claims.

[0034] The terms used herein are for describing the embodiments and are not intended to limit the scope of the claims herein. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. As used herein, "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the components mentioned.

[0035] Throughout the specification, the same reference numerals refer to the same components, and "and / or" includes each of the mentioned components and all combinations of one or more thereof. Although terms such as "first," "second," etc., are used to describe various components, they are not limited by these terms. These terms are used merely to distinguish one component from another. Accordingly, the first component mentioned below may be the second component within the scope of the technical concept of the present invention.

[0036] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by a person skilled in the art to which this specification pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0037] The method for generating image situation information according to the present specification may be implemented in the form of a computer program. When the computer program is executed by one or more processors, it may be stored on a non-volatile computer-readable storage medium as computer-executable instructions written to enable one or more processors to perform each step of the method for generating image situation information according to the present specification.

[0038] The above computer program may include code encoded in a computer language such as C / C++, C#, JAVA, Python, or machine language, which can be read by the computer's processor (CPU) through the computer's device interface, in order for the computer to read the program and execute the methods implemented in the program. Such code may include functional code related to functions that define the necessary functions for executing the methods, and may include control code related to execution procedures necessary for the computer's processor to execute the functions according to a predetermined procedure. Additionally, such code may further include memory reference code regarding where (address) additional information or media necessary for the computer's processor to execute the functions should be referenced in the computer's internal or external memory. In addition, if the processor of the computer needs to communicate with any other computer or server located remotely in order to execute the above functions, the code may further include communication-related code regarding how to communicate with any other computer or server located remotely using the communication module of the computer, and what information or media to transmit or receive during communication.

[0039] The above-mentioned storage medium refers to a medium that stores data semi-permanently and is readable by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, examples of the above-mentioned storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the above-mentioned program may be stored on various recording media on various servers that the computer can access, or on various recording media on the user's computer. Additionally, the above-mentioned medium may be distributed across networked computer systems, and computer-readable code may be stored in a distributed manner.

[0040] An Artificial Neural Network (ANN) implements artificial intelligence by connecting artificial neurons that mathematically model the neurons constituting the human brain.

[0041] In this specification, the term "artificial neural network model" may consist of a set of interconnected computational units that may generally be referred to as nodes. These nodes may also be referred to as neurons. A neural network is composed of at least one node. The nodes (or neurons) constituting the neural networks may be interconnected by one or more links.

[0042] In this specification, "inputting" data into an artificial neural network model means that a value is input to the initial input node. In this specification, "obtaining a value," "outputting data," "obtaining information," etc., from an artificial neural network means that data is output from the final output node.

[0043] In this specification, "learning" of an artificial neural network model means that the neural network updates the connection weights of each node so that the error of the output is minimized, and "learning" according to this specification is not limited by a specific learning method.

[0044] Information regarding the nodes and weights of the artificial neural network model can be stored in the storage medium.

[0045] In this specification, the term "processor" may be composed of one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), or a neural processing unit (NPU) of a computing device. The processor may read a computer program stored in memory and perform data processing for machine learning according to one embodiment of this specification. According to one embodiment of this specification, the processor may perform computations for training a neural network. The processor may perform computations for training a neural network, such as processing input data for training in deep learning (DL), extracting features from input data, calculating errors, and updating the weights of the neural network using backpropagation. At least one of the CPU, GPGPU, TPU, and NPU of the processor may process the training of a network function. For example, the CPU and GPGPU may together process the training of a network function and data classification using the network function. In addition, in one embodiment of the present specification, processors of a plurality of computing devices may be used together to process the learning of a network function and data classification using a network function. In addition, a computer program executed on a computing device according to one embodiment of the present specification may be a CPU, GPGPU, TPU, or NPU executable program.

[0046] When the artificial neural network model is executed by the above processor, information regarding the nodes and weights of the artificial neural network model can be loaded into the memory of a GPGPU, TPU, or NPU for execution.

[0047] In addition, the processor may include a general-purpose processor, an ASIC (application-specific integrated circuit), other chipsets, logic circuits, registers, communication modems, data processing devices, etc., known in the art to which the present invention belongs, to execute output and various control logic. Furthermore, when the logic to be described below is implemented in software, the software may be stored in a memory device and executed by a processor.

[0048] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0049] FIG. 1 is a flowchart of a method for generating image situation information according to one embodiment of the present specification, and FIG. 2 is an example diagram showing integrated information displayed on a display device.

[0050] Referring to FIGS. 1 and FIGS. 2, in step S10, the processor can input a plurality of videos received from a plurality of video capturing devices (100) installed at preset locations to an artificial neural network model (110) to obtain situation analysis information for each video according to preset time intervals.

[0051] The plurality of video recording devices (100) mentioned above may refer to CCTV cameras that record closed-circuit television (CCTV) video. The CCTV cameras are examples and are not limited thereto.

[0052] The above-mentioned plurality of video recording devices (100) can digitize an optical signal using an image sensor and an image signal processor, packetize it into a video stream, and transmit it through a network. The processor can receive video in real time from the above-mentioned plurality of video recording devices (100). At this time, the processor can also receive identification information such as the IP address, MAC address, and access URL of each video recording device. Since the technology for receiving video from a CCTV camera is widely known among those skilled in the art, a detailed description is omitted.

[0053] In the present specification, the situation analysis information may refer to information generated in text form from the analysis of the input video by the artificial neural network model (110).

[0054] The processor can input the plurality of videos to the artificial neural network model (110) to obtain situation analysis information according to each video.

[0055] The artificial neural network model (110) above may correspond to a Vision-Language Model (VLM) that generates text information using video as input data. The Vision-Language Model can generate text regarding situations occurring in a specific area or the entire area of ​​the video by utilizing visual prompting technology. Since visual prompting is a technology widely known among those skilled in the art, a detailed description is omitted.

[0056] According to one embodiment of the present specification, the artificial neural network model (110) can be trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to the situation information for an input video using training data including at least one of a plurality of videos, a situation type for each video, situation information, and an object related to each situation information.

[0057] The processor can train the artificial neural network model (110) to generate situation analysis information including at least one of a situation type, situation information, and an object related to the situation information for an input video using training data including a plurality of videos, a situation type for each video, situation information, and an object related to each situation information.

[0058] The above situation type may refer to classification information regarding situations within the video. The above situation type may include normal situations and multiple pre-set abnormal situations such as traffic accidents, assault, loitering, and fire.

[0059] The above situational information may refer to detailed information regarding the situation included in the video. For example, regarding a traffic accident, the situational information may include details such as the vehicles involved, whether there are injured persons, and the possibility of additional accidents.

[0060] The object related to the above situation information may refer to at least one object among a plurality of objects included in the video that is required to generate the above situation information. For example, regarding a traffic accident situation, the object related to the situation information may include the accident vehicle, driver, injured person, etc.

[0061] Figure 3 is an example of training data for an artificial neural network model.

[0062] In the example of Fig. 3, the video may include a traffic accident situation that occurred at an intersection.

[0063] In the example of Fig. 3, the situation type may correspond to a traffic accident. The situation information may include detailed information about the situation, such as, "After a white SUV (Sport Utility Vehicle) and a silver sedan collided at an intersection, the two vehicles are stopped in the center of the intersection. The driver has gotten out of the vehicle and is checking the accident scene. Nearby vehicles are passing through the accident scene, and pedestrians are staying near the scene. No additional collisions or hazards were observed after the accident."

[0064] Objects related to the above situation information may include a white SUV, a silver sedan, a driver, surrounding vehicles, and pedestrians. Objects related to the above situation information may be displayed together in the image frames of the video.

[0065] Through this, the artificial neural network model can generate situation analysis information more precise than conventional methods by comprehensively analyzing the correlations of each object within the video.

[0066] Referring again to FIG. 1, in step S11, the processor can select a recommended action information corresponding to each situation analysis information among the plurality of recommended action information and set a risk level for each situation analysis information according to a preset standard, thereby generating a plurality of integrated information including situation analysis information, recommended action information, and risk level.

[0067] According to one embodiment of the present specification, the storage medium may pre-store information on multiple recommended actions according to multiple situation types. The information on recommended actions may include details that control personnel or relevant agencies must take depending on each situation type. For example, regarding a traffic accident situation type, the information on recommended actions may include reporting to emergency agencies, details of actions to be taken at the scene, and contact information for relevant agencies, such as "immediately dispatch to the scene to suppress the situation and rescue injured persons."

[0068] The processor can select recommended action information corresponding to the situation analysis information according to the situation type of the situation analysis information.

[0069] In addition, the processor can set a risk level for each situation analysis information according to preset criteria.

[0070] The above risk level may refer to information classifying the level of risk regarding the situation within the video. For example, the above risk level may be assigned from level 1 to level 10. Levels 1 and 2 may refer to normal conditions with no unusual circumstances. Levels 3 through 5 may refer to conditions requiring careful observation. Levels 6 and 7 may correspond to situations requiring immediate verification by control personnel. Levels 8 through 10 may refer to situations requiring immediate verification and rapid response by control personnel. The examples for classifying the above risk levels are examples only and are not limited thereto.

[0071] The above processor can set the risk level according to preset criteria.

[0072] For example, the above-mentioned preset criteria may correspond to a situation type. The processor may set a higher risk level when the situation type is a traffic accident than when it is illegal parking. This is an example and is not limited thereto.

[0073] As another example, the above-mentioned preset criteria may correspond to specific keywords. When the above-mentioned situation type is a traffic accident, the processor may set a higher risk level when the keyword "occurrence of fatalities and serious injuries" is present than when the keyword "connection accident" is present. This is an example and is not limited thereto.

[0074] As another example, the processor can set the risk level by using both the situation type and specific keywords.

[0075] Subsequently, the processor may generate multiple integrated information including situation analysis information, recommended action information corresponding to the situation analysis information, and the risk level of the situation analysis information.

[0076] Figure 4 is an example of integrated information.

[0077] In the example of FIG. 4, the processor may input a video of a traffic accident situation to the artificial neural network model (110). The artificial neural network model (110) may generate situation analysis information including at least one of the situation type, the situation information, and an object related to the situation information. The processor may select recommended action information corresponding to the situation analysis information and set a risk level. Subsequently, the processor may generate integrated information (200) including the situation analysis information, recommended action information, and risk level.

[0078] In addition, the processor may generate the integrated information (200) by further including identification information of the video recording device that captured the video.

[0079] The processor can obtain situation analysis information for each video from the artificial neural network model (110) according to a preset time interval. For example, the processor can obtain the situation analysis information from the artificial neural network model (110) every 30 seconds. The time is an example and is not limited thereto.

[0080] According to one embodiment of the present specification, the artificial neural network model (110) can be trained using a video of a preset length as training data.

[0081] The above processor can train the artificial neural network model (110) using multiple videos of a preset length as training data.

[0082] The processor can train the artificial neural network model (110) by segmenting the original video into multiple segment videos. For example, the processor can generate multiple segment videos by segmenting the original video every 30 seconds. The processor can train the artificial neural network model (110) using the multiple segment videos. The length of the segment videos is an example and is not limited.

[0083] At this time, the artificial neural network model (110) can be trained by further utilizing situation analysis information for the previous video segment as training data when generating situation analysis information for the target video segment.

[0084] When the processor generates situation analysis information for the target segment video, it can further use situation analysis information for the previous segment video as training data to train the artificial neural network model (110).

[0085] For example, the processor can generate a first segment video divided from 0 seconds to 30 seconds of the original video and a second segment video divided from 30 seconds to 60 seconds. When the target segment video is the second segment video, the processor can train the artificial neural network model (110) by using the situation analysis information for the first segment video as additional training data.

[0086] The processor may store situation analysis information regarding the previous video segment in a storage medium. Subsequently, the processor may train the artificial neural network model (110) to generate situation analysis information regarding the target analysis video by further utilizing the situation analysis information regarding the previous video segment stored in the storage medium as training data. Through this, the artificial neural network model (110) may be trained to generate situation analysis information regarding the current situation by analyzing the preceding situation together with the video received in real time from the video capturing device. Since this is a technology widely known among those skilled in the art, a detailed explanation is omitted.

[0087] In the above step S10, the processor may input the received video to the artificial neural network model whenever the length of the video received from each video capturing device reaches a preset length. At this time, the processor may further input situation analysis information of the previous video segment stored in the storage medium to the artificial neural network model.

[0088] For example, the processor may input the received video to the artificial neural network model when the length of the received video reaches 30 seconds. Subsequently, when the length of the received video reaches 60 seconds, the processor may input the video corresponding to the interval from 30 seconds to 60 seconds to the artificial neural network model. At this time, the processor may further input situation analysis information regarding the video in the first 30-second interval to the artificial neural network model. The video length is an example and is not limited thereto.

[0089] Through this, the processor can acquire situation analysis information regarding the video received from the plurality of video capturing devices according to preset time intervals.

[0090] In the above step S11, the processor can generate integrated information according to a preset time interval using situation analysis information obtained according to a preset time interval.

[0091] According to one embodiment of the present specification, the artificial neural network model (110) may be trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to each situation in an input video, using training data that includes at least one of a situation type, situation information, and an object related to each situation, when at least two of a plurality of situation types are included in at least one of a plurality of pre-set situation types in at least one of the plurality of videos.

[0092] When at least two of a plurality of preset situation types are included in at least one of the plurality of videos, the processor can train the artificial neural network model (110) to generate situation analysis information including a situation type, situation information, and at least one of an object related to each situation in an input video using training data that includes a situation type, situation information, and at least one of an object related to each situation according to each situation included in the video.

[0093] For example, a single video may capture both a traffic accident situation and a fire situation. The processor may generate a first data set containing at least one of a situation type, situation information, and an object associated with each situation regarding the traffic accident situation in the video. Additionally, the processor may generate a second data set containing at least one of a situation type, situation information, and an object associated with each situation regarding the fire situation in the video. The processor may train the artificial neural network model (110) using the video, the first data set, and the second data set.

[0094] In step S10 above, if at least one of the plurality of videos input to the artificial neural network model (110) contains at least two of the plurality of preset situation types, the processor can obtain situation analysis information for each situation type in the video.

[0095] In the above step S11, the processor can generate integrated information based on each situation analysis information in the corresponding video.

[0096] Referring again to FIGS. 1 and FIGS. 2, in step S12, the processor can display the plurality of integrated information on a display device (120) according to a preset priority.

[0097] For example, the processor may display the plurality of integrated information in the form of a list (130) according to a preset priority. The processor may display the information with higher priority at the top of the list. The form of the list is an example and is not limited thereto.

[0098] According to one embodiment of the present specification, in step S11, the processor may preferentially display integrated information with a relatively high risk level. As described above, the risk level may be set to any one of levels 1 to 10. The processor may display integrated information in a list in order from level 10 to level 1.

[0099] At this time, in step S12, the processor may display the integrated information satisfying a preset risk level condition on the display device. For example, the processor may display the integrated information corresponding to a risk level of 8 or higher on the display device. The processor may not display the integrated information corresponding to a risk level of 8 or lower on the display device. The risk level condition is an example and is not limited thereto.

[0100] According to one embodiment of the present specification, in step S12, if there are multiple pieces of integrated information with the same risk level, the processor may display the multiple pieces of integrated information according to a situation type that is pre-set to be displayed preferentially.

[0101] For example, among the plurality of situation types mentioned above, it may be configured to display them preferentially in the order of traffic accident, fire, assault, and loitering. The processor may display integrated information with a situation type of traffic accident as the highest priority among the plurality of integrated information. The processor may display integrated information with a situation type of traffic accident at the very top of the list. The processor may display integrated information with a situation type of traffic accident, fire, assault, and loitering in the order of being positioned at the top of the list. The situation types configured to be displayed preferentially are examples and are not limited thereto.

[0102] According to one embodiment of the present specification, in step S10, the processor may further generate time information for acquiring each situation analysis information. The processor may generate the time information by reading time from a Real Time Clock (RTC).

[0103] In the above step S11, the processor may generate integrated information by further including time information at which each situation analysis information was acquired.

[0104] In step S12 above, if there are multiple integrated information files with the same risk level, the processor may display the most recently created integrated information first. The processor may display the most recently created integrated information at the top of the list.

[0105] The most recently generated integrated information may refer to the integrated information for which the time at which the aforementioned situation analysis information was acquired is closest to the current time. In other words, the most recently generated integrated information may refer to the most recently generated integrated information.

[0106] The above processor can display in a list the integrated information in the order of situation analysis information that includes the acquired time closest to the current time.

[0107] According to one embodiment of the present specification, in step S12, if there are multiple integrated informations having the same risk level and situation type, the processor may preferentially display the most recently created integrated information.

[0108] Figure 5 is another example of integrated information displayed on a display device.

[0109] Referring to FIG. 5, in step S12, the processor may further display location information where a video recording device corresponding to each integrated information is installed. Each integrated information and the video recording device corresponding to each integrated information may refer to a video recording device that has captured a video for each integrated information.

[0110] The storage medium may further store identification information for each image capturing device and location information where each image capturing device is installed. The processor may use the received identification information to extract the location information where each image capturing device is installed and display it on the display device.

[0111] In addition, the processor can further display the shortest distance between the location where each video recording device is installed and related institutions such as nearby fire stations, police stations, and hospitals.

[0112] The processor can display each of the integrated information and the corresponding image capturing devices (101 to 103) on a map. Additionally, the processor can match each of the integrated information and the corresponding image capturing device and display them on the display device.

[0113] According to one embodiment of the present specification, in step S12, the processor may further output an alarm output signal to the speaker device according to a preset risk level condition.

[0114] For example, the processor may output an additional alarm output signal to the speaker device when the risk level is 8 or higher. The risk level condition is an example and is not limited thereto.

[0115] According to one embodiment of the present specification, in step S12, the processor may display at least one piece of integrated information that is not classified as a false positive according to a preset criterion among a plurality of integrated information on the display device.

[0116] The processor can determine whether each of the integrated information is a false detection based on preset criteria.

[0117] In this specification, false detection may refer to the artificial neural network model generating situation analysis information that differs from the actual situation due to environmental factors within the video, changes in lighting, weather conditions, etc. For example, if a traffic accident occurs but the situation analysis information is generated as simple parking, this may constitute a false detection.

[0118] The processor can determine whether it is a false detection by verifying the semantic consistency of the integrated information generated from each video according to preset criteria.

[0119] As described above, the processor can generate integrated information by acquiring situation analysis information for each video according to a preset time interval.

[0120] According to one embodiment of the present specification, the processor may classify the first integrated information as a false detection if the similarity between the first integrated information generated from any one video and the second integrated information generated earlier than the first integrated information is less than a preset threshold value.

[0121] The above second integrated information may be integrated information created immediately before or earlier than the above first integrated information.

[0122] The second integrated information may include text such as "occurrence of a traffic accident," and the first integrated information may include text such as "road control due to the occurrence of a traffic accident." In this case, the semantic similarity between the first integrated information and the second integrated information may be greater than or equal to a preset threshold.

[0123] The second integrated information may include text such as "occurrence of a traffic accident," and the first integrated information may include text such as "traffic congestion due to illegal parking." In this case, the semantic similarity between the first integrated information and the second integrated information may be less than a preset threshold.

[0124] The processor can map the first integrated information and the second integrated information into a high-dimensional vector space using algorithms such as Word2Vec or SbERT, and calculate similarity by calculating the cosine similarity, Euclidean distance, etc., between the two vectors. Since this is a technique widely known among those skilled in the art, a detailed explanation is omitted.

[0125] The processor can store integrated information classified as false detections in the storage medium. The integrated information classified as false detections can be used as retraining data (feedback loop) for autonomous performance improvement of the artificial neural network model.

[0126] FIG. 6 is a flowchart of a method for generating image situation information according to another embodiment of the present specification.

[0127] Referring to FIG. 6, steps S20 to S22 are identical to steps S10 to S12, so a repetitive description is omitted.

[0128] In step S23, when integrated information having a preset risk level is repeatedly generated for a preset time, the processor can progressively increase the risk level for said integrated information.

[0129] For example, the processor may gradually increase the risk level of the maritime situation analysis information when the situation analysis information with a risk level of 8 or lower is repeatedly generated for a preset period of time. In this case, if the situation analysis information with a risk level of 5 or lower is repeatedly generated for 6 hours or more, the processor may increase the risk level of the situation analysis information by one level. The conditions regarding the risk level and the conditions regarding time are examples and are not limited thereto.

[0130] Video footage may be received from a video recording device showing that the driver has not gotten out of the vehicle after the vehicle has been parked. The situation in which the vehicle is parked may correspond to the first stage. If the driver does not get out of the vehicle for 6 hours, the processor may raise the risk level for the relevant integrated information to the second stage. Subsequently, if the driver does not get out of the vehicle for 6 hours, the processor may raise the risk level for the relevant integrated information to the third stage.

[0131] A situation where a vehicle is parked may be considered a normal situation. However, if the driver does not get out of the vehicle for an extended period, it may constitute a situation where an unusual circumstance has occurred to the driver.

[0132] According to one embodiment of the present specification, in step S23, when the risk level of the integrated information increases, the processor may further change the situation type for the integrated information.

[0133] For example, a situation in which the vehicle is parked may correspond to the normal situation. A situation in which the driver does not get out of the vehicle for a long period of time may correspond to any one of the abnormal situations. The processor may further change the situation type for the integrated information depending on the situation within the video.

[0134] Through this, the method for generating video situation information according to the present specification can update integrated information by analyzing the context of the preceding and succeeding situations within a video received from a video capturing device.

[0135] According to one embodiment of the present specification, in step S23, when the risk level for the integrated information becomes greater than or equal to a preset risk level, the processor may further display the integrated information on the display device.

[0136] For example, when the risk level of the integrated information is 8 or higher, the processor may further display the integrated information on the display device. The risk level condition is an example and is not limited thereto.

[0137] At this time, the processor may further display location information where an image capturing device corresponding to the integrated information with a higher risk level is installed.

[0138] According to one embodiment of the present specification, in step S23, when the risk level of the integrated information becomes greater than or equal to a preset risk level, the processor may further output an alarm output signal to the speaker device.

[0139] A non-volatile computer-readable storage medium according to the present specification can be used as a component of a screening control system that displays multiple integrated information generated from video images received from multiple video capturing devices installed at preset locations on a display device.

[0140] Although embodiments of this specification have been described above with reference to the attached drawings, those skilled in the art to which this specification pertains will understand that the present invention may be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.

Claims

Claim 1 A non-volatile computer-readable storage medium storing recommended action information according to a plurality of situation types, wherein the computer-executable instructions are written to perform each step of a video situation information generation method when executed by one or more processors, and the computer-executable instructions include: (a) a step of inputting a plurality of videos received from a plurality of video capturing devices installed at a preset location into an artificial neural network model to obtain situation analysis information for each video according to a preset time interval; (b) a step of selecting recommended action information corresponding to each situation analysis information among the plurality of recommended action information, setting a risk level for each situation analysis information according to a preset criterion, and generating a plurality of integrated information including situation analysis information, recommended action information, and risk level; and (c) a step of displaying the plurality of integrated information on a display device according to a preset priority; wherein step (c) includes: (c-1) a step of classifying the plurality of integrated information as false detection according to a preset criterion; and (c-2) a step of displaying at least one integrated information that is not classified as a false detection on a display device according to the preset priority; wherein, in step (c-1), if the similarity between the first integrated information generated from any one video and the second integrated information generated prior to the first integrated information is less than a preset threshold value, the first integrated information is classified as a false detection; and the artificial neural network model is a non-volatile computer-readable storage medium trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to the situation information for an input video using training data including at least one of a plurality of videos, a situation type for each video, situation information, and an object related to each situation information. Claim 2 A non-volatile computer-readable storage medium according to claim 1, wherein the artificial neural network model is trained to generate situation analysis information including at least one of a situation type, situation information, and an object related to each situation in an input video using training data including at least one of a situation type, situation information, and an object related to each situation according to each situation included in at least one of a plurality of situations included in at least one of a plurality of situations included in at least one of a plurality of situations included in at least one of a plurality of situations included in at least one of a plurality of situations included in at least one of a plurality of situations included in at least one of a plurality of situations included in the input video, and wherein, in step (a), the artificial neural network model acquires situation analysis information for each situation type in the video. Claim 3 A non-volatile computer-readable storage medium according to claim 1, characterized in that, in step (c-2), the integrated information having a relatively high risk level is preferentially displayed. Claim 4 A non-volatile computer-readable storage medium according to claim 3, wherein, in step (c-2), if there are multiple integrated information pieces of the same risk level, the plurality of integrated information pieces are displayed according to a situation type that is pre-set to be displayed preferentially. Claim 5 A non-volatile computer-readable storage medium according to claim 3, characterized in that, in step (a), time information for acquiring each situation analysis information is further generated, and in step (c-2), if there are multiple integrated informations with the same risk level, the most recently generated integrated information is prioritized for display. Claim 6 A non-volatile computer-readable storage medium according to claim 4, characterized in that, in step (a), time information for acquiring each situation analysis information is further generated, and in step (c-2), if there are multiple integrated informations with the same risk level and situation type, the most recently generated integrated information is displayed preferentially. Claim 7 A non-volatile computer-readable storage medium according to claim 1, characterized in that, in step (c-2), location information where a corresponding image capturing device is installed is further displayed for each integrated information. Claim 8 A non-volatile computer-readable storage medium according to claim 1, characterized in that, in step (c-2), an alarm output signal is further output to a speaker device according to a preset risk level condition. Claim 9 delete Claim 10 delete Claim 11 A non-volatile computer-readable storage medium according to claim 1, wherein the computer-executable instruction further comprises: (d) a risk level adjustment step of progressively increasing the risk level of said integrated information when integrated information having a preset risk level is repeatedly generated for a preset time. Claim 12 A non-volatile computer-readable storage medium according to claim 11, wherein the computer-executable instruction further comprises: (e) changing the situation type for said integrated information when the risk level for said integrated information increases. Claim 13 A non-volatile computer-readable storage medium according to claim 11, wherein the computer-executable instruction further comprises: (f) displaying the integrated information to the display device when the risk level for the integrated information becomes greater than or equal to a preset risk level. Claim 14 A non-volatile computer-readable storage medium according to claim 11, wherein the computer-executable instruction further comprises: (g) outputting an alarm output signal to a speaker device when the risk level for the integrated information becomes greater than or equal to a preset risk level. Claim 15 A screening control system comprising: a non-volatile computer-readable storage medium according to any one of claims 1 to 8 and 11 to 14; a plurality of image capturing devices installed at preset locations; and a display device on which the plurality of integrated information is displayed.

Citation Information

Patent Citations

  • Method for managing dangerous article transport vehicle and apparatus for managing dangerous article transport vehicle by method

    KR1020140015045A

  • Artificial intelligence-based safe privacy zone management server and method and privacy detection device

    KR1020250061975A

  • Workplace risk assessment system based on AI video analysis

    KR1020250070721A

  • Method and Apparatus for controlling Situation Room of Traffic Information Center in Specific Situations

    KR102593845B1

  • Method and system for prioritizing displays of surveillance system

    WO2013085377A1