System for de-identified textualization and dynamic prioritization of public safety events via semantic analysis of multi-source video data

KR103025488B1Active Publication Date: 2026-09-29주식회사 드제이
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020260091052
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-09-29
Estimated Expiration
2046-05-20

Smart Images

  • Figure 112026060972882-PAT00001_ABST
    Figure 112026060972882-PAT00001_ABST
Patent Text Reader

Abstract

A system for de-identifying textualizing public safety events and determining dynamic priorities through semantic analysis of multi-source video data is disclosed. According to the present invention, by integrating and recognizing individual events within multi-source video collected independently from multiple cameras and different locations into a single organic event unit based on spatiotemporal correlations, complex public safety situations occurring in a wide area can be identified quickly and accurately without distortion. This significantly reduces misjudgments and situational misinterpretations that may occur from the analysis of fragmentary metadata from individual cameras, and by clearly identifying causal relationships between multiple signs of danger, it has the effect of greatly improving the ability to disseminate information and respond to complex situations requiring emergency rescue or large-scale disasters.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The following disclosure relates to a system for the anonymization of public safety events and dynamic priority determination through semantic analysis of multi-source video data. Background Technology

[0003] Recently, to enhance public safety and security, the deployment of multi-source video collection infrastructure, including security CCTVs, is rapidly expanding in various physical spaces such as train stations, intersections, hospitals, and schools. Consequently, AI-based Intelligent Video Analytics technology is being widely utilized to efficiently monitor the vast amount of collected video data, and research is actively underway to generate metadata within videos from multiple channels using deep learning models such as object detection, action recognition, and object tracking. In particular, semantic event recognition technology, which aims to extract temporal, spatial, and object information from individual video streams acquired in a distributed manner in a multi-camera environment and derive correlations between events based on this information to understand the situation, is establishing itself as a core technology for intelligent control systems.

[0004] In addition, due to the advancement of Vision-Language Models (VLMs) that combine vision recognition and natural language processing technologies, technologies are being introduced to automatically generate text descriptions in natural language that allow humans to intuitively understand complex situations within video data. However, when the text describing a dangerous situation generated by a control system includes a combination of a specific person's physical characteristics, clothing, and detailed time and location information, this can become a new pathway for privacy infringement capable of re-identifying specific individuals, just as effectively as the raw video. Therefore, in the fields of data security and privacy protection, research is being conducted from various angles on natural language-based de-identification and filtering technologies that go beyond masking or blurring raw video to identify sensitive information within the generated semantic text and remove identifying elements while maintaining the utility of the data for control.

[0005] Meanwhile, in environments where a large volume of hazardous events are simultaneously detected by multi-channel cameras within an extensive control area, visual fatigue and cognitive load on control personnel increase sharply, making rapid response to emergencies difficult. To address this, there are ongoing attempts to maximize system efficiency by applying dynamic priority scheduling algorithms, previously utilized in data communication and network traffic control, to the fields of disaster safety and public control. Specifically, active technical discussions are continuing regarding dynamic event management technology that readjusts processing priorities in real time by comprehensively calculating variable environmental factors—such as the risk type of the event itself, spatial characteristics of the occurrence location, density by time period, and predicted risk—along with the reliability of the detection model for multiple detected events. The problem to be solved

[0007] The problem that this invention aims to solve is to provide an integrated semantic event recognition technology capable of precisely analyzing the spatiotemporal correlations between multiple source video streams acquired from multiple cameras and various different environments, thereby grouping dispersed individual events into a single organic event unit for consistent identification. In other words, by integrating and recognizing multiple danger signs (e.g., falls, crowding) detected independently at different locations as one interconnected, large-scale public safety situation, this invention seeks to overcome the limitations of existing control systems, which were limited to the level of individual metadata analysis, and to minimize the possibility of situational misinterpretation.

[0008] In addition, we aim to provide a technical solution that fundamentally blocks new forms of personal information leakage and re-identification risks that may occur during the process of converting dangerous situations within video into intuitive natural language descriptive text using a Vision-Language Model (VLM). To fill the gaps in existing technologies that were skewed toward structural de-identification processing, such as masking raw data, we aim to thoroughly protect personal information while preserving the minimum amount of information essential for monitoring by implementing a text conversion technology based on secondary de-identification, filtering, and summarization rules that detects sensitive identification elements in real time—such as combinations of a person's appearance, clothing, and detailed time and place—within generated natural language sentences and automatically refines them.

[0009] Finally, we aim to provide a dynamic priority algorithm that can quantitatively resolve the issues of severe control fatigue and cognitive overload experienced by control personnel due to numerous public safety events detected simultaneously within a wide control area. By calculating variable environmental factors in real time—such as the spatial characteristics of the location of occurrence, population density by time of day, and the probability of false positives (reliability) of the detection model, as well as the inherent risk type of the event—and dynamically readjusting the scheduling, we intend to maximize the operational efficiency of the system and the speed of emergency response by summarizing and selectively providing only the highest risk situations to control personnel. means of solving the problem

[0011] A system for de-identifying text and dynamically prioritizing public safety events through semantic analysis of multi-source video data according to one embodiment disclosed in this document comprises: a video collection unit that receives video data from a plurality of cameras placed in different locations and adds source metadata including a camera identifier, a location identifier, and a time of capture to each of the video data; an event detection unit that detects objects and events from the video data and classifies the detected events into one of the event types among falling, driving in reverse, unauthorized entry, abandoned objects, crowd congestion, and dangerous behavior, and generates a unit event including the event type, the source metadata, and detection reliability; an integrated event recognition unit that recognizes two or more unit events that are spatiotemporally related among the plurality of unit events by grouping them into a single integrated semantic event; and a de-identifying text generation unit that generates de-identifying situational text for the integrated semantic event by replacing an identifying expression that enables the re-identification of an individual with a de-identifying expression. It may include a priority determination unit that calculates the dynamic priority of the integrated semantic event based on the urgency, public risk, and detection reliability of the integrated semantic event, and readjusts the dynamic priority whenever a new unit event is generated; and a control provision unit that provides control information, including the non-identifiable situation text, the urgency, and recommended measures, to a control terminal for the top N integrated semantic events in order of high dynamic priority.

[0012] According to one embodiment, when one of the plurality of unit events is designated as a first unit event and the other as a second unit event, the integrated event recognition unit determines the first unit event and the second unit event as a temporal candidate pair if the difference between the time of shooting of the first unit event and the time of shooting of the second unit event is within a preset time window, and maintains the first unit event and the second unit event as separate integrated semantic events if the difference exceeds the time window, and determines the temporal candidate pair as a spatiotemporal candidate pair if the location identifier of the first unit event and the location identifier of the second unit event forming the temporal candidate pair belong to a predefined same control area, and groups the spatiotemporal candidate pair into one integrated semantic event if the combination of the event type of the first unit event and the event type of the second unit event forming the spatiotemporal candidate pair corresponds to a predefined simultaneous occurrence rule, and maintains the first unit event and the second unit event as separate integrated semantic events if they do not correspond to the simultaneous occurrence rule.

[0013] According to one embodiment, the integrated event recognition unit calculates a spatiotemporal correlation for the spatiotemporal candidate pairs, wherein the value is larger when the difference between the shooting times of the first unit event and the second unit event is small, and the value is larger when the degree of overlap between the field of view of the camera corresponding to the camera identifier of the first unit event and the camera corresponding to the camera identifier of the second unit event is large; and, only when the calculated spatiotemporal correlation is greater than or equal to a reference value, the spatiotemporal candidate pairs are grouped into one integrated semantic event, wherein if two or more unit events are included in one integrated semantic event, the unit event that maximizes the product of the detection reliability and the risk weight pre-assigned for each event type among the included unit events is selected as the representative unit event, the event type of the representative unit event is set as the representative type of the integrated semantic event, and the public risk of the integrated semantic event can be adjusted to be higher as the number of unit events included in the integrated semantic event increases.

[0014] According to one embodiment, the non-identifying text generation unit generates a primary situation text describing a situation using a vision-language model from the integrated semantic event, decomposes the primary situation text into an appearance description element indicating the appearance of a person, a personal description element indicating the demographic attributes of a person, a location description element indicating the location of occurrence, a time description element indicating the time of occurrence, and a situation description element indicating the type of situation, classifies the appearance description element and the personal description element as identification attributes that enable re-identification of an individual, classifies the situation description element as a control-essential attribute that is not an identification attribute, performs a substitution to generalize a first identification attribute among the description elements classified as identification attributes that has a re-identification contribution less than a preset threshold to a predefined super-concept term, performs a substitution to delete the corresponding description element among the description elements classified as identification attributes that has a re-identification contribution greater than or equal to the threshold, maintains the original text for the description elements classified as control-essential attributes, and outputs the result of recombining the substituted or maintained description elements as the non-identifying situation text.

[0015] According to one embodiment, the non-identification text generation unit maps the location description element and the time description element to spatiotemporal buckets corresponding to the place identifier and the time of shooting, calculates a re-identification risk having a larger value as the number of estimated users belonging to the spatiotemporal bucket decreases and a larger value as the number of identification attributes remaining in the non-identification situation text increases, and if the calculated re-identification risk exceeds a reference risk, performs additional generalization to generalize the remaining identification attributes, the location description element, and the time description element by one level from a low generalization level to a high generalization level, and then recalculates the re-identification risk, repeating this process until the re-identification risk becomes less than or equal to the reference risk, and if the re-identification risk does not exceed the reference risk, does not perform the additional generalization, and if there are multiple generalization results in which the re-identification risk is less than or equal to the reference risk, selects the generalization result that maximizes the amount of information preserved from the control essential attribute among them and outputs it as the non-identification situation text.

[0016] According to one embodiment, the priority determination unit assigns a first location risk coefficient when the location identifier of the integrated semantic event corresponds to any one of a predefined group of high-risk locations including a railway, a roadway, an emergency room, and a school entrance, and assigns a second location risk coefficient smaller than the first location risk coefficient when it does not correspond to any of the high-risk location groups; assigns a first time zone weight when the time of capture of the integrated semantic event falls within a predefined congestion time zone, and assigns a second time zone weight smaller than the first time zone weight when it does not fall within the congestion time zone; calculates a density correction coefficient having a larger value as the density of the area where the integrated semantic event occurred increases, calculates a predicted risk having a larger value as the deterioration probability increases based on a pre-learned deterioration probability for the representative type of the integrated semantic event, calculates a false positive probability having a larger value as the detection reliability decreases, and calculates a reliability correction coefficient having a smaller value as the false positive probability increases, and the urgency, the public risk, the first location risk coefficient, or the second location The dynamic priority is calculated by multiplying the reliability correction factor by the weighted sum of the risk factor, the first time zone weight or the second time zone weight, the density correction factor, and the predicted risk; when a new unit event is incorporated into any one of the integrated semantic events, the dynamic priority of the said integrated semantic event is readjusted upward; and when no new unit event is incorporated into any one of the integrated semantic events during a preset maintenance time, the dynamic priority of the said integrated semantic event is readjusted downward by applying a decay function according to the passage of time; the number of control information items provided per unit time to the control terminal is defined as the control fatigue index; and the N of the top N items is determined variably so that the control fatigue index does not exceed a preset allowable load,If the above control fatigue index exceeds the above allowable load, N is reduced so that the integrated semantic event with high dynamic priority is provided first, and if the above control fatigue index does not exceed the above allowable load, N can be increased within the above allowable load range. Effects of the invention

[0018] According to the present invention, by integrating and recognizing individual events within multi-source images collected independently from multiple cameras and different locations into a single organic event unit based on spatiotemporal correlations, complex public safety situations occurring in a wide area can be identified quickly and accurately without distortion. This significantly reduces misjudgments and situational misinterpretations that may occur from the analysis of fragmentary metadata from individual cameras, and by clearly identifying causal relationships between multiple risk signs, it has the effect of greatly improving the ability to disseminate information and respond to complex situations requiring emergency rescue or large-scale disasters.

[0019] In addition, by applying secondary de-identification, filtering, and summarization rules that filter and refine combinations of personal appearance, clothing, and specific spatiotemporal information in real time to text describing dangerous situations generated by a Vision-Language Model (VLM), it is possible to fundamentally block new forms of personal information leakage and re-identification risks resulting from the introduction of video captioning technology. This precisely preserves only the core situational information essential for on-site control and dissemination to related agencies while excluding the possibility of infringing on the privacy of specific individuals, thereby satisfying strict personal information protection regulations and simultaneously ensuring the usability and safety of public data.

[0020] Finally, by dynamically readjusting processing priorities through real-time calculations of risk types, spatial risk levels, population density by time period, and detection reliability for multiple detected events, it is possible to select critical situations requiring top priority response in real time from among the flood of monitoring data. By providing monitoring personnel with a selection and summary of only the highest-risk situations, the accumulated fatigue and cognitive overload resulting from a 24-hour monitoring environment can be quantitatively reduced. Furthermore, by concentrating limited monitoring personnel and resources on emergency situations, it provides the effect of maximizing the speed of initial response for preventing public safety accidents. Brief explanation of the drawing

[0022] FIGS. 1 and FIGS. 2 are schematic block diagrams of a system according to one embodiment. FIG. 3 is a block diagram illustrating the operation of an integrated event recognition unit according to one embodiment. FIG. 4 is a flowchart illustrating the operation of an integrated event recognition unit according to one embodiment. FIG. 5 is a block diagram illustrating the operation of a non-identifiable text generation unit according to one embodiment. FIG. 6 is a flowchart illustrating the operation of a non-identifiable text generation unit according to one embodiment. FIG. 7 is a block diagram illustrating the operation of a priority determination unit according to one embodiment. FIG. 8 is a schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented. Specific details for implementing the invention

[0023] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, actual implementations are not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or substitutions included in the technical concept described by the embodiments.

[0024] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.

[0025] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or joined to that other component, or that there may be other components in between.

[0026] Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” may each include any one of the items listed together with the corresponding phrase, or all possible combinations thereof. In this specification, terms such as “comprising” or “having” are intended to designate the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0027] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0028] As used herein, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0029] As used in this document, the term "part" refers to software or hardware components, such as FPGAs or ASICs, and the "part" performs certain roles. However, the meaning of "part" is not limited to software or hardware. The "part" may be configured to reside in an addressable storage medium or configured to operate one or more processors. For example, the "part" may include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and "parts" may be combined into a smaller number of components and "parts" or further separated into additional components and "parts." Furthermore, the components and "parts" may be implemented to operate one or more CPUs within a device or secure multimedia card. Additionally, '~part' may include one or more processors.

[0030] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.

[0032] FIGS. 1 and FIGS. 2 are schematic block diagrams of a system according to one embodiment.

[0033] Referring to FIGS. 1 and FIGS. 2, the system (100) may be a system that de-identifies public safety events into text and determines dynamic priorities through semantic analysis of multi-source image data.

[0034] Here, "multi-source video data" may refer to a set of video data that is independently captured by multiple cameras placed in different locations, such as stations, terminals, intersections, hospitals, and schools, and input into the system (100). For example, if images captured by camera C1 installed on the platform of subway station A, camera C2 installed in the waiting room of the same station A, and camera C3 installed at a nearby intersection are each collected into a single system (100), the set of these images may be multi-source video data.

[0035] "Semantic analysis" refers to an analysis that extracts the meaning of objects, such as people, vehicles, and physical things, as well as the actions or situations they are in, from the pixel information of an image. In other words, it is an analysis that interprets "what is currently happening" rather than simply "how the pixels have changed."

[0036] "Public safety event" refers to an event that may affect the safety of an unspecified number of people, and may include falling, driving against traffic, trespassing, abandoned objects, crowd congestion, and dangerous behavior as described below.

[0037] "De-identification text" may refer to a process of describing a situation perceived from an image in human-readable natural language, but processing it so that the identity of a specific individual cannot be re-identified from the natural language description alone.

[0038] "Dynamic priority determination" may refer to a process in which the processing order of multiple already detected events is not assigned as a fixed value, but rather the order is variably updated according to the passage of time and new inputs.

[0039] As illustrated in FIG. 2, the system (100) may include an image collection unit (110), an event detection unit (120), an integrated event recognition unit (130), a non-identifiable text generation unit (140), a priority determination unit (150), and a control provision unit (160), and each unit may be integrated into one physical device or distributed across multiple devices connected by a network.

[0040] The image acquisition unit (110) receives image data from a plurality of cameras placed in different locations and can add source metadata including a camera identifier, a location identifier, and a time of shooting to each of the image data.

[0041] "Multiple cameras" may refer to two or more shooting devices that provide images to the system (100), and "different places" may refer to physically separated points or zones. "Image data" may refer to a set of consecutive frames or a stream thereof that are captured and transmitted by the camera.

[0042] "Camera identifier" may refer to a value that uniquely distinguishes an individual camera (e.g., C1, C2), "place identifier" may refer to a value that distinguishes the place or jurisdiction where the camera is installed (e.g., Station A-Platform, Station A-Waiting Room), and "time of capture" may refer to a value indicating the time when the video data was captured (e.g., 2026-05-18 08:12:31). "Source metadata" may refer to a combination of the camera identifier, place identifier, and time of capture indicating which camera, at which place, and when the video data was captured.

[0043] For example, the video collection unit (110) can add {camera identifier=C1, location identifier=Station A-Platform, shooting time=08:12:31} to a frame received from camera C1, and add {camera identifier=C2, location identifier=Station A-Waiting Room, shooting time=08:12:33} to a frame received from camera C2. By adding source metadata to the video data, the video collection unit (110) can track which event occurred when and where in a subsequent step.

[0044] The event detection unit (120) detects objects and events from the image data and classifies the detected events into one of the event types, such as falling, driving in reverse, unauthorized entry, abandoned objects, crowd density, and dangerous behavior, and can generate a unit event including the event type, the source metadata, and the detection reliability.

[0045] "Object" may refer to entities such as people, vehicles, or objects detected within an image, and "Event" may refer to a meaningful occurrence generated by one or more objects. "Event Type" may refer to the category to which a detected event belongs, and each type can be defined as follows.

[0046] "Falling" refers to a state in which a person falls to the ground from a standing or walking position and maintains an abnormal posture for a certain period of time or longer; "reverse movement" refers to a state in which a vehicle or person moves in a direction opposite to the designated direction of travel; "unauthorized entry" refers to the act of entering a restricted area without authorization; "abandoned object" refers to an object that remains in a specific location for a certain period of time or longer without a manager; "crowding" refers to a state in which the number of people per unit area exceeds a standard value; and "dangerous behavior" refers to actions that may cause harm to oneself or others, such as assault, fighting, or approaching the tracks.

[0047] "Detection confidence" may refer to a value (e.g., a probability value between 0 and 1) indicating the degree of certainty that the detection and classification result corresponds to an actual event. "Unit event" may refer to a minimum processing unit in which the event type, the source metadata, and the detection confidence are combined for a single event detected by a single camera.

[0048] For example, if the event detection unit (120) detects a scene in the video of camera C1 where a person suddenly falls to the ground while walking and then does not move, it can classify this as "falling down" and generate unit event E1 = {Type=Falling down, Source metadata=(C1, Station A-Platform, 08:12:31), Detection reliability=0.92}. In addition, if it detects a scene in the video of camera C2 at the same time where a large number of people quickly gather in a narrow area, it can classify this as "crowd density" and generate unit event E2 = {Type=Crowd density, Source metadata=(C2, Station A-Waiting Room, 08:12:33), Detection reliability=0.81}.

[0049] The integrated event recognition unit (130) can recognize two or more unit events that are spatially and temporally related among the plurality of unit events by grouping them into one integrated semantic event.

[0050] "Spatiotemporally related" may mean that two or more unit events occur in close proximity in time and in the same or adjacent jurisdictions in space, and are in a relationship where they can be considered to have originated from a single event. "Integrated semantic event" may refer to a higher-level concept of an event expressed by grouping multiple unit events detected individually by different cameras into a "single event unit."

[0051] For example, the integrated event recognition unit (130) recognizes that the previously generated unit events E1 (falling, Station A, 08:12:31) and E2 (crowd density, Station A, 08:12:33) occurred at almost the same Station A, and can group them into a single integrated semantic event U1, which is "a fall that occurred at Station A platform and the resulting crowd density and rescue need situation." By grouping the unit events into an integrated semantic event, the integrated event recognition unit (130) can prevent the same event from being duplicated as multiple individual notifications.

[0052] The non-identification text generation unit (140) can generate non-identification situation text for the integrated semantic event by replacing an identification expression that enables re-identification of an individual with a non-identification expression.

[0053] "Re-identification of an individual" may mean that a specific individual is identified again through a combination of appearance, clothing, demographic attributes, time, and place, even if the face is not directly exposed in the video. "Identifying expression" may mean natural language expressions that enable such re-identification (e.g., "male in his 30s wearing a green jumper," "in front of Exit 3," "08:12"), and "non-identifying expression" may mean expressions that convey the same situation but are generalized or deleted so as not to contribute to individual identification (e.g., "one user," "near the entrance"). "Non-identifying contextual text" may mean human-readable contextual descriptions resulting from the replacement of identifying expressions with non-identifying expressions.

[0054] For example, the non-identifiable text generation unit (140) can generate a non-identifiable situation text for the integrated semantic event U1 by replacing the primary description “a man in his 30s wearing a green jumper collapsed in front of Exit 3 of Station A at 08:12” with “one user collapsed near the entrance of Station A”, and at this time, the category of the occurrence location and the type of situation, which are minimum information required for control, can be maintained.

[0055] The priority determination unit (150) can calculate the dynamic priority of the integrated semantic event based on the urgency, public risk, and detection reliability of the integrated semantic event, and readjust the dynamic priority whenever a new unit event is generated.

[0056] This operation is an operation that "calculates dynamic priorities based on urgency, public risk, and detection reliability," and the values ​​serving as the basis and the values ​​being calculated can be defined as follows.

[0057] "Urgency" may refer to the degree of temporal urgency regarding the impact of the relevant integrated semantic event on human life or safety (e.g., falling on the tracks is high urgency, simple abandonment on the outskirts is low urgency). "Public risk" may refer to the magnitude of harm that the relevant integrated semantic event may cause to an unspecified number of people (e.g., crowd density on a crowded platform is high public risk). "Detection confidence" may refer to the degree of certainty that the relevant event is a real occurrence, as previously defined; a lower detection confidence may be treated as indicating a higher probability of false positives. "Dynamic priority" may refer to a numerical value calculated by combining the above values ​​to determine the processing order.

[0058] For example, the priority determination unit (150) can calculate the said urgency by reflecting weights based on the directness of harm to human life (e.g., whether it is a high-risk location such as a track or road) and the speed of deterioration of condition (e.g., duration of non-response after falling) to a predefined standard urgency value for each event type. Similarly, the priority determination unit (150) can calculate the said public risk to have a larger value as the estimated number of people who may be affected and the possibility of spatial spread of the event increase.

[0059] For example, the priority determination unit (150) can determine the dynamic priority of the integrated semantic event U1 as the highest based on the fact that the urgency and public risk are both high and the detection reliability is also high, as the collapse and crowd density within the station during rush hour are combined. On the other hand, for a short unauthorized intrusion detected by an outer camera during the early morning hours, the dynamic priority can be determined as the lower based on the fact that the urgency and public risk are low. Additionally, if a new unit event is additionally created and incorporated into an integrated semantic event, the priority determination unit (150) can readjust the dynamic priority of the integrated semantic event, considering that the scale of the event has increased. By updating the dynamic priority, the processing order can be changed in real time according to the development of the situation.

[0060] The control provision unit (160) can provide control information including the non-identifiable situation text, the urgency, and the recommended measures to the control terminal for the top N integrated semantic events in order of high dynamic priority.

[0061] "Top N events" may refer to N integrated semantic events that are located at the top when the calculated dynamic priority is sorted in descending order, where N is an integer greater than or equal to 1. "Recommended action" may refer to information guiding the action that a control operator should take regarding the relevant situation (e.g., "recommendation to dispatch station staff and contact 119"). "Control information" may refer to a bundle of information provided to a control terminal, including non-identifiable situation text, urgency, and recommended action. "Control terminal" may refer to a device through which a control operator checks the above control information.

[0062] For example, the control provider (160) can create an item corresponding to the representative type and urgency of the integrated semantic event by referring to a response manual in which recommended actions are predefined for each event type or representative type (e.g., collapse -> dispatch of rescue personnel and linkage with 119, unauthorized entry -> dispatch of security personnel, abandoned object -> instruction to check the site) and selecting the item as the recommended action.

[0063] For example, the control provision unit (160) can provide control information to the control terminal for the integrated semantic event U1, which has the highest dynamic priority, consisting of {non-identified situation text="One user collapsed near Station A entrance", urgency=high, recommended action="Recommendation to dispatch station staff and link with 119"}. At this time, by the control provision unit (160) selecting and providing only the top N cases, the control personnel do not have to check all events occurring from multiple cameras one by one, thereby reducing the fatigue of the control personnel.

[0065] FIG. 3 is a block diagram illustrating the operation of an integrated event recognition unit according to one embodiment.

[0066] FIG. 3 may be a flowchart illustrating the process of an integrated event recognition unit (130) combining two or more unit events into one integrated semantic event. The integrated event recognition unit (130) may compare two unit events as a pair by designating one of the multiple unit events generated by the event detection unit (120) as the first unit event and the other as the second unit event.

[0067] Here, "first unit event" may refer to a unit event arbitrarily designated to refer to one of the two unit events being compared, and "second unit event" may refer to the other unit event being compared. That is, the expressions "first" and "second" are merely names to distinguish the two unit events and do not limit the order of occurrence or importance. Each unit event may include time (time of capture), place (place identifier), and type (event type), as illustrated in FIG. 3.

[0068] For example, the integrated event recognition unit (130) can compare two unit events as a pair by making unit event E1 detected by camera C1 {Type = collapse, (C1, Station A-Platform, 08:12:31), detection reliability = 0.92} the first unit event and unit event E2 detected by camera C2 {Type = crowd density, (C2, Station A-Waiting Room, 08:12:33), detection reliability = 0.81} the second unit event.

[0069] The integrated event recognition unit (130) determines the first unit event and the second unit event as a temporal candidate pair when the difference between the time of shooting of the first unit event and the time of shooting of the second unit event is within a preset time window, and when it exceeds the time window, the first unit event and the second unit event can be maintained as separate integrated semantic events.

[0070] This operation is an operation that "determines temporal candidate pairs based on whether the difference in shooting times is within a time window," and the value serving as the basis for the judgment and the determined result can be defined as follows.

[0071] "Difference in shooting time" may refer to the time interval between the shooting time of the first unit event and the shooting time of the second unit event. "Pre-set time window" refers to an allowable time range in which two unit events can be considered to have originated from the same event in time, and may refer to a preset value (e.g., 30 seconds), such as T_win indicated in FIG. 3. "Temporal candidate pair" may refer to a pair of two unit events that are initially determined to have the potential to be grouped into a single event because the difference in shooting time is within the time window. "Maintain as separate integrated semantic events" refers to not grouping the two unit events and leaving each as an independent event, i.e., the "maintain separate events (integrated semantic event undecided)" state in FIG. 3.

[0072] For example, the integrated event recognition unit (130) can determine E1 and E2 as a temporal candidate pair based on the fact that the difference between the time of capture of the first unit event E1 at 08:12:31 and the time of capture of the second unit event E2 at 08:12:33 is 2 seconds and this value is within a time window of 30 seconds. On the other hand, if the difference between the time of capture of the first unit event E1 at 08:12:31 and the time of capture of another unit event E3 at 08:40:00 is about 27 minutes and exceeds a time window of 30 seconds, the integrated event recognition unit (130) can maintain E1 and E3 as separate integrated semantic events without grouping them together.

[0073] The integrated event recognition unit (130) can determine the temporal candidate pair as a spatiotemporal candidate pair if the location identifier of the first unit event and the location identifier of the second unit event forming the temporal candidate pair belong to the same predefined control zone. As shown in FIG. 3, if the two location identifiers do not belong to the same control zone (No), they are not determined as a spatiotemporal candidate pair, and the first unit event and the second unit event can be maintained as separate integrated semantic events.

[0074] This operation is an operation that "determines spatiotemporal candidate pairs based on whether two location identifiers belong to the same control area," and the value serving as the basis for the judgment and the determined result can be defined as follows.

[0075] "Location identifier" may refer to a value that distinguishes the location or jurisdiction where a unit event is detected, as defined in Claim 1. "Predefined same control zone" may refer to a zone that is pre-grouped so that even if they are different location identifiers, they are managed together as a single control unit. For example, location identifiers "Station A-Platform" and "Station A-Waiting Room" are different location identifiers, but they may both be pre-defined to belong to the control zone "Station A". "Spatial-temporal candidate pair" may refer to a pair of two unit events that satisfy both time and location conditions, are temporally close, and are determined to have occurred in the same control zone.

[0076] For example, the integrated event recognition unit (130) can determine E1 (place identifier = Station A-Platform) and E2 (place identifier = Station A-Waiting Room), which are determined as temporal candidate pairs, as spatiotemporal candidate pairs based on the fact that both place identifiers belong to the same control area "Station A". On the other hand, even if the first unit event E1 (Station A) and the unit event detected by the nearby intersection camera C3 form a temporal candidate pair, if the two place identifiers belong to different control areas (Station A and the intersection), the integrated event recognition unit (130) can maintain the two unit events as separate integrated semantic events.

[0077] The integrated event recognition unit (130) can group the spacetime candidate pair into one integrated semantic event if the combination of the event type of the first unit event and the event type of the second unit event forming the spacetime candidate pair corresponds to a predefined simultaneous occurrence rule, and can maintain the first unit event and the second unit event as separate integrated semantic events if they do not correspond to the simultaneous occurrence rule.

[0078] This operation is an operation that "determines an integrated semantic event based on whether a combination of two event types corresponds to a co-occurrence rule," and the value serving as the basis for the judgment and the determined result can be defined as follows.

[0079] "Combination of event types" may refer to a pair formed by combining the event type of a first unit event with the event type of a second unit event (e.g., falling + crowding). "Predefined simultaneous occurrence rules" may refer to a set of rules that predetermine which event types should be integrated into a single event when they occur together. For example, as shown in the rule example in Fig. 3, the combination of "falling + crowding" can be predefined as an integration target. "Bundling into an integrated semantic event" may refer to combining two unit events forming a spatiotemporal candidate pair into a single event unit, and as illustrated in Fig. 3, a single integrated semantic event may be generated as a result.

[0080] For example, the integrated event recognition unit (130) can group E1 (event type = falling) and E2 (event type = crowd density), which are determined as spatiotemporal candidate pairs, into a single integrated semantic event U1 based on the fact that the combination of "falling + crowd density" corresponds to the simultaneous occurrence rule. In this case, the integrated semantic event U1 can be recognized as a single event unit, such as "a fall that occurred at Station A platform and the resulting crowd density and need for rescue." On the other hand, if the event types of the two unit events forming the spatiotemporal candidate pair are the combination of "falling + abandoned object" and this combination does not correspond to the simultaneous occurrence rule, the integrated event recognition unit (130) can not group the two unit events and maintain them as separate integrated semantic events.

[0081] By sequentially applying time conditions, location conditions, and type conditions to the integrated event recognition unit (130) to group unit events into integrated semantic events, the system (100) can recognize events detected individually from different cameras as a single event unit, prevent the same event from being duplicated as multiple notifications, and simultaneously suppress the incorrect combination of unrelated events through step conditions.

[0083] FIG. 4 is a flowchart illustrating the operation of an integrated event recognition unit according to one embodiment.

[0084] FIG. 4 may be a flowchart illustrating the process in which the integrated event recognition unit (130) groups spatiotemporal candidate pairs into an integrated semantic event and determines the representative type and public risk level.

[0085] The integrated event recognition unit (130) can calculate a spatiotemporal correlation for spatiotemporal candidate pairs, having a larger value as the difference between the shooting time of the first unit event and the second unit event becomes smaller, and having a larger value as the degree of overlap between the camera corresponding to the camera identifier of the first unit event and the camera corresponding to the camera identifier of the second unit event becomes larger (S410).

[0086] This operation is an operation that "calculates spatiotemporal correlation based on the difference in shooting time and the degree of overlap of viewing angles," and the value serving as the basis for the calculation and the calculated value can be defined as follows.

[0087] "Difference in shooting time" may refer to the time interval between the shooting time of the first unit event and the shooting time of the second unit event, as defined in Claim 2. "Camera corresponding to the camera identifier" may refer to the actual shooting device pointed to by the camera identifier included in the source metadata of the unit event. "Field of view overlap" may refer to the degree of spatial overlap between the field of view ranges captured by each of the two cameras; it may have a larger value when the two cameras illuminate the same area together, and a smaller value when they illuminate only different areas. The field of view overlap may be calculated and stored in advance based on the installation position, orientation, and field of view of each camera. "Spatiotemporal correlation" is a numerical value representing the probability that two unit events originated from a single identical event; it may have the characteristic of having a larger value when the difference in shooting time is smaller and a larger value when the field of view overlap is larger.

[0088] That is, the integrated event recognition unit (130) can calculate a larger spatiotemporal correlation value by considering that the two unit events are likely to be the same event as the closer the two unit events are to each other in time and as much as the two cameras have captured an area that overlaps.

[0089] For example, the integrated event recognition unit (130) can calculate a spatiotemporal correlation value of a relatively large value, such as 0.86, based on the fact that the difference in shooting time is small (2 seconds) and the degree of overlap of the field of view between camera C1 and camera C2 is relatively high for the first unit event E1 ((C1, Station A-Platform, 08:12:31)} and the second unit event E2 ((C2, Station A-Waiting Room, 08:12:33)} forming a spatiotemporal candidate pair. On the other hand, for other spatiotemporal candidate pairs where the difference in shooting time is larger (25 seconds) and the two cameras only illuminate different areas and the degree of overlap of the field of view is low, the integrated event recognition unit (130) can calculate a spatiotemporal correlation value of a relatively small value, such as 0.30.

[0090] The integrated event recognition unit (130) can group a pair of spatiotemporal candidate pairs into a single integrated semantic event only when the calculated spatiotemporal correlation is greater than or equal to a reference value (S420).

[0091] This operation is an operation that "determines an integrated semantic event based on whether the spatiotemporal correlation is greater than or equal to a threshold value," and the value serving as the basis for such determination and the determined result can be defined as follows.

[0092] "Criterion value" may refer to a pre-set threshold value to determine whether to group a pair of spatiotemporal candidate pairs into a single event. "Grouping only when greater than or equal to threshold value" may mean combining a pair of spatiotemporal candidate pairs into a unified semantic event only when the spatiotemporal correlation is greater than or equal to the threshold value, and not combining them when the spatiotemporal correlation is less than the threshold value.

[0093] For example, if the reference value is set to 0.70, the integrated event recognition unit (130) may group a spatiotemporal candidate pair (E1, E2) with a spatiotemporal correlation of 0.86 into a single integrated semantic event U1 based on the fact that the spatiotemporal correlation is greater than or equal to the reference value. On the other hand, for another spatiotemporal candidate pair with a spatiotemporal correlation of 0.30, it may not group them into an integrated semantic event based on the fact that the spatiotemporal correlation is less than the reference value. Even if the spatiotemporal candidate pair satisfies all the conditions of claim 2, the integrated event recognition unit (130) may combine them only when the spatiotemporal correlation is greater than or equal to the reference value, thereby further preventing accidental simultaneous occurrences that are not filtered out by time, place, and type conditions alone from being wrongly grouped into a single event.

[0094] The integrated event recognition unit (130) can perform an operation to select a unit event that maximizes the product of detection reliability and a risk weight assigned by event type among the included unit events as a representative unit event when two or more unit events are included in one integrated semantic event, and then set the event type of the representative unit event as the representative type of the integrated semantic event (S430).

[0095] This operation is an operation that "selects the unit event that maximizes the product of detection reliability and risk weight as the representative unit event," and the value serving as the basis for calculation and the selected and set result can be defined as follows.

[0096] "Detection reliability" may refer to a value indicating the degree of certainty that the detection and classification results correspond to actual events, as defined in Claim 1. "Pre-assigned risk weights by event type" refers to values ​​pre-assigned to each type based on the severity of the impact each event type has on safety; for example, relatively large risk weights may be assigned to collapses and dangerous behaviors, while relatively small risk weights may be assigned to abandoned objects. "Unit event maximizing the product of detection reliability and risk weights" may refer to the unit event with the largest product when the product of detection reliability and the risk weight of that event type is calculated for each unit event included in the integrated semantic event. "Representative unit event" may refer to a unit event selected in this manner, and "Representative type" may refer to setting the event type of the representative unit event as a type that represents the entire integrated semantic event.

[0097] For example, if unit events E1 (type = collapse, detection reliability = 0.92) and E2 (type = crowd density, detection reliability = 0.81) are included in the integrated semantic event U1, and the risk weight for collapse is 0.9 and the risk weight for crowd density is 0.6, the integrated event recognition unit (130) can calculate the product of E1 as 0.92 × 0.9 = 0.828 and the product of E2 as 0.81 × 0.6 = 0.486. The integrated event recognition unit (130) can select E1, which has a larger product, as the representative unit event and set the event type of E1, "collapse," as the representative type of the integrated semantic event U1. By setting a representative type by the integrated event recognition unit (130), even if an event consists of multiple unit events, a type that represents the event in a single word can be determined, and subsequently, the non-identifiable text generation unit (140) and the priority determination unit (150) can operate consistently based on the representative type.

[0098] The integrated event recognition unit (130) can adjust the public risk level of the integrated semantic event higher as the number of unit events included in the integrated semantic event increases (S440).

[0099] This operation is an operation that "adjusts the public risk level based on the number of included unit events," and the value serving as the basis for the adjustment and the value being adjusted can be defined as follows.

[0100] "Number of included unit events" may mean the total number of unit events grouped into a single integrated semantic event. "Public risk" may mean the magnitude of harm that an integrated semantic event may cause to an unspecified number of people, as defined in Claim 1. "Adjusted higher as the number increases" may mean that the more unit events grouped into the same event, the larger the scale of the event is considered to be, and thus the public risk is adjusted to a higher value.

[0101] For example, the integrated event recognition unit (130) can adjust the public risk level to a higher value for an integrated semantic event that includes four unit events, in which additional unit events detected by adjacent cameras are incorporated, compared to an integrated semantic event that includes only two unit events (E1, E2). By adjusting the public risk level according to the number of unit events, the integrated event recognition unit (130) can determine that an event large enough to be captured simultaneously at multiple points has a higher dynamic priority in the priority determination unit (150).

[0102] By having the integrated event recognition unit (130) perform additional verification based on spatiotemporal correlation, set representative types by maximizing the product, and correct public risk based on the number of unit events, the system (100) can more precisely exclude accidental simultaneous occurrences, consistently process events composed of multiple unit events by representing them as the most critical type, and increase the accuracy of subsequent priority calculation by reflecting the scale of the event in the public risk.

[0104] FIG. 5 is a block diagram illustrating the operation of a non-identifiable text generation unit according to one embodiment.

[0105] FIG. 5 may be a block diagram illustrating the process of a non-identifiable text generation unit (140) generating non-identifiable situational text from an integrated semantic event. The non-identifiable text generation unit (140) can generate primary situational text describing the situation using a non-language model from an integrated semantic event.

[0106] This operation is an operation that "generates primary context text with a vision-language model based on integrated semantic events," and the input values, means used, and generated results can be defined as follows.

[0107] "Integrated semantic event" may refer to a higher-level concept event in which multiple unit events are grouped into a single event unit, as defined in Claim 1. "Vision-language model" may refer to a model that receives an image or information obtained from an image as input and describes its content in natural language sentences. "Primary context text" may refer to a natural language context description text that has been primarily generated by the vision-language model for the integrated semantic event and has not yet undergone de-identification processing.

[0108] For example, the non-identifiable text generation unit (140) can generate primary situational text such as “there is another person running toward the police while lying down near the park around 3 p.m.” using a vision-language model for the integrated semantic event U1. The primary situational text is a human-readable description, but it may be in a state where there is a risk of a specific individual being re-identified by combining appearance, demographic attributes, time, and location.

[0109] The non-identifiable text generation unit (140) can decompose the primary situation text into an appearance description element indicating the appearance of a person, a personal description element indicating the demographic attributes of a person, a location description element indicating the location of occurrence, a time description element indicating the time of occurrence, and a situation description element indicating the type of situation.

[0110] Each narrative element can be defined as follows: "Appearance narrative element" may refer to expressions indicating a character's outward appearance, such as clothing, color, or possessions (e.g., "red jacket"). "Personal narrative element" may refer to expressions indicating demographic attributes, such as a character's age or gender (e.g., "middle-aged male"). "Location narrative element" may refer to expressions indicating the place where an event occurred (e.g., "near the park"). "Time narrative element" may refer to expressions indicating the time of day when an event occurred (e.g., "3:00 PM"). "Situational narrative element" may refer to expressions indicating the type of event or the situation itself (e.g., "falling down," "running toward the police").

[0111] For example, the non-identifiable text generation unit (140) can break down the preceding primary situation text into an appearance description element "red jacket", a personal description element "middle-aged male", a location description element "near the park", a time description element "3 PM", and a situation description element "collapsed".

[0112] The non-identifiable text generation unit (140) can classify appearance description elements and personal description elements as identification attributes that enable re-identification of an individual, and situation description elements as control essential attributes that are not identification attributes.

[0113] "Identification attribute" may refer to an attribute that allows a specific individual to be identified again by its expression alone or in combination with other expressions. "Control essential attribute" may refer to an attribute that does not contribute to individual identification but must be maintained for a control officer to assess the situation and respond. Appearance descriptive elements (e.g., "red jacket") and personal descriptive elements (e.g., "middle-aged male") can be classified as identification attributes because they contribute to identifying an individual, and situational descriptive elements (e.g., "falling down") can be classified as control essential attributes because they are necessary for response, even though they do not identify an individual. As illustrated in FIG. 5, location descriptive elements and visual descriptive elements are subject to original text preservation and can be preserved as they are.

[0114] The non-identification text generation unit (140) performs a substitution that generalizes a first identification attribute, which has a re-identification contribution less than a preset threshold among the descriptive elements classified as identification attributes, into a predefined higher concept word, performs a substitution that deletes the corresponding descriptive element for a second identification attribute, which has a re-identification contribution greater than or equal to the threshold among the descriptive elements classified as identification attributes, and can maintain the original text for descriptive elements classified as control essential attributes.

[0115] This operation is an operation that "replaces identification attributes based on whether the re-identification contribution is below or above a threshold," and the value serving as the basis for such judgment and the processing performed can be defined as follows.

[0116] "Re-identification contribution" is a value indicating the degree to which a descriptive element contributes to the re-identification of an individual; it may have a larger value as the expression becomes rarer or more unique. "Pre-set threshold" is a value set in advance to determine whether to generalize or delete an identification attribute, and may correspond to the threshold indicated in Fig. 5. "First identification attribute" may refer to a descriptive element classified as an identification attribute whose re-identification contribution is less than the threshold, and "second identification attribute" may refer to a descriptive element classified as an identification attribute whose re-identification contribution is greater than or equal to the threshold. "Substitution that generalizes to a pre-defined superordinate concept" may refer to a process that replaces a specific expression with a more general expression that encompasses it (e.g., "red jacket" -> "colored clothing"). "Substitution that deletes" may refer to a process that removes the corresponding descriptive element from the resulting text. "Maintain original text" may refer to a process that leaves the corresponding descriptive element as is without modification.

[0117] For example, the non-identification text generation unit (140) may calculate the re-identification contribution such that the lower the frequency of appearance of the corresponding descriptive element expression calculated from statistics of the same area and same time period, the greater the value, or the greater the reduction in the number of people narrowed down to identification candidates when the corresponding descriptive element is combined with other descriptive elements, the greater the value.

[0118] For example, if the re-identification contribution of the appearance description element "red jacket" is 0.4 and the threshold is set to 0.6, the non-identification text generation unit (140) can perform a substitution that generalizes "red jacket" to the predefined super-concept term "colored clothing" by considering it as a first identification attribute with a re-identification contribution below the threshold. On the other hand, if the re-identification contribution of the personal description element "middle-aged male" is 0.7 or higher than the threshold, the non-identification text generation unit (140) can perform a substitution that deletes the corresponding description element by considering it as a second identification attribute. Meanwhile, since the situation description element "falling down" is classified as a control essential attribute, the non-identification text generation unit (140) can maintain it as the original text without change, and as shown in FIG. 5, the original text can also be maintained for the location description element "near the park" and the time description element "3 PM".

[0119] The non-identifiable text generation unit (140) can output the result of recombining substituted or retained descriptive elements as non-identifiable situational text.

[0120] "Recombination" may mean a process of recombining the remaining descriptive elements, excluding generalized substituted descriptive elements and deleted descriptive elements, and the original text preserved descriptive elements into a single natural language sentence. "Non-identifiable contextual text" may mean a human-readable contextual description resulting from the replacement of an identifiable expression with a non-identifiable expression, as defined in Claim 1.

[0121] For example, the non-identifiable text generation unit (140) can output non-identifiable situational text such as “a person wearing colored clothing collapsed near the park around 3 PM” by recombining the generalized appearance description element “colored clothing”, the deleted personal description element, the original text preserved location description element “near the park”, the original text preserved time description element “3 PM”, and the original text preserved situational description element “collapsed”. By the non-identifiable text generation unit (140) generalizing or deleting identification attributes according to the re-identification contribution while maintaining the essential attributes for control, the system (100) can prevent the natural language description generated by the vision-language model itself from becoming new personal information while maintaining the minimum information necessary for control.

[0122] By having a non-identification text generation unit (140) decompose primary situation text into descriptive elements and classify each descriptive element into identification attributes or essential monitoring attributes, and then applying differential generalization, deletion, and retention based on a comparison of re-identification contribution and threshold, the system (100) can mitigate the secondary personal information risk of generated text that is not resolved by structural non-identification of raw images alone, and at the same time minimize the loss of situation information necessary for monitoring response.

[0124] FIG. 6 is a flowchart illustrating the operation of a non-identifiable text generation unit according to one embodiment.

[0125] FIG. 6 may be a flowchart illustrating the process in which a non-identifying text generation unit (140) further generalizes and outputs non-identifying situation text based on the risk of re-identification. Claim 5 may be a configuration that performs a quantitative re-identification risk assessment and iterative additional generalization, taking into account spatiotemporal information, on the result of the primary non-identification processing according to Claim 4.

[0126] The non-identifiable text generation unit (140) can map location description elements and time description elements into spatiotemporal buckets corresponding to a place identifier and a time of shooting (S610).

[0127] This operation is an operation that "maps location descriptive elements and time descriptive elements to spatiotemporal buckets by matching them to place identifiers and shooting times," and the input values ​​and mapped results can be defined as follows.

[0128] "Location description element" and "time description element" may refer to expressions indicating the location and time of occurrence, respectively, as defined in Claim 4. "Place identifier" and "time of capture" may refer to values ​​included in the source metadata of a unit event, as defined in Claim 1. "Spatiotemporal bucket" may refer to dividing continuous places and times into specific sections, grouping a combination of a spatial section and a time section to which a certain location or time belongs into a single unit. For example, space may be divided into zone units such as "near Station A Entrance," and time into time zone units such as "3:00 PM to 4:00 PM," and the combination may form a spatiotemporal bucket.

[0129] For example, the non-identifiable text generation unit (140) can map the location description element "near the park" and the time description element "3:00 PM" to the area indicated by the place identifier and the time period to which the time of shooting belongs, thereby mapping them to the spatiotemporal bucket B={Area=Station A-Entrance, Time period=15:00}. By mapping the location and time to the spatiotemporal bucket, the non-identifiable text generation unit (140) can quantitatively evaluate the risk of re-identification based on how many people usually exist in the same spatiotemporal bucket.

[0130] The non-identification text generation unit (140) can calculate a re-identification risk that has a larger value as the number of estimated users belonging to the spatiotemporal bucket decreases and a larger value as the number of identification attributes remaining in the non-identification situation text increases (S620).

[0131] This operation is an operation that "calculates the risk of re-identification based on the estimated number of users and the number of remaining identification attributes," and the value serving as the basis for the calculation and the calculated value can be defined as follows.

[0132] "Estimated number of users" may refer to the number of people presumed to typically exist in the relevant spatiotemporal bucket, and may be calculated in advance from statistics of the same area and time period. In spatiotemporal buckets with a small estimated number of users, it is difficult for one person to be hidden by mixing with others, so the possibility of a specific individual being identified even with the same description may increase. "Number of remaining identifying attributes" may refer to the number of identifying attributes that still remain in the de-identified situation text even after the first de-identification processing of Claim 4. "Re-identification risk" is a numerical value representing the probability that a specific individual will be re-identified from the de-identified situation text, and may have the characteristic of having a larger value as the estimated number of users decreases and a larger value as the number of remaining identifying attributes increases.

[0133] For example, the non-identification text generation unit (140) can calculate the re-identification risk as a relatively large value, such as 0.78, based on the fact that the estimated number of users of spatiotemporal bucket B is relatively small at 8 people and there are 2 remaining identification attributes (e.g., generalized appearance expression and location expression) in the non-identification situation text. On the other hand, for other spatiotemporal buckets where the estimated number of users is large at 500 people and there are 0 remaining identification attributes, the re-identification risk can be calculated as a small value, such as 0.10.

[0134] The non-identification text generation unit (140) may repeat the process of recalculating the re-identification risk level after performing additional generalization to generalize the remaining identification attributes, location description elements, and visual description elements from a low generalization level to a high generalization level when the calculated re-identification risk level exceeds the standard risk level, until the re-identification risk level becomes lower than or equal to the standard risk level, and may not perform additional generalization when the re-identification risk level does not exceed the standard risk level (S630).

[0135] This operation is an operation that "repeats additional generalization based on whether the re-identification risk level exceeds the reference risk level," and the value serving as the basis for the judgment and the processing performed can be defined as follows.

[0136] "Reference risk level" may refer to a threshold value of the re-identification risk level set in advance to determine a level at which non-identifiable situation text can be output. "Generalization level one step from a low expression to a high expression" refers to a process of raising the expression one step at a time from a more specific to a more comprehensive one, for example, raising the location expression one step in the order of "Station A Exit 3" -> "Near Station A Entrance" -> "Near Station A", and raising the time expression one step in the order of "3:12 PM" -> "Around 3 PM" -> "Afternoon time." "Additional generalization" may refer to generalization performed one step further after the first non-identification processing of claim 4. "Repeat until below reference risk level" may mean checking whether the result is below the reference risk level each time additional generalization and recalculation of the re-identification risk level are performed, and if it is not below, repeating the process of generalizing one step further until the condition is satisfied.

[0137] For example, if the reference risk level is set to 0.50, the non-identifying text generation unit (140) can obtain 0.61 by recalculating the re-identifying risk level based on the fact that the re-identifying risk level exceeds the reference risk level by 0.78, and by generalizing the remaining identification attributes and location and time description elements by one step (e.g., "colored clothing" is left as a remaining identification attribute, but the location "Station A Exit 3" is raised to "near Station A Entrance" and the time "3 PM" is raised to "around 3 PM"). Since 0.61 still exceeds the reference risk level, the non-identifying text generation unit (140) can recalculate the re-identifying risk level to 0.42 by generalizing by one step again (e.g., the location is raised to "near Station A" and the time is raised to "afternoon time"), and since 0.42 is less than or equal to the reference risk level of 0.50, the repetition can be terminated. Meanwhile, if the initial re-identification risk level does not exceed the standard risk level from the beginning, such as 0.42, the non-identification text generation unit (140) can use the result of claim 4 as is without performing additional generalization.

[0138] The non-identification text generation unit (140) can select a generalization result that maximizes the amount of information preserved from the control essential attribute among multiple generalization results in which the re-identification risk level is lower than or equal to the standard risk level and output it as a non-identification situation text (S640).

[0139] This operation is an operation that "selects a generalized result that maximizes the amount of information preserved for essential control attributes," and the value serving as the basis for comparison and the selected result can be defined as follows.

[0140] "In cases where there are multiple generalized results in which the re-identification risk is below the reference risk level" may mean a state in which two or more output candidates exist, such that different combinations of generalization (e.g., a result in which only location is further generalized and a result in which only time is further generalized) all satisfy the reference risk level or below. "Amount of information preserved from control-essential attributes" may mean an amount indicating how specifically the descriptive elements classified as control-essential attributes in Claim 4 remain in the generalized result, and may have a larger value as the situational information useful for control response remains more specifically. "Generalized result that maximizes the amount of preserved information" may mean the candidate among multiple output candidates in which the specificity of the control-essential attributes is maintained most significantly.

[0141] For example, as output candidates satisfying a threshold risk level or lower, there exist both candidate P1, which has a relatively specific situation description such as "there is a person moving quickly while lying down" by strongly generalizing the location, and candidate P2, which has a situation description that is only "an abnormal situation occurred" by lumping together the situation description. In this case, the non-identification text generation unit (140) can select candidate P1, which has a larger amount of information preserved for essential control attributes, and output it as non-identification situation text. By the non-identification text generation unit (140) selecting the candidate with the most control information among equally safe candidates, the system (100) can minimize the loss of information necessary for a control officer to judge the situation while lowering the risk of re-identification to below the threshold.

[0142] By the non-identifying text generation unit (140) quantifying the risk of re-identification based on spatiotemporal buckets, and gradually adding generalization of expressions until the risk level becomes below the reference level, and by selecting the result that maximizes the amount of information preserving essential control attributes among equally safe results, the system (100) can suppress re-identification by the generated text based on quantitative criteria even in quiet times and places with few estimated users, and at the same time, maintain information necessary for control response as specifically as possible.

[0144] FIG. 7 is a block diagram illustrating the operation of a priority determination unit according to one embodiment.

[0145] FIG. 7 may be a block diagram illustrating the process in which a priority determination unit (150) calculates and readjusts the dynamic priority of integrated semantic events and variably determines the top N cases to be provided to the control terminal.

[0146] The priority determination unit (150) may assign a first location risk factor if the location identifier of the integrated semantic event corresponds to any one of the predefined high-risk location groups including a railway, a roadway, an emergency room, and a school entrance, and may assign a second location risk factor smaller than the first location risk factor if it does not correspond to any of the high-risk location groups.

[0147] This operation is an operation that "assigns a location risk factor based on whether a location identifier corresponds to a high-risk location group," and the value serving as the basis for the judgment and the assigned value can be defined as follows.

[0148] "High-risk location group" refers to a set of locations designated in advance as high-risk due to a high probability of resulting in casualties in the event of an accident, and may include railway tracks, roadways, emergency rooms, and school entrances. "Location risk factor" may refer to a coefficient for reflecting the risk of a location where an integrated semantic event occurred in dynamic priority. "First location risk factor" may refer to a value assigned when a location identifier corresponds to a high-risk location group, and "Second location risk factor" may refer to a value smaller than the first location risk factor assigned when it does not correspond to a high-risk location group.

[0149] For example, the priority determination unit (150) may assign a first location risk factor of 1.5 to an integrated semantic event where the location identifier is "Station A-Track" based on the fact that the track belongs to a high-risk location group, and assign a second location risk factor of 1.0 to an integrated semantic event where the location identifier is "Station A-Waiting Room" based on the fact that it does not belong to a high-risk location group.

[0150] The priority determination unit (150) may assign a first time zone weight if the time of shooting of the integrated semantic event falls within a predefined congested time zone, and assign a second time zone weight smaller than the first time zone weight if it does not fall within a congested time zone.

[0151] This operation is an operation that "assigns time zone weights based on whether the shooting time falls within a congested time zone," and the value serving as the basis for the judgment and the assigned value can be defined as follows.

[0152] "Congestion time zone" refers to a time period when there is typically a high volume of users and a high likelihood of damage escalating in the event of an accident; it may be predefined, for example, as commuting hours. "Time zone weight" may refer to a coefficient for reflecting the congestion level at the time of an incident in dynamic priority. "First time zone weight" may refer to a value assigned when the time of recording falls within a congestion time zone, and "Second time zone weight" may refer to a value smaller than the first time zone weight assigned when the time does not fall within a congestion time zone.

[0153] For example, the priority determination unit (150) may assign a first time zone weight of 1.3 to an integrated semantic event with a shooting time of 08:12 based on the fact that the time falls within the morning rush hour, and assign a second time zone weight of 1.0 to an integrated semantic event with a shooting time of 03:40 based on the fact that it does not fall within the rush hour.

[0154] The priority determination unit (150) can calculate a density correction factor that has a larger value as the density of the area where the integrated semantic event occurred increases.

[0155] "Zone density" may refer to the number of people per unit area currently existing in the zone where an integrated semantic event has occurred. "Density correction factor" is a coefficient for reflecting the degree of congestion in the zone to dynamic priority, and may have a characteristic of having a larger value as the density increases. For example, the priority determination unit (150) may calculate a density correction factor of 1.4 for an integrated semantic event with high platform density and a density correction factor of 1.0 for an integrated semantic event with low density.

[0156] The priority determination unit (150) can calculate a predicted risk level that has a larger value as the probability of deterioration increases, based on a pre-learned deterioration probability for a representative type of integrated semantic event.

[0157] This operation is an operation that "calculates predicted risk based on the deterioration probability of representative types," and the value serving as the basis for the calculation and the calculated value can be defined as follows.

[0158] "Representative type" may refer to an event type that represents an integrated semantic event as defined in claim 3. "Pre-learned aggravation probability" may refer to a value learned in advance from past data regarding the probability that an event of a representative type will progress to a more severe state over time. "Predicted risk" is a value used to reflect the likelihood of an event worsening in the dynamic priority, and may have the characteristic of having a larger value as the aggravation probability increases.

[0159] For example, the priority determination unit (150) can calculate a predicted risk value as a large value, such as 0.8, based on a high pre-learned deterioration probability for an integrated semantic event where the representative type is "falling," and a predicted risk value as a small value, such as 0.2, based on a low deterioration probability for an integrated semantic event where the representative type is "abandoned."

[0160] The priority determination unit (150) can calculate a false positive probability that has a larger value as the detection reliability is lower, and a reliability correction coefficient that has a smaller value as the false positive probability is higher.

[0161] This operation is an operation that "calculates the probability of a false positive based on the detection reliability and calculates a reliability correction factor based on the probability of a false positive," and each basis value and result value can be defined as follows.

[0162] "Detection reliability" may refer to the degree of certainty that the detection and classification result corresponds to an actual event, as defined in Claim 1. "False positive probability" refers to the probability that the detection result is not actually an event, and may have a characteristic of having a larger value as the detection reliability decreases. "Reliability correction factor" refers to a coefficient for dampening dynamic priority so that events with a high probability of false positive do not receive an excessively high priority, and may have a characteristic of having a smaller value as the probability of false positive increases.

[0163] For example, the priority determination unit (150) can calculate the false positive probability as 0.05 and the reliability correction factor as 0.95 for integrated semantic events with a high detection reliability of 0.95, and the false positive probability as 0.45 and the reliability correction factor as 0.60 for integrated semantic events with a low detection reliability of 0.55.

[0164] The priority determination unit (150) can calculate a dynamic priority by multiplying a reliability correction factor by the weighted sum of the urgency, public risk, first location risk coefficient or second location risk coefficient, first time zone weight or second time zone weight, density correction factor and predicted risk.

[0165] "Weighted sum" may refer to a value obtained by multiplying each value by a predetermined weight and adding them together. "Calculated by multiplying by a reliability correction factor" may refer to a process in which the value obtained from the weighted sum is multiplied by a reliability correction factor to proportionally lower the priority of events with a high probability of false positives.

[0166] For example, the priority determination unit (150) can calculate a dynamic priority of 80 for an integrated semantic event that obtains a value of 80 by weighted summing of urgency, public risk, location risk coefficient, time zone weight, density correction coefficient, and prediction risk, and 80 × 0.95 = 76 if the reliability correction coefficient is 0.95, and can calculate a dynamic priority of 80 × 0.60 = 48 even with the same weighted sum value if the reliability correction coefficient is 0.60. By multiplying the weighted sum value by the reliability correction coefficient, the priority determination unit (150) can suppress the incorrect placement of events with a high probability of false positives at the top.

[0167] The priority determination unit (150) can readjust the dynamic priority of an integrated semantic event upward when a new unit event is incorporated into an integrated semantic event, and readjust the dynamic priority of an integrated semantic event downward by applying a decay function over time when a new unit event is not incorporated into an integrated semantic event for a preset maintenance time.

[0168] "Incorporation of a new unit event" may mean that a unit event satisfying the integration conditions of claims 2 and 3 is additionally bound to an already formed integrated semantic event. "Pre-set retention time" may mean a time set in advance to determine the time maintained without the incorporation of a new unit event. "Damping function" may mean a function that gradually lowers the dynamic priority as time elapses.

[0169] For example, the priority determination unit (150) may readjust the dynamic priority upward when a new unit event detected by an adjacent camera is incorporated into an integrated semantic event and the event scale increases, and may gradually readjust the dynamic priority downward by applying a damping function when a new unit event is not incorporated into another integrated semantic event for a preset retention time. By the priority determination unit (150) readjusting the dynamic priority upward and downward according to the development of the situation, the processing order can be updated according to the real-time change of the event.

[0170] The priority determination unit (150) defines the number of control information items provided to the control terminal per unit time as a control fatigue index, and variably determines N of the top N items so that the control fatigue index does not exceed a preset allowable load, and if the control fatigue index exceeds the allowable load, N is reduced so that integrated semantic events with high dynamic priority are provided first, and if the control fatigue index does not exceed the allowable load, N is increased within the allowable load range.

[0171] This operation is an operation that "variably determines N based on whether the control fatigue index exceeds the allowable load," and the value serving as the basis for the judgment and the determined value can be defined as follows.

[0172] "Control fatigue index" is a value that quantifies the burden received by a control operator and can be defined as the number of control information items provided to a control terminal per unit of time. "Pre-set allowable load" may refer to a value set in advance as an upper limit on the number of control information items per unit of time that a control operator can process without strain. "N of the top N items" may refer to the number of integrated semantic events provided to the control terminal in order of high dynamic priority, as defined in Claim 1. "Variably determining N" may mean increasing or decreasing N based on the comparison result between the control fatigue index and the allowable load, rather than fixing N.

[0173] For example, if the allowable load is set to 10 cases per unit time, the priority determination unit (150) can reduce N based on the fact that the control fatigue index is 14 cases, which exceeds the allowable load, so that integrated semantic events with high dynamic priority are provided to the control terminal in priority. On the other hand, if the control fatigue index is 6 cases, which does not exceed the allowable load, the priority determination unit (150) can increase N within the allowable load range so that more integrated semantic events are provided so that safety-related information is not omitted. By the priority determination unit (150) variably determining N according to the control fatigue index, the system (100) can quantitatively manage the fatigue of control personnel below the allowable load while ensuring that important events are not omitted.

[0174] The priority determination unit (150) calculates a dynamic priority by a weighted summation and multiplication operation that combines a location risk coefficient, a time zone weight, a density correction coefficient, a predicted risk, and a reliability correction coefficient, readjusts it upward or downward according to the development of events, and variably determines the number of cases N provided based on a comparison of the control fatigue index and the allowable load, thereby the system (100) can maintain a priority in real time that reflects all of the event type, location, time zone, density, predicted risk, and false positive possibility, and can process high-risk events first while quantitatively reducing the fatigue of control personnel.

[0176] FIG. 8 is a schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented.

[0177] Although the present disclosure has been described as generally being implementable by a system (100) (e.g., a server), those skilled in the art will know that the present disclosure may be implemented in combination with computer-executable instructions and / or other program modules that can be executed on one or more computers and / or as a combination of hardware and software.

[0178] Operation according to an embodiment of the present disclosure may be performed by a server (802) (e.g., a computing device). An exemplary environment for implementing various aspects of the present disclosure including a server (802) is shown in FIG. 8, and the server (802) includes a processing unit (804), a system memory (806), and a system bus (808). The system bus (808) connects system components, including the system memory (806) (but not limited thereto), to the processing unit (804). The processing unit (804) may be any processor among various commercial processors. Dual processors and other multiprocessor architectures may also be used as the processing unit (804). The system bus (808) may be any of several types of bus structures that can additionally be interconnected to a memory bus, a peripheral bus, and a local bus using any of various commercial bus architectures. System memory (806) includes read-only memory (ROM) (810) and random access memory (RAM) (812). The basic input / output system (BIOS) is stored in non-volatile memory (810), such as ROM, EPROM, EEPROM, etc., and this BIOS includes basic routines that help transfer information between components within the server (802) at times such as during startup. The RAM (812) may also include high-speed RAM, such as static RAM, for caching data. The server (802) also includes an internal hard disk drive (HDD) (814) (e.g., EIDE, SATA)—this internal hard disk drive (814) may also be configured for external use within a suitable chassis—a magnetic floppy disk drive (FDD) (816) (e.g., for reading from or writing to a removable diskette (818)), and an optical disk drive (820) (e.g., for reading from a CD-ROM disk (822) or reading from or writing to other high-capacity optical media such as a DVD).The hard disk drive (814), magnetic disk drive (816) and optical disk drive (820) can each be connected to the system bus (808) via the hard disk drive interface (824), magnetic disk drive interface (826), and optical drive interface (828).

[0179] The interface (824) for implementing an external drive includes at least one or both of USB (Universal Serial Bus) and IEEE 1394 interface technologies.

[0180] These drives and associated computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, etc. In the case of a server (802), the drives and media correspond to storing any data in a suitable digital format. Although the description of computer-readable media above refers to HDDs, removable magnetic disks, and removable optical media such as CDs or DVDs, those skilled in the art will know that other types of computer-readable media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, etc., may also be used in exemplary operating environments and that any of these media may contain computer-executable instructions for performing the methods of the present disclosure.

[0181] A number of program modules, including an operating system (830), one or more application programs (832), other program modules (834), and program data (836), may be stored in the drive and RAM (812). All or part of the operating system, application, module, and / or data may also be cached in RAM (812). It will be well known that the present disclosure may be implemented in various commercially available operating systems or combinations of operating systems. A user may input commands and information to the server (802) through one or more wired / wireless input devices, such as pointing devices like a keyboard (838) and a mouse (840). Other input devices may include a microphone, IR remote control, joystick, game pad, stylus pen, touch screen, etc. These and other input devices are often connected to the processing unit (804) via an input device interface (842) connected to the system bus (808), but may also be connected via other interfaces such as a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR interface, etc. A monitor (844) or other type of display device is also connected to the system bus (808) via an interface such as a video adapter (846). In addition to the monitor (844), the computer generally includes other peripheral output devices such as speakers, a printer, etc.

[0182] The server (802) may operate in a networked environment using a logical connection to one or more remote computers, such as remote computer(s) (848), via wired and / or wireless communication. The remote computer(s) (848) may be a workstation, a computing device computer, a router, a personal computer, a portable computer, a microprocessor-based entertainment device, a peer device, or other conventional network node, and generally include many or all of the components described for the server (802), but for brevity, only the memory storage device (850) is shown. The logical connection shown includes a wired / wireless connection to a local area network (LAN) (852) and / or a larger network, e.g., a wide area network (WAN) (854). Such LAN and WAN networking environments are common in offices and companies, facilitating enterprise computer networks such as intranets, all of which may be connected to a global computer network, e.g., the Internet. When used in a LAN networking environment, the server (802) is connected to a local network (852) via a wired and / or wireless communication network interface or adapter (856). The adapter (856) may facilitate wired or wireless communication to the LAN (852), and the LAN (852) may also include a wireless access point installed therein to communicate with the wireless adapter (856). When used in a WAN networking environment, the server (802) may include a modem (858), be connected to a communication computing device on the WAN (854), or have other means to establish communication over the WAN (854), such as through the Internet. The modem (858), which may be internal or external and wired or wireless, is connected to the system bus (808) via a serial port interface (842). In a networked environment, the program modules described for the server (802) or parts thereof may be stored in a remote memory / storage device (850).

[0183] It will be well known that the described network connection is exemplary and that other means of establishing communication links between computers may be used. The server (802) performs the operation of communicating with any wireless device or object deployed and operating via wireless communication, such as a printer, scanner, desktop and / or portable computer, PDA (portable data assistant), communication satellite, any equipment or place associated with a wireless detectable tag, and a telephone.

[0184] This includes at least Wi-Fi and Bluetooth wireless technologies. Accordingly, communication may be a predefined structure as in conventional networks, or simply ad hoc communication between at least two devices. The exemplary embodiments of this document and the terms used herein are not intended to limit the technical features described herein to specific embodiments and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components.

[0185] Where any (e.g., first) component is referred to as "coupled" or "connected" to another (e.g., second) component, with or without the terms "functionally" or "communicationly," it means that said component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component. The term "module" as used in the exemplary embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or part thereof that performs one or more functions. For example, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0186] The exemplary embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., internal memory or external memory) readable by a machine (e.g., an electronic device). For example, a processor (e.g., a processor) of the machine (e.g., an electronic device) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium. According to exemplary embodiments, the method according to the exemplary embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones).In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0188] A sensing device (e.g., an electronic device) according to the embodiments disclosed in this document may be of various forms. The sensing device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The sensing device according to the embodiments of this document is not limited to the aforementioned devices.

[0189] The embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0190] The term “module” as used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0191] One embodiment of the present document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., internal memory or external memory) readable by a machine (e.g., an electronic device). For example, a processor (e.g., a processor) of the machine (e.g., an electronic device) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0192] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0193] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

Claim 1 A system for de-identifying text and dynamically prioritizing public safety events through semantic analysis of multi-source video data, comprising: a video collection unit that receives video data from multiple cameras placed in different locations and adds source metadata including a camera identifier, a location identifier, and a time of capture to each of the video data; an event detection unit that detects objects and events from the video data, classifies the detected events into one of the event types such as falling, driving against traffic, trespassing, abandoned objects, crowd congestion, and dangerous behavior, and generates a unit event including the event type, the source metadata, and detection reliability; an integrated event recognition unit that recognizes two or more unit events that are spatiotemporally related among the multiple unit events by grouping them into a single integrated semantic event; a de-identifying text generation unit that generates de-identifying situational text for the integrated semantic event by replacing an identifying expression that enables the re-identification of an individual with a de-identifying expression; and a priority that calculates the dynamic priority of the integrated semantic event based on the urgency, public risk level, and detection reliability of the integrated semantic event, and readjusts the dynamic priority whenever a new unit event is generated. The system includes a decision unit; and a control provision unit that provides control information, including the non-identifiable situation text, the urgency, and the recommended measures, to a control terminal for the top N integrated semantic events in order of highest dynamic priority; wherein the integrated event recognition unit determines the first unit event and the second unit event as a temporal candidate pair when one of the plurality of unit events is designated as a first unit event and the other as a second unit event, if the difference between the time of capture of the first unit event and the time of capture of the second unit event is within a preset time window, and maintains the first unit event and the second unit event as separate integrated semantic events when the difference exceeds the time window.If the location identifier of the first unit event and the location identifier of the second unit event forming the aforementioned temporal candidate pair belong to the same predefined control zone, the aforementioned temporal candidate pair is determined as a spatiotemporal candidate pair; if the combination of the event type of the first unit event and the event type of the second unit event forming the spatiotemporal candidate pair corresponds to a predefined simultaneous occurrence rule, the spatiotemporal candidate pair is grouped into a single integrated semantic event; if it does not correspond to the simultaneous occurrence rule, the first unit event and the second unit event are each maintained as separate integrated semantic events; and the integrated event recognition unit calculates a spatiotemporal correlation for the spatiotemporal candidate pair, having a larger value as the difference between the shooting times of the first unit event and the second unit event decreases and a larger value as the degree of overlap of the field of view between the camera corresponding to the camera identifier of the first unit event and the camera corresponding to the camera identifier of the second unit event increases; and only if the calculated spatiotemporal correlation is greater than or equal to a reference value, the spatiotemporal candidate pair is grouped into a single integrated semantic event, wherein a single integrated A system characterized by, when two or more unit events are included in a semantic event, performing an operation to select a unit event that maximizes the product of the detection reliability and the risk weights pre-assigned for each event type among the included unit events as a representative unit event, setting the event type of the representative unit event as the representative type of the integrated semantic event, and adjusting the public risk of the integrated semantic event to be higher as the number of unit events included in the integrated semantic event increases. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 In claim 1, the priority determination unit assigns a first location risk coefficient when the location identifier of the integrated semantic event corresponds to any one of a predefined group of high-risk locations including railway tracks, roadways, emergency rooms, and school entrances, and assigns a second location risk coefficient smaller than the first location risk coefficient when it does not correspond to any of the high-risk location groups; assigns a first time zone weight when the time of capture of the integrated semantic event falls within a predefined congestion time zone, and assigns a second time zone weight smaller than the first time zone weight when it does not fall within the congestion time zone; calculates a density correction coefficient having a larger value as the density of the area where the integrated semantic event occurred increases; calculates a predicted risk having a larger value as the deterioration probability increases, based on a pre-learned deterioration probability for the representative type of the integrated semantic event; calculates a false positive probability having a larger value as the detection reliability decreases, and calculates a reliability correction coefficient having a smaller value as the false positive probability increases; and the urgency, the public risk, the first location risk coefficient, or the second The dynamic priority is calculated by multiplying the reliability correction factor by the weighted sum of the location risk factor, the first time zone weight or the second time zone weight, the density correction factor, and the predicted risk; when a new unit event is incorporated into any one of the integrated semantic events, the dynamic priority of the said integrated semantic event is readjusted upward; and when no new unit event is incorporated into any one of the integrated semantic events during a preset maintenance time, the dynamic priority of the said integrated semantic event is readjusted downward by applying a decay function according to the passage of time; the number of control information items provided per unit time to the control terminal is defined as the control fatigue index; and the N of the top N items is determined variably so that the control fatigue index does not exceed a preset allowable load,A system characterized by reducing N when the control fatigue index exceeds the allowable load to provide priority to integrated semantic events with high dynamic priority, and increasing N within the allowable load range when the control fatigue index does not exceed the allowable load.

Citation Information

Patent Citations

  • Method and Device for Selective Surveillance Based on Event Risk by Scene Categories

    KR102959954B1