Method and computer program for detecting text prompt-based event
By defining events as text prompts and analyzing video content for similarity, the method addresses the inefficiencies in existing video surveillance technologies, enhancing event detection accuracy and reducing human monitoring fatigue.
Patent Information
- Application Number
- PCT/KR2024/016841
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-29
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Existing video surveillance technologies face challenges in event detection, including high manager fatigue due to continuous monitoring and the need for extensive rule-setting in rule-based systems, which can lead to inefficient event detection.
The method involves defining events to be detected as text prompts and analyzing images by comparing these prompts with the video content, using techniques such as creating event vectors and section vectors in a potential space, and determining event strength based on similarity.
This approach allows for flexible and accurate event detection, reducing the burden on human monitors and improving the efficiency of event identification by enabling the analysis of images with human-like flexibility and accuracy.
Smart Images

Figure KR2024016841_08052025_PF_FP_ABST
Abstract
Description
Event detection method and computer program based on text prompts
[0001] The present invention relates to a method and a computer program for detecting an event based on a text prompt defining the event content.
[0002] Advances in information and communication technology are driving the adoption of artificial intelligence (AI) in many applications. AI is also actively being adopted in the field of video surveillance, and as a result, tasks previously performed by humans are increasingly being replaced by AI.
[0003] In the past, to detect specific events in videos, administrators monitored the videos or used techniques to detect predefined events in the videos based on rules.
[0004] In the case of technology where the administrator monitors the video, there was a problem that the administrator's continuous video monitoring was required, which resulted in high administrator fatigue, and there was a problem that gaps in monitoring could occur depending on the administrator's status, such as the administrator's absence.
[0005] In the case of rule-based technology, not only does it require setting rules for all situations that may occur, but it also requires setting multiple rules to respond to various cases for a single event, which is cumbersome and has problems with low event detection accuracy.
[0006] The present invention aims to solve the above-described problem by defining an event to be detected as text and comparing it with the situation of the image, thereby analyzing the image with human-like flexibility and accuracy.
[0007] An event detection method based on a text prompt according to one embodiment of the present invention may include the steps of: generating an event vector, which is a vector in a latent space, for each of one or more event prompts defined in natural language; extracting a feature for each of a plurality of sections constituting an image; generating an interval vector, which is a vector in the latent space for each of the plurality of sections, based on each of the extracted features; generating image analysis data based on a similarity between an interval vector and one or more event vectors in the latent space for each of the plurality of sections; and providing an analysis result of the image based on the image analysis data.
[0008] The step of generating the above image analysis data may include: determining an event intensity of each of one or more events for each of the plurality of sections based on a similarity between a section vector for each of the plurality of sections and one or more event vectors; grouping time points at which the event intensity exceeds a predetermined threshold to generate one or more upper event groups; and generating upper event information of the upper event group based on the content of an event prompt corresponding to each of one or more time points belonging to the same group and an occurrence order of the one or more time points for each of the one or more upper event groups.
[0009] The step of providing the above analysis result may include a step of providing an analysis screen including a first area that shows the event intensity of each of one or more events for each of the plurality of intervals in a time series manner, wherein the event intensity is calculated based on a similarity between an interval vector for each of the plurality of intervals and one or more event vectors.
[0010] The step of providing the above analysis result may include a step of providing an analysis screen including a second area that collects points in time when an event intensity exceeds a predetermined threshold for each of the one or more event prompts and displays them in a time series for each of the one or more events, wherein the event intensity is calculated based on a similarity between an interval vector for each of the plurality of intervals and one or more event vectors.
[0011] The step of providing the above analysis result may include a step of collecting points in time when the event intensity exceeds a predetermined threshold and providing an analysis screen including a third area that shows the event in a time series based on the point in time when the event occurs, wherein the event intensity is calculated based on a similarity between a section vector for each of the plurality of sections and one or more event vectors.
[0012] The step of providing the above analysis result may include the step of providing an event list in which one or more event details are displayed; the step of providing an event image corresponding to a first event selected from the event list in a fourth area; and the step of providing a slider bar in a fifth area that indicates the relative position of individual frames displayed in the fourth area within the event image according to playback of the event image and controls the displayed frame according to a user's input.
[0013] The step of providing the slider bar includes the step of providing a first object corresponding to at least a portion of the event image to be displayed; and the step of providing one or more thumbnails to be displayed on the first object, wherein a frame at a point in time when an event intensity exceeds a predetermined threshold is provided to be displayed as the thumbnail; and the event intensity may be calculated based on a similarity between a section vector for each of one or more sections constituting the event image and one or more event vectors.
[0014] The step of providing the above slider bar may further include the step of providing an object representing an event prompt related to one or more thumbnails so that the object is displayed in association with the one or more thumbnails.
[0015] The step of providing the above slider bar may further include the step of providing an object representing upper event group information including an event corresponding to each of the one or more thumbnails so as to be displayed in association with the one or more thumbnails.
[0016] The step of providing the above slider bar may include: a step of providing an image corresponding to a point in time of one of the one or more thumbnails so that it is displayed in the fourth area according to a user's selection of the one or more thumbnails; and a step of providing an image corresponding to a first event among one or more events belonging to the upper event group so that it is displayed in the fourth area according to a user's selection of an object representing the upper event group information.
[0017] According to the present invention, by defining an event to be detected as text and comparing it with the situation of the image, an image can be analyzed with human-like flexibility and accuracy.
[0018] In addition, the present invention provides an event video so that the user can view major events within the event video at once through a slider bar without having to view the entire video, and also allows the user to review events by event group.
[0019] FIG. 1 is a diagram schematically illustrating the configuration of an image analysis system according to one embodiment of the present invention.
[0020] FIG. 2 is a diagram schematically illustrating the configuration of a server (100) according to one embodiment of the present invention.
[0021] FIG. 3 is a diagram schematically illustrating the configuration of a user terminal (200) according to one embodiment of the present invention.
[0022] FIG. 4 is a diagram illustrating a process in which a server (100) generates a vector from an event prompt and an image according to one embodiment of the present invention.
[0023] FIG. 5 is a diagram illustrating an example of a graph representing image analysis data generated by a server (100) according to one embodiment of the present invention.
[0024] Figure 6 is a diagram illustrating an exemplary event group.
[0025] FIG. 7 is a drawing illustrating an exemplary image analysis result provision screen (600) displayed on a user terminal (200).
[0026] FIG. 8 is a drawing illustrating an exemplary image analysis result provision screen (700) displayed on a user terminal (200).
[0027] FIG. 9 is a flowchart illustrating an event detection method based on a text prompt performed by a server (100) according to one embodiment of the present invention. The method is described below with reference to FIGS. 1 to 8.
[0028] An event detection method based on a text prompt according to one embodiment of the present invention may include the steps of: generating an event vector, which is a vector in a latent space, for each of one or more event prompts defined in natural language; extracting a feature for each of a plurality of sections constituting an image; generating an interval vector, which is a vector in the latent space for each of the plurality of sections, based on each of the extracted features; generating image analysis data based on a similarity between an interval vector and one or more event vectors in the latent space for each of the plurality of sections; and providing an analysis result of the image based on the image analysis data.
[0029] The present invention is capable of various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention, as well as the methods for achieving them, will become clearer with reference to the embodiments described in detail below, along with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be implemented in various forms.
[0030] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same drawing reference numerals, and redundant descriptions thereof will be omitted.
[0031] In the following examples, terms such as first, second, etc. are not used in a limiting sense, but are used for the purpose of distinguishing one component from another. In the following examples, singular expressions include plural expressions unless the context clearly indicates otherwise. In the following examples, terms such as include or have mean that a feature or component described in the specification exists, and do not exclude in advance the possibility that one or more other features or components may be added. In the drawings, the sizes of components may be exaggerated or reduced for convenience of explanation. For example, the sizes and shapes of each component shown in the drawings have been arbitrarily shown for convenience of explanation, and therefore, the present invention is not necessarily limited to what is shown.
[0032]
[0033] FIG. 1 is a diagram schematically illustrating the configuration of an image analysis system according to one embodiment of the present invention.
[0034] An image analysis system according to one embodiment of the present invention can detect events within an image using one or more event prompts defined in natural language. Furthermore, the image analysis system according to one embodiment of the present invention can generate image analysis data based on the similarity between the event prompts and the image in latent space and provide the data to the user.
[0035] In the present invention, a 'prompt' may refer to an input value input by a user for the operation of a trained artificial neural network (or model). In addition, an 'event prompt' in the present invention may refer to an input value indicating an event to be detected using a trained artificial neural network (or model). Furthermore, an 'event prompt defined in natural language' may refer to an input value indicating an event to be detected using a trained artificial neural network (or model), and may refer to a value written in a language understandable to humans. For example, an event prompt defined in natural language may be a sentence written in natural language to detect fire and smoke, such as 'There is fire and smoke in the building.' However, such prompts are exemplary and the spirit of the present invention is not limited thereto.
[0036] In the present invention, 'latent space' may mean a space in which the latent characteristics of event prompts and images are quantified.
[0037] An image analysis system according to one embodiment of the present invention may include a server (100), a user terminal (200), an image storage device (300), an image acquisition device (400), and a communication network (500) as illustrated in FIG. 1.
[0038] According to one embodiment of the present invention, a server (100) can detect events within an image using one or more event prompts defined in natural language. Furthermore, according to one embodiment of the present invention, the server (100) can generate image analysis data based on the similarity between the event prompts and the image in latent space and provide the data to the user.
[0039]
[0040] FIG. 2 is a diagram schematically illustrating the configuration of a server (100) according to one embodiment of the present invention. Referring to FIG. 2, the server (100) according to one embodiment of the present invention may include a communication unit (110), a first processor (120), a memory (130), and a second processor (140). In addition, although not illustrated in the drawing, the server (100) according to one embodiment of the present invention may further include an input / output unit, a program storage unit, etc.
[0041] The communication unit (110) may be a device including hardware and software necessary for the server (100) to transmit and receive signals such as control signals or data signals through a wired or wireless connection with another network device such as a user terminal (200) and / or an image storage device (300).
[0042] The first processor (120) may be a device that controls a series of processes for detecting events in received images. For example, the first processor (120) may determine the similarity between a vector corresponding to an event prompt in latent space and a vector corresponding to a section of the image, and generate image analysis data based on the similarity.
[0043] Additionally, the first processor (120) may be a device that controls a series of processes for generating output data from input data using trained artificial neural networks. For example, the first processor (120) may be a device that controls a process of extracting features from an event prompt using a text model, or a process of extracting features from an image using a vision-language model.
[0044] Here, the processor may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program, for example. Examples of such data processing devices built into hardware include processing devices such as a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), and a field programmable gate array (FPGA), but the scope of the present invention is not limited thereto.
[0045] The memory (130) performs the function of temporarily or permanently storing data processed by the server (100). The memory may include a magnetic storage medium or a flash storage medium, but the scope of the present invention is not limited thereto. For example, the memory (130) may temporarily and / or permanently store data (e.g., coefficients) constituting a trained artificial neural network. Of course, the memory (130) may also store training data for training the artificial neural network or image data received from an image acquisition device (400). However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0046] The second processor (140) may refer to a device that performs operations under the control of the first processor (120) described above. In this case, the second processor (140) may be a device having higher computational capabilities than the first processor (120) described above. For example, the second processor (140) may be configured with a GPU (Graphics Processing Unit) and / or an NPU (Neural Processing Unit). However, this is merely exemplary and the spirit of the present invention is not limited thereto. In one embodiment of the present invention, the number of second processors (140) may be plural or singular.
[0047]
[0048] A user terminal (200) according to one embodiment of the present invention may be a device that provides an image analysis result provided by a server (100) to a user.
[0049] FIG. 3 is a diagram schematically illustrating the configuration of a user terminal (200) according to one embodiment of the present invention. Referring to FIG. 3, the user terminal (200) according to one embodiment of the present invention may include a communication unit (210), a third processor (220), a memory (230), and a fourth processor (240). In addition, although not illustrated in the drawing, the user terminal (200) according to one embodiment of the present invention may further include an input / output unit, a program storage unit, etc.
[0050] The communication unit (210) may be a device including hardware and software necessary for the user terminal (200) to transmit and receive signals such as control signals or data signals through a wired or wireless connection with another network device such as a server (100) and / or an image storage device (300).
[0051] In one embodiment of the present invention, the third processor (220) may provide the user with the image analysis results provided by the server (100). Furthermore, the third processor (220) may transmit requests based on user input to the server (100). For example, the third processor (220) may receive an event list from the server (100), provide it to the user, and request the server (100) for an image corresponding to an event selected by the user from the list.
[0052] Here, the processor may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program, for example. Examples of such data processing devices built into hardware include processing devices such as a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), and a field programmable gate array (FPGA), but the scope of the present invention is not limited thereto.
[0053] The memory (230) performs the function of temporarily or permanently storing data processed by the user terminal (200). The memory may include a magnetic storage medium or a flash storage medium, but the scope of the present invention is not limited thereto. For example, the memory (230) may temporarily and / or permanently store images received from the server (100). However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0054] The fourth processor (240) may refer to a device that performs operations under the control of the third processor (220) described above. In this case, the fourth processor (240) may be a device having higher computational capabilities than the third processor (220) described above. For example, the fourth processor (240) may be configured with a GPU (Graphics Processing Unit) and / or an NPU (Neural Processing Unit). However, this is merely exemplary and the spirit of the present invention is not limited thereto. In one embodiment of the present invention, the number of fourth processors (240) may be plural or singular.
[0055] A user terminal (200) according to one embodiment of the present invention may mean a portable terminal (201, 202, 203) or a computer (204), as shown in FIG. 1.
[0056] According to one embodiment of the present invention, a user terminal (200) may further include a display means for displaying content and an input means for obtaining user input regarding such content to perform the aforementioned functions. The input means and the display means may be configured in various ways. For example, the input means may include, but are not limited to, a keyboard, mouse, trackball, microphone, buttons, or touch panel.
[0057]
[0058] In one embodiment of the present invention, the image storage device (300) may be a device that temporarily or permanently stores images acquired by the image acquisition device (400). In addition, the image storage device (400) may be a device that provides stored images upon request from another device.
[0059] In another embodiment of the present invention, the server (100) and the image storage device (300) described above may be configured as one unit. In such an embodiment, the server (100) can detect events in an image and simultaneously store the image.
[0060]
[0061] An image acquisition device (400) according to one embodiment of the present invention may be a device that acquires images of a surveillance target environment or a surveillance target object and transmits them to another network device. Such image acquisition devices (400) may be singular or plural.
[0062]
[0063] A communication network (500) according to one embodiment of the present invention may refer to a communication network that mediates data transmission and reception between each component of an image analysis system. For example, the communication network (500) may encompass wired networks such as LANs (Local Area Networks), WANs (Wide Area Networks), MANs (Metropolitan Area Networks), and ISDNs (Integrated Service Digital Networks), or wireless networks such as wireless LANs, CDMA, Bluetooth, and satellite communication, but the scope of the present invention is not limited thereto.
[0064] Below, the process of detecting an event based on a text prompt by the server (100) is described.
[0065]
[0066] FIG. 4 is a diagram illustrating a process in which a server (100) generates a vector from an event prompt and an image according to one embodiment of the present invention.
[0067] A server (100) according to one embodiment of the present invention can generate an event vector, which is a vector in latent space, for each of one or more event prompts defined in natural language. More specifically, the server (100) according to one embodiment of the present invention can extract features from the prompts and generate an event vector, which is a vector in latent space, based on the extracted features. At this time, the server (100) according to one embodiment of the present invention can extract features using various text models.
[0068] For example, the server (100) can extract the first event feature (Event Feature 1) from the first event prompt (Event Prompt 1) and use it to generate the first event vector (EV1). Of course, the server (100) can also generate event vectors for the remaining event prompts through the same process.
[0069]
[0070] A server (100) according to one embodiment of the present invention can generate a segment vector for each of a plurality of segments constituting an image. More specifically, the server (100) according to one embodiment of the present invention can divide an image into a plurality of segments, extract features for each of the divided segments, and generate a segment vector, which is a vector in a latent space for each of the plurality of segments, based on each of the extracted features. At this time, the server (100) according to one embodiment of the present invention can extract features from the image using a vision-language model.
[0071] For example, the server (100) can generate a first section (Section 1) from an image and generate a first section feature (Section 1 Feature) from this. Furthermore, the server (100) can generate a first section vector (S1F) using the first section feature (Section 1 Feature). Of course, the server (100) can generate section vectors for the remaining sections through the same process.
[0072]
[0073] A server (100) according to one embodiment of the present invention can generate image analysis data based on the similarity between an interval vector and one or more event vectors in a latent space.
[0074] FIG. 5 is a diagram illustrating an example of a graph representing image analysis data generated by a server (100) according to one embodiment of the present invention. In FIG. 5, the Time axis is an axis that lists multiple sections in time series (however, in FIG. 5, the multiple sections are not displayed to be distinct from each other), and the Intensity axis is an axis that indicates the intensity of an event in each section.
[0075] According to one embodiment of the present invention, the server (100) can determine the event intensity of one or more events for each of a plurality of sections constituting an image based on the similarity between the section vector for each of the plurality of sections and one or more event vectors. For example, as illustrated in FIG. 5, the server (100) can determine the event intensity for each section.
[0076] According to one embodiment of the present invention, the server (100) can verify the content and time point of an event prompt at which the event intensity exceeds a predetermined threshold value (I_th). For example, the server (100) can verify that the event intensity of Event Prompt 1 exceeds a predetermined threshold value (I_th) for the T5 section (time point).
[0077] In an optional embodiment of the present invention, a predetermined threshold value may be set differently for each event prompt. For example, the threshold values for Event Prompt 1 and Event Prompt 2 may be set differently, taking into account the content and importance of the event.
[0078] According to one embodiment of the present invention, the server (100) may group points in time at which an event intensity exceeds a predetermined threshold value to generate one or more upper event groups. Furthermore, according to one embodiment of the present invention, the server (100) may generate upper event information of an event group based on the content of an event prompt corresponding to each of one or more points in time within the same group and the occurrence order of one or more points in time, for each of the one or more generated upper event groups.
[0079]
[0080] Figure 6 is a diagram illustrating an exemplary event group.
[0081] A server (100) according to one embodiment of the present invention can group a series of individual events to create a higher event group, and can create higher event information based on the content, duration, and occurrence order of individual events belonging to the created higher event group.
[0082] For example, the server (100) can group individual events of running (Event 2), falling (Event N), running (Event 2), running (Event 2), and falling (Event 1) into one upper event group (Event Group X) and generate upper event information of the group as 'pursuit'.
[0083] In this way, the server (100) according to one embodiment of the present invention can derive higher-level event information based on a combination of a series of individual events.
[0084] According to one embodiment of the present invention, a server (100) can generate video analysis data including the event intensity of each of one or more events, a higher event group, and the content of the higher event for each higher event group for each of the plurality of sections generated according to the above-described process. The generated video analysis data can be provided to a user terminal (200) or the like according to the process described below.
[0085]
[0086] According to one embodiment of the present invention, a server (100) may provide image analysis results based on image analysis data. For example, the server (100) may transmit image analysis results to a user terminal (200).
[0087]
[0088] FIG. 7 is a drawing illustrating an exemplary image analysis result provision screen (600) displayed on a user terminal (200).
[0089] According to one embodiment of the present invention, the server (100) may provide a first region (610) that shows the event intensity of each of one or more events for each of a plurality of sections in a time series manner through the analysis screen (600). At this time, the event intensity may be calculated based on the similarity between the section vector for each of the plurality of sections and one or more event vectors, as described above. For example, the server (100) may provide the event intensity (611) for each section, i.e., the time flow of the first prompt (Prompt1), in the form of a graph, as illustrated in FIG. 7. Of course, the server (100) may also provide the event intensity for the remaining prompts (Prompt2, Prompt3) in the form of a graph according to the time flow.
[0090] According to one embodiment of the present invention, the server (100) may further provide an entity (612) corresponding to a predetermined reference value that serves as a criterion for determining whether an event has occurred. In an optional embodiment of the present invention, if the degree of an event occurrence is determined in multiple stages, there may be multiple entities (612). However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0091]
[0092] According to one embodiment of the present invention, the server (100) may collect points in time at which the event intensity exceeds a predetermined threshold for each of one or more event prompts through the analysis screen (600) and provide a second area (620) that displays the points in time series for each of one or more events. For example, the server (100) may provide points in time at which the event intensity of the first prompt (Prompt1) exceeds a predetermined threshold as objects (621, 622) together with the identification information of the first prompt (Prompt1). In this case, of course, the event intensity may be calculated based on the similarity between the section vector for each of a plurality of sections and one or more event vectors.
[0093]
[0094] A server (100) according to one embodiment of the present invention can collect points in time when an event intensity exceeds a predetermined threshold value through an analysis screen (600) and provide a third area (630) that displays events in a time series based on the time of occurrence of the event.
[0095] For example, in the third area (630), individual events corresponding to the time points displayed in the second area (620) may be expressed in the form of objects corresponding to each time point and event. For example, objects (631, 632) corresponding to the first prompt (Prompt1) may be expressed in order based on the time point of event occurrence. However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0096]
[0097] FIG. 8 is a drawing illustrating an exemplary image analysis result provision screen (700) displayed on a user terminal (200).
[0098] A server (100) according to one embodiment of the present invention can provide an image analysis result provision screen (700) including an area (710) where an event image is displayed, an area (720) where an event list is displayed, and an area (730) where a slider bar is displayed, such as a screen (700).
[0099] A server (100) according to one embodiment of the present invention may provide an event list in which one or more event details are displayed via an area (720). For example, a server (100) according to one embodiment of the present invention may provide a list of upper event groups generated according to the process described in FIG. 6.
[0100] According to one embodiment of the present invention, the server (100) may provide an event video corresponding to a first event selected from an event list displayed in an area (720) in an area (710). In addition, the server (100) may provide a slider bar in an area (730) that indicates the relative position of individual frames provided in an area (710) within the event video according to playback of the event video and controls the displayed frame according to a user's input. For example, when a user selects a first event (721) from the list, the server (100) may provide an event video corresponding to the first event (721) in an area (710) and provide a slider bar for controlling the event video in an area (730).
[0101] According to one embodiment of the present invention, the server (100) may provide a first object (736) corresponding to at least a portion of an event image to be displayed. In addition, the server (100) may provide one or more thumbnails (731, 732, 733, 734) to be displayed on the first object (736). In this case, the server (100) may provide a frame at a point in time when the event intensity exceeds a predetermined threshold to be displayed as one or more thumbnails (731, 732, 733, 734).
[0102] According to one embodiment of the present invention, a server (100) may provide an object representing an event prompt related to one or more thumbnails (731, 732, 733, 734) to be displayed in association with the one or more thumbnails. For example, the server (100) may display a text object such as "Prompt 1 Scene" in association with (e.g., overlapping) the first thumbnail (731).
[0103] In addition, the server (100) according to one embodiment of the present invention may provide an object (735) representing upper event group information including an event corresponding to each of one or more thumbnails (731, 732, 733, 734) so as to be displayed in association with one or more thumbnails (731, 732, 733, 734). For example, as illustrated in FIG. 8, the server (100) may provide one or more thumbnails (731, 732, 733, 734) and an object (735) representing upper event group information for Event Group 1 in correspondence with each other on a slider bar.
[0104] However, this display format is exemplary, and any display format that can display the relationship between one or more thumbnails (731, 732, 733, 734) and an object (735) representing upper event group information may be used without limitation.
[0105] Meanwhile, the text on the object (735) representing the upper event group information may be text representing the content of the corresponding upper event. For example, if one or more thumbnails (731, 732, 733, 734) each relate to running, running, falling, and falling, the text on the object (735) representing the upper event group information may be "pursuit," which is a higher concept of the event. However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0106] According to one embodiment of the present invention, the server (100) may provide an image corresponding to a point in time of one or more thumbnails (731, 732, 733, 734) to be displayed in an area (710) based on a user's selection of the thumbnail. In addition, according to one embodiment of the present invention, the server (100) may provide an image corresponding to the first event among one or more events belonging to the upper event group to be displayed in an area (710) based on a user's selection of an object (735) representing upper event group information.
[0107] In an optional embodiment of the present invention, only one or more thumbnails (731, 732, 733, 734) representing upper event group information may be provided to be sequentially displayed in the area (710) over time, based on the user's selection of an object (735). In this case, each of the one or more thumbnails (731, 732, 733, 734) may be a frame at a point in time when the event intensity exceeds a predetermined threshold value, as described above.
[0108] Accordingly, the present invention provides an event video so that the user can view the main events within the event video at once through a slider bar without having to view the entire video, and also allows the user to review events by event group.
[0109]
[0110] FIG. 9 is a flowchart illustrating an event detection method based on a text prompt performed by a server (100) according to one embodiment of the present invention. The method is described below with reference to FIGS. 1 to 8.
[0111] A server (100) according to one embodiment of the present invention can generate an event vector, which is a vector in a latent space for each of one or more event prompts defined in natural language. (S910)
[0112] FIG. 4 is a diagram illustrating a process in which a server (100) generates a vector from an event prompt and an image according to one embodiment of the present invention.
[0113] A server (100) according to one embodiment of the present invention can extract features from a prompt and generate an event vector, which is a vector in a latent space, based on the extracted features. At this time, the server (100) according to one embodiment of the present invention can extract features using various text models.
[0114] For example, the server (100) can extract the first event feature (Event Feature 1) from the first event prompt (Event Prompt 1) and use it to generate the first event vector (EV1). Of course, the server (100) can also generate event vectors for the remaining event prompts through the same process.
[0115]
[0116] A server (100) according to one embodiment of the present invention can generate a section vector for each of a plurality of sections constituting an image.
[0117] More specifically, the server (100) according to one embodiment of the present invention divides an image into a plurality of sections, extracts features for each of the divided sections (S920), and generates a section vector, which is a vector in a latent space for each of the plurality of sections, based on each of the extracted features (S930). At this time, the server (100) according to one embodiment of the present invention can extract features from the image using a vision-language model.
[0118] For example, the server (100) can generate a first section (Section 1) from an image and generate a first section feature (Section 1 Feature) from this. Furthermore, the server (100) can generate a first section vector (S1F) using the first section feature (Section 1 Feature). Of course, the server (100) can generate section vectors for the remaining sections through the same process.
[0119]
[0120] A server (100) according to one embodiment of the present invention can generate image analysis data based on the similarity between an interval vector and one or more event vectors in a latent space (S940).
[0121] FIG. 5 is a diagram illustrating an example of a graph representing image analysis data generated by a server (100) according to one embodiment of the present invention. In FIG. 5, the Time axis is an axis that lists multiple sections in time series (however, in FIG. 5, the multiple sections are not displayed to be distinct from each other), and the Intensity axis is an axis that indicates the intensity of an event in each section.
[0122] According to one embodiment of the present invention, the server (100) can determine the event intensity of one or more events for each of a plurality of sections constituting an image based on the similarity between the section vector for each of the plurality of sections and one or more event vectors. For example, as illustrated in FIG. 5, the server (100) can determine the event intensity for each section.
[0123] According to one embodiment of the present invention, the server (100) can verify the content and time point of an event prompt at which the event intensity exceeds a predetermined threshold value (I_th). For example, the server (100) can verify that the event intensity of Event Prompt 1 exceeds a predetermined threshold value (I_th) for the T5 section (time point).
[0124] In an optional embodiment of the present invention, a predetermined threshold value may be set differently for each event prompt. For example, the threshold values for Event Prompt 1 and Event Prompt 2 may be set differently, taking into account the content and importance of the event.
[0125] According to one embodiment of the present invention, the server (100) may group points in time at which an event intensity exceeds a predetermined threshold value to generate one or more upper event groups. Furthermore, according to one embodiment of the present invention, the server (100) may generate upper event information of an event group based on the content of an event prompt corresponding to each of one or more points in time within the same group and the occurrence order of one or more points in time, for each of the one or more generated upper event groups.
[0126]
[0127] Figure 6 is a diagram illustrating an exemplary event group.
[0128] A server (100) according to one embodiment of the present invention can group a series of individual events to create a higher event group, and can create higher event information based on the content, duration, and occurrence order of individual events belonging to the created higher event group.
[0129] For example, the server (100) can group individual events of running (Event 2), falling (Event N), running (Event 2), running (Event 2), and falling (Event 1) into one upper event group (Event Group X) and generate upper event information of the group as 'pursuit'.
[0130] In this way, the server (100) according to one embodiment of the present invention can derive higher-level event information based on a combination of a series of individual events.
[0131] According to one embodiment of the present invention, a server (100) can generate video analysis data including the event intensity of each of one or more events, a higher event group, and the content of the higher event for each higher event group for each of the plurality of sections generated according to the above-described process. The generated video analysis data can be provided to a user terminal (200) or the like according to the process described below.
[0132]
[0133] A server (100) according to one embodiment of the present invention can provide an image analysis result based on image analysis data. (S950) For example, the server (100) can transmit an image analysis result to a user terminal (200).
[0134]
[0135] FIG. 7 is a drawing illustrating an exemplary image analysis result provision screen (600) displayed on a user terminal (200).
[0136] According to one embodiment of the present invention, the server (100) may provide a first region (610) that shows the event intensity of each of one or more events for each of a plurality of sections in a time series manner through the analysis screen (600). At this time, the event intensity may be calculated based on the similarity between the section vector for each of the plurality of sections and one or more event vectors, as described above. For example, the server (100) may provide the event intensity (611) for each section, i.e., the time flow of the first prompt (Prompt1), in the form of a graph, as illustrated in FIG. 7. Of course, the server (100) may also provide the event intensity for the remaining prompts (Prompt2, Prompt3) in the form of a graph according to the time flow.
[0137] According to one embodiment of the present invention, the server (100) may further provide an entity (612) corresponding to a predetermined reference value that serves as a criterion for determining whether an event has occurred. In an optional embodiment of the present invention, if the degree of an event occurrence is determined in multiple stages, there may be multiple entities (612). However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0138]
[0139] According to one embodiment of the present invention, the server (100) may collect points in time at which the event intensity exceeds a predetermined threshold for each of one or more event prompts through the analysis screen (600) and provide a second area (620) that displays the points in time series for each of one or more events. For example, the server (100) may provide points in time at which the event intensity of the first prompt (Prompt1) exceeds a predetermined threshold as objects (621, 622) together with the identification information of the first prompt (Prompt1). In this case, of course, the event intensity may be calculated based on the similarity between the section vector for each of a plurality of sections and one or more event vectors.
[0140]
[0141] A server (100) according to one embodiment of the present invention can collect points in time when an event intensity exceeds a predetermined threshold value through an analysis screen (600) and provide a third area (630) that displays events in a time series based on the time of occurrence of the event.
[0142] For example, in the third area (630), individual events corresponding to the time points displayed in the second area (620) may be expressed in the form of objects corresponding to each time point and event. For example, objects (631, 632) corresponding to the first prompt (Prompt1) may be expressed in order based on the time point of event occurrence. However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0143]
[0144] FIG. 8 is a drawing illustrating an exemplary image analysis result provision screen (700) displayed on a user terminal (200).
[0145] A server (100) according to one embodiment of the present invention can provide an image analysis result provision screen (700) including an area (710) where an event image is displayed, an area (720) where an event list is displayed, and an area (730) where a slider bar is displayed, such as a screen (700).
[0146] A server (100) according to one embodiment of the present invention may provide an event list in which one or more event details are displayed via an area (720). For example, a server (100) according to one embodiment of the present invention may provide a list of upper event groups generated according to the process described in FIG. 6.
[0147] According to one embodiment of the present invention, the server (100) may provide an event video corresponding to a first event selected from an event list displayed in an area (720) in an area (710). In addition, the server (100) may provide a slider bar in an area (730) that indicates the relative position of individual frames provided in an area (710) within the event video according to playback of the event video and controls the displayed frame according to a user's input. For example, when a user selects a first event (721) from the list, the server (100) may provide an event video corresponding to the first event (721) in an area (710) and provide a slider bar for controlling the event video in an area (730).
[0148] According to one embodiment of the present invention, the server (100) may provide a first object (736) corresponding to at least a portion of an event image to be displayed. In addition, the server (100) may provide one or more thumbnails (731, 732, 733, 734) to be displayed on the first object (736). In this case, the server (100) may provide a frame at a point in time when the event intensity exceeds a predetermined threshold to be displayed as one or more thumbnails (731, 732, 733, 734).
[0149] According to one embodiment of the present invention, a server (100) may provide an object representing an event prompt related to one or more thumbnails (731, 732, 733, 734) to be displayed in association with the one or more thumbnails. For example, the server (100) may display a text object such as "Prompt 1 Scene" in association with (e.g., overlapping) the first thumbnail (731).
[0150] In addition, the server (100) according to one embodiment of the present invention may provide an object (735) representing upper event group information including an event corresponding to each of one or more thumbnails (731, 732, 733, 734) so as to be displayed in association with one or more thumbnails (731, 732, 733, 734). For example, as illustrated in FIG. 8, the server (100) may provide one or more thumbnails (731, 732, 733, 734) and an object (735) representing upper event group information for Event Group 1 in correspondence with each other on a slider bar.
[0151] However, this display format is exemplary, and any display format that can display the relationship between one or more thumbnails (731, 732, 733, 734) and an object (735) representing upper event group information may be used without limitation.
[0152] Meanwhile, the text on the object (735) representing the upper event group information may be text representing the content of the corresponding upper event. For example, if one or more thumbnails (731, 732, 733, 734) each relate to running, running, falling, and falling, the text on the object (735) representing the upper event group information may be "pursuit," which is a higher concept of the event. However, this is merely exemplary and the scope of the present invention is not limited thereto.
[0153] According to one embodiment of the present invention, the server (100) may provide an image corresponding to a point in time of one or more thumbnails (731, 732, 733, 734) to be displayed in an area (710) based on a user's selection of the thumbnail. In addition, according to one embodiment of the present invention, the server (100) may provide an image corresponding to the first event among one or more events belonging to the upper event group to be displayed in an area (710) based on a user's selection of an object (735) representing upper event group information.
[0154] In an optional embodiment of the present invention, only one or more thumbnails (731, 732, 733, 734) representing upper event group information may be provided to be sequentially displayed in the area (710) over time, based on the user's selection of an object (735). In this case, each of the one or more thumbnails (731, 732, 733, 734) may be a frame at a point in time when the event intensity exceeds a predetermined threshold value, as described above.
[0155] Accordingly, the present invention provides an event video so that the user can view the main events within the event video at once through a slider bar without having to view the entire video, and also allows the user to review events by event group.
[0156]
[0157] The embodiments of the present invention described above may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may be a medium that stores a program that can be executed by a computer. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and media configured to store program instructions, including ROMs, RAMs, and flash memories.
[0158] Meanwhile, the computer program may be specifically designed and constructed for the present invention, or may be one known and available to those skilled in the computer software field. Examples of computer programs may include not only machine language code, such as that generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
[0159] The specific implementations described in the present invention are exemplary embodiments and do not limit the scope of the present invention in any way. For the sake of brevity, descriptions of conventional electronic components, control systems, software, and other functional aspects of the systems may be omitted. In addition, the lines connecting or connecting members between components depicted in the drawings are merely representative of functional connections and / or physical or circuit connections, and may be replaced or represented as various additional functional connections, physical connections, or circuit connections in an actual device. In addition, unless specifically mentioned as "essential," "important," etc., a component may not be absolutely necessary for the application of the present invention.
[0160] Therefore, the idea of the present invention should not be limited to the embodiments described above, and not only the scope of the patent claims described below but also all scopes equivalent to or equivalently modified from the scope of the patent claims are considered to fall within the scope of the idea of the present invention.
Claims
1. In an event detection method based on a text prompt, A step of generating an event vector, which is a vector in the latent space for each of one or more event prompts defined in natural language; A step of extracting features for each of a plurality of sections constituting an image; A step of generating an interval vector, which is a vector in the latent space for each of the plurality of intervals, based on each of the extracted features; A step of generating image analysis data based on the similarity between the interval vector and one or more event vectors in the latent space for each of the plurality of intervals; and An event detection method based on a text prompt, comprising: a step of providing an analysis result of the image based on the image analysis data; 2. In claim 1 The step of generating the above image analysis data is A step of determining an event intensity of each of one or more events for each of the plurality of intervals based on a similarity between an interval vector for each of the plurality of intervals and one or more event vectors; A step of grouping points at which the event intensity exceeds a predetermined threshold to create one or more upper event groups; and An event detection method based on a text prompt, comprising: a step of generating upper event information of an upper event group based on the contents of an event prompt corresponding to each of one or more time points belonging to the same group for each of one or more upper event groups and the occurrence order of the one or more time points; 3. In claim 1 The step of providing the above analysis results is A method for detecting events based on a text prompt, comprising: a step of providing an analysis screen including a first region that time-seriesly represents the event intensity of each of one or more events for each of the plurality of intervals, wherein the event intensity is calculated based on a similarity between an interval vector for each of the plurality of intervals and one or more event vectors; 4. In claim 1 The step of providing the above analysis results is A method for detecting events based on text prompts, comprising: a step of providing an analysis screen including a second area that collects points in time when an event intensity exceeds a predetermined threshold for each of the one or more event prompts and displays them in a time series for each of the one or more events, wherein the event intensity is calculated based on a similarity between an interval vector for each of the plurality of intervals and one or more event vectors.
5. In claim 1 The step of providing the above analysis results is A method for detecting an event based on a text prompt, comprising: a step of collecting points in time when an event intensity exceeds a predetermined threshold and providing an analysis screen including a third area that represents events in a time series based on the point in time when the event occurs, wherein the event intensity is calculated based on a similarity between a section vector for each of the plurality of sections and one or more event vectors.
6. In claim 1 The step of providing the above analysis results is A step for providing a list of events in which one or more event details are displayed; A step of providing an event image corresponding to a first event selected from the above event list in a fourth area; and A step of providing a slider bar in a fifth area that indicates the relative position of individual frames displayed in the fourth area within the event image according to playback of the event image, and controls the displayed frame according to a user's input; The step of providing the above slider bar is A step of providing a first object corresponding to at least a portion of the event image to be displayed; and A step of providing one or more thumbnails to be displayed on the first object, comprising: a step of providing a frame at a point in time when an event intensity exceeds a predetermined threshold to be displayed as the thumbnail; An event detection method based on a text prompt, wherein the event intensity is calculated based on the similarity between a segment vector for each of one or more segments constituting the event image and one or more event vectors.
7. In claim 6 The step of providing the above slider bar is A method for detecting an event based on a text prompt, further comprising: providing an object representing an event prompt associated with one or more thumbnails so that the object is displayed in association with the one or more thumbnails.
8. In claim 6 The step of providing the above slider bar is A method for detecting events based on a text prompt, further comprising: providing an object representing information about a higher event group, each of which includes an event corresponding to one or more thumbnails, so that the object is displayed in association with the one or more thumbnails.
9. In claim 8 The step of providing the above slider bar is A step of providing an image corresponding to a point in time of a thumbnail to be displayed in the fourth area according to a user's selection of one or more of the thumbnails; and A method for detecting an event based on a text prompt, comprising: providing an image corresponding to the first event among one or more events belonging to the upper event group so that the image is displayed in the fourth area according to a user's selection of an object representing the upper event group information.
10. A computer program stored on a medium for executing any one of the methods of clauses 1 to 9 using a computer.
Citation Information
Patent Citations
Content processing device, method, and program
KR1020130038820A
Metal-organic framework (MOFs) nanoparticles coated with a fusion protein of Glutathione S-transferase and peptides targeting diseased cells and use thereof
KR1020240086776A
Ultrasonic humidifier
KR1020250024440A
Refill type lipstick container
KR1020250037876A
KR20230085103A