Information processing device, method, and program

The information processing device enhances event-related object recognition by selecting and analyzing images for event detection areas and object positions, improving accuracy in identifying relevant objects within event images.

WO2026038442A1PCT designated stage Publication Date: 2026-02-19NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/026022
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-14
Filing Date
2025-07-23
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing technologies lack sufficient accuracy in identifying objects related to events from a set of multiple images that include the time period when the event occurred.

Method used

An information processing device and method that selects images where an event has been detected, estimates the area within the image contributing to the event detection, and identifies the position of objects to enhance the accuracy of event-related object recognition, using a combination of image analysis techniques and trained models.

Benefits of technology

Improves the accuracy of identifying objects related to events by utilizing the area and position of objects within detected images, facilitating effective event analysis and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025026022_19022026_PF_FP_ABST
    Figure JP2025026022_19022026_PF_FP_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to improve the accuracy with which an object related to an event is specified within an image. This information processing device comprises a selection means that selects an image in which a prescribed event has been detected from a plurality of image sequences captured during a prescribed period, a region estimation means that estimates a region that contributed to the detection of the event in the selected image, an object estimation means that estimates the positions of objects included in at least the selected image from among the plurality of image sequences, and a specifying means that specifies an event-related object that is related to the event on the basis of the region and the positions of the objects estimated from the selected image.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, method, and program

[0001] The present disclosure relates to an information processing device, a method, and a program.

[0002] In recent years, there has been a demand for technology that analyzes events such as accidents by analyzing video data.

[0003] For example, Patent Literature 1 discloses a technology related to an image processing device that detects objects and events from images. The image processing device detects a person (object) from image data, and performs tracking processing on other image data using the person detection result to detect events such as leaving a bag behind or being taken away.

[0004] JP 2011-91705 A

[0005] However, the above-mentioned technology has a problem in that it does not have sufficient accuracy in identifying an object related to an event from a set of multiple images that include the time period when the event occurred.

[0006] In view of the above-described problems, an object of the present disclosure is to provide an information processing device, a method, and a program for improving the accuracy of identifying an object related to an event in an image.

[0007] The information processing device according to the present disclosure comprises a selection means for selecting an image in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period of time; an area estimation means for estimating an area within the selected image that contributed to the detection of the event; an object estimation means for estimating a position of an object included in at least the selected image from the sequence of multiple images; and an identification means for identifying an event-related object related to the event based on the position and area of ​​the object estimated from the selected image.

[0008] The information processing method according to the present disclosure includes a computer selecting images in which a predetermined event has been detected from a sequence of multiple images taken over a predetermined period of time, estimating an area within the selected images that contributed to the detection of the event, estimating a position of an object included in at least the selected images from the sequence of multiple images, and identifying an event-related object related to the event based on the estimated position and area of ​​the object from the selected images.

[0009] The information processing program according to the present disclosure causes a computer to execute the following processes: a selection process for selecting images in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period; an event estimation process for estimating an area in the selected images that contributed to the detection of the event; an object estimation process for estimating a position of an object included in at least the selected images from the sequence of multiple images; and an identification process for identifying an event-related object related to the event based on the position and area of ​​the object estimated from the selected images.

[0010] The present disclosure allows for improved accuracy in identifying objects associated with an event in an image.

[0011] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flowchart showing a flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 4 is a flowchart showing a flow of an information processing method according to the present disclosure. FIG. 5 is a block diagram showing an overall configuration of an information processing system according to the present disclosure. FIG. 6 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 7 is a flowchart showing a flow of an event-related object identification process according to the present disclosure. FIG. 8 is a diagram showing an example of an event visualization screen according to the present disclosure. FIG. 9 is a diagram showing an example of a seek bar in the event visualization screen according to the present disclosure. FIG. 10 is a diagram showing an example of initial evaluation information displayed in an evaluation information display editing area according to the present disclosure. FIG. 11 is a diagram showing an example of another frame image displayed on the event visualization screen according to the present disclosure. FIG. 12 is a diagram showing an example of another frame image displayed on the event visualization screen according to the present disclosure. FIG. 13 is a diagram for explaining the relationship between event detection, object estimation, and event-related object identification in each image according to the present disclosure. FIG. 14 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 15 is a diagram showing an example of an event visualization screen according to the present disclosure. FIG. 1 is a diagram showing an example of an event visualization screen after a change in the designation of an analysis viewpoint according to the present disclosure. FIG. 2 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 3 is a flowchart showing a flow of an initial evaluation information generation and display process according to the present disclosure. FIG. 4 is a diagram showing an example of a layout of an event visualization screen according to the present disclosure. FIG. 5 is a diagram showing an example of initial evaluation information displayed in a user operation reception area and an evaluation information generation result display area according to the present disclosure. FIG. 6 is a diagram showing an example of automatically generated event information displayed in an event information display and editing area according to the present disclosure.FIG. 1 is a diagram showing an example of case information candidates extracted for the first time and displayed in a case information candidate display selection area according to the present disclosure. FIG. 2 is a diagram showing an example of an input prompt when generating evaluation information for the first time according to the present disclosure. FIG. 3 is a flowchart showing a flow of evaluation information update display processing according to the present disclosure. FIG. 4 is a diagram showing an example of case information candidates re-extracted in response to correction of event information according to the present disclosure. FIG. 5 is a diagram showing an example of an update of an input prompt according to the present disclosure. FIG. 6 is a diagram showing an example of an update of evaluation information according to the present disclosure. FIG. 7 is a block diagram showing a hardware configuration of an information processing device according to the present disclosure.

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and for clarity of explanation, duplicate explanations will be omitted as necessary.

[0013] 1 is a block diagram showing the configuration of an information processing device 1. The information processing device 1 is a computer system that identifies an event-related object related to a predetermined event from a sequence of multiple images captured during a predetermined period. The information processing device 1 includes a selection unit 11, a region estimation unit 12, an object estimation unit 13, and an identification unit 14.

[0014] The selection unit 11 selects images in which a predetermined event has been detected from a sequence of multiple images captured over a predetermined period of time. Note that "a sequence of multiple images captured over a predetermined period of time" refers to a collection of frame images in chronological order contained in video data. Specifically, the selection unit 11 performs a process of detecting a predetermined event for each image in the sequence of images. If a predetermined event is detected, the selection unit 11 selects the detected image. In other words, the selection unit 11 classifies each image in the sequence of images into images in which an event has been detected and images in which an event has not been detected. Note that the process of detecting a predetermined event from any image may be realized by a predetermined image analysis technique or a trained model that uses an image as input and outputs whether or not an event has been detected.

[0015] The region estimation unit 12 estimates a region that contributed to the detection of an event within the image selected by the selection unit 11. Here, the "region that contributed to the detection of an event" can also be said to be a region used to determine event detection in the event detection process. Note that the process of estimating a region that contributed to the detection of an event within an image may be realized by the predetermined image analysis technology or a technology for estimating a gaze region during deep learning of the trained model.

[0016] The object estimation unit 13 estimates the position of an object included in at least the images selected by the selection unit 11 from the plurality of image sequences. Note that the process of estimating the position of an object from an image may be realized by a predetermined image analysis technique or the like.

[0017] The identification unit 14 identifies an event-related object related to the event based on the position of the object estimated by the object estimation unit 13 from the selected image and the area estimated by the area estimation unit 12 from the selected image.

[0018] 2 is a flowchart showing the flow of the information processing method. First, the selection unit 11 selects images in which a predetermined event has been detected from a sequence of multiple images captured during a predetermined period (S1). Next, the region estimation unit 12 estimates regions in the selected images that contributed to the detection of the event (S2). Then, the object estimation unit 13 estimates the position of an object included in at least the selected images from the sequence of multiple images (S3). Thereafter, the identification unit 14 identifies an event-related object related to the event based on the position of the object estimated in step S3 from the images selected in step S1 and the region estimated in step S2 from the images selected in step S1 (S4).

[0019] Note that steps S2 and S3 may be processed in a reversed order or in parallel. Also, step S3 may be processed independently of steps S1 and S2. At least in steps S2 and S3, the region estimation unit 12 and the object estimation unit 13 perform processing on the same image for which an event was selected in step S1.

[0020] Here, in a typical object detection process, a large number of objects may be detected from a particular image. Therefore, it is difficult to narrow down the detected objects to those related to an event. Furthermore, in an event detection process, it is possible to detect an image in which an event occurred from a sequence of multiple images. However, it is difficult to identify which objects in an image are related to the event. Therefore, a user is forced to visually identify event-related objects from the images. Therefore, it is difficult to accurately identify objects related to the event from a set of multiple images that include the time period in which the event occurred. In contrast, in the technology disclosed herein, when an event is detected in a particular image, the area in the image that contributed to the detection of the event and the position where the object was detected in the same image are used. This improves the accuracy of identifying objects related to the event in the image.

[0021] Second Embodiment FIG. 3 is a block diagram showing the configuration of an information processing device 1a. The information processing device 1a includes a display control unit 15a. When a sequence of multiple images captured during a predetermined period is played back in chronological order and displayed on a predetermined screen, the display control unit 15a displays an area corresponding to an event-related object related to an event occurring in the image in a identifiable manner within the image. Here, an "event-related object" refers to an object identified based on an area that contributed to the detection of a predetermined event in an image from a sequence of multiple images captured during a predetermined period. Alternatively, an "event-related object" may refer to an object related to an event identified based on an area that contributed to the detection of the event in the image. Alternatively, an "event-related object" may refer to an object related to an event identified based on an area that contributed to the detection when the predetermined event is detected in the image. For example, an "event-related object" may be identified using the information processing device 1 according to the first embodiment. Furthermore, a "screen" refers to a display area of ​​a display device built into the information processing device 1a or a display device connected to the information processing device 1a (e.g., via a communication network). Furthermore, "displaying an area in an identifiable manner within an image" may mean, for example, displaying the relevant area within the image by enclosing it, or displaying the relevant area or its borders in a color different from the surrounding area, etc. However, the method of "displaying an area in an identifiable manner within an image" is not limited to these.

[0022] The display control unit 15a also displays on the screen an event area bar indicating an event section in which an event was detected during a predetermined period. Here, the "event section" refers to, for example, a set of playback positions within a predetermined period corresponding to one or more images selected by the selection unit 11 according to the first embodiment. The "event section" refers to a time period corresponding to a sequence of chronologically consecutive images selected by the selection unit 11. The "event area bar" is display information indicating the length of the event section compared to the length of the predetermined period. For example, when the predetermined period is displayed as one-dimensional bar-shaped display information, the "event area bar" may be one that displays the event section as one-dimensional bar-shaped display information with a length shorter than the predetermined period. However, the "event area bar" is not limited to these.

[0023] 4 is a flowchart showing the flow of the information processing method. The display control unit 15a displays, on a predetermined screen, an area corresponding to an event-related object in a manner that allows it to be identified within an image in which the event has been detected, and displays, on the screen, an event area bar indicating an event section in which the event has been detected during a predetermined period (S5a). As described above, an "event-related object" is an object related to a predetermined event that has been identified based on an area that contributed to the detection of the event in an image in which the event has been detected among a sequence of multiple images captured during a predetermined period. The display control unit 15a may display the area corresponding to the event-related object and the event area bar at different times. Therefore, the display order of the area corresponding to the event-related object and the event area bar is not limited.

[0024] In this way, the technology disclosed herein displays the area of ​​an event-related object on a screen in a manner that allows it to be identified within an image, and displays an event section on the screen. Therefore, when playing back and displaying video data on a screen, a user can easily identify the area related to the event within a frame image of the event scene. Furthermore, when playing back and displaying video data on a screen, a user can easily recognize the event section within a predetermined period of time when the video data was captured by using the event area bar. Therefore, when playing back and displaying video data including the time before and after an event that occurred, analysis of the event can be effectively supported.

[0025] Third Embodiment FIG. 5 is a block diagram showing the configuration of an information processing device 1b. The information processing device 1b includes a display control unit 15b. The display control unit 15b displays, on a predetermined screen, event information including an event-related object, case information indicating examples related to the event, and evaluation information related to the event. Here, the "event-related object" refers to an object related to an event identified based on an area that contributed to the detection of the event in an image. Alternatively, the "event-related object" may refer to an object related to an event identified based on an area that contributed to the detection when the event is detected from an image. Alternatively, the "event-related object" may be an object identified based on an area that contributed to the detection in an image in which the event is detected from a sequence of multiple images captured during a predetermined period. For example, the "event-related object" may be identified using the information processing device 1 according to the first embodiment. Furthermore, the "evaluation information" refers to evaluation information related to the event output from a predetermined trained model using input information based on the event information and case information. Furthermore, the "screen" is a display area of ​​a display device built into the information processing device 1b or a display device connected to the information processing device 1b (for example, via a communication network).

[0026] 6 is a flowchart showing the flow of the information processing method. The display control unit 15b displays, on a predetermined screen (S5b), event information including an event-related object, case information indicating cases related to the event, and evaluation information related to the event. As described above, the "event-related object" is, for example, an object related to the event identified based on an area that contributed to the detection of the event in an image. As described above, the "evaluation information" is information evaluating the event, output from a predetermined trained model using input information based on the event information and case information.

[0027] Generally, various data sources are required to accurately grasp the circumstances of an event such as an accident. However, it may be difficult to acquire a wide variety of data sources related to an event. In response to this, the technology disclosed herein uses images from the data sources in which the event was detected or a sequence of images captured during a predetermined period including the time when the event occurred. The technology disclosed herein uses these images to display, on a predetermined screen, event information including event-related objects, case information showing examples related to the event, and evaluation information related to the event. Therefore, by presenting highly accurate determination results, etc., related to the event that occurred, generated using event information obtained within the constraints of the data sources, it is possible to assist in grasping the circumstances of the event.

[0028] 7 is a block diagram showing the overall configuration of an information processing system 1000. The information processing system 1000 includes an information processing device 100 and a display terminal 200. The information processing device 100 and the display terminal 200 are communicably connected to each other via a communication network N. Here, the communication network N is a communication line network that may be wired or wireless.

[0029] The display terminal 200 is an information processing device operated by a user and has a screen 201. In response to a user's operation, the display terminal 200 transmits information related to the playback of video data and event analysis to the information processing device 100 via the communication network N. The display terminal 200 then receives images, display information, etc. from the information processing device 100 via the communication network N. The display terminal 200 displays the received images, display information, etc. on the screen 201. The screen 201 is, for example, a display area of ​​a liquid crystal display, an organic EL (organic electroluminescence) display, etc.

[0030] FIG. 8 is a block diagram showing the configuration of an information processing device 100. The information processing device 100 is an example of the information processing devices 1, 1a, and 1b. The information processing device 100 may be realized as a computer system in which functions are distributed or made redundant by a plurality of computer devices. The information processing device 100 includes a storage unit 110, an event detection unit 121, an area estimation unit 122, an object detection and tracking unit 123, an identification unit 124, an acquisition unit 125, a display control unit 126, and a reception unit 127. The event detection unit 121 is an example of the selection unit 11 described above. The area estimation unit 122 is an example of the area estimation unit 12 described above. The object detection and tracking unit 123 is an example of the object estimation unit 13 described above. The identification unit 124 is an example of the identification unit 14 described above. The display control unit 126 is an example of the display control unit 15a or 15b described above.

[0031] The storage unit 110 includes, for example, a non-volatile storage device such as a hard disk or flash memory, and a memory such as a RAM (Random Access Memory), i.e., a volatile storage device. The storage unit 110 stores video data 111, event-related object information 1121, ... 112m (m is a natural number equal to or greater than 1), and evaluation information 113.

[0032] The video data 111 includes a sequence of multiple images captured during a predetermined period of time that includes the period during which the event to be detected occurred. The video data 111 is, for example, video captured by an in-vehicle camera or the like. Specifically, the video data 111 includes images 1111, ... 111n (n is a natural number equal to or greater than 2). Each of the images 1111, etc., is associated with at least the capture time of the frame image. Therefore, the images 1111, etc., are an image sequence that can be played and displayed in the order of their capture times, that is, in chronological order, or can be played and displayed from any specified time.

[0033] Each piece of event-related object information 1121, etc. is information relating to a different event-related object. Each piece of event-related object information 1121, etc. includes identification information of the event-related object, the type of object, identification information of the image in which the object is detected, and the position within the image in which the object is detected (information indicating an area or a group of two-dimensional coordinates). Furthermore, each piece of event-related object information 1121, etc. is information relating to the same event-related object detected in one or more images within the video data 111. Therefore, each piece of event-related object information 1121, etc. includes a set of identification information of one or more images and their positions within the corresponding images.

[0034] The evaluation information 113 is an example of the above-mentioned "evaluation information related to an event." The evaluation information 113 includes text data such as a sentence explaining the content of the detected event and a sentence evaluating the event. For example, if the event is a traffic accident, the evaluation information 113 may be a sentence explaining the fault ratio or the accident situation. Alternatively, if the event is an optional task, the evaluation information 113 may be a sentence reporting the work. Note that the evaluation information 113 is not limited to these. Specific examples of the evaluation information 113 will be described later.

[0035] The event detection unit 121 performs an event detection process for each image in the video data 111. Then, the event detection unit 121 selects an image in which an event has been detected. For example, the event detection unit 121 calculates a score indicating the likelihood of an event occurring for the image to be processed by image analysis processing or the like, and selects the image as an image in which an event has been detected if the score is equal to or greater than a threshold. Here, the event detection unit 121 may use a first trained model that calculates and outputs a score indicating the likelihood of an event occurring for an input image. The first trained model may be, for example, an AI (Artificial Intelligence) model trained using a technique based on Action Segmentation or Video Moment Retrieval. The first trained model may be, for example, a model trained by machine learning using a dataset in which each image of the video data has been annotated with temporal, spatial, and categorical annotations. Alternatively, the first trained model may input a sequence of images included in the video data 111, calculate a score for each image, and output whether or not the score is equal to or greater than a threshold, that is, whether or not an event has been detected, for each image. In this case, the event detection unit 121 may select the input image when the first trained model outputs that an event has been detected as the selected image.

[0036] The region estimation unit 122 estimates an event area, which is a region that contributed to the detection of an event within the image selected by the event detection unit 121. For example, the region estimation unit 122 estimates a region that contributed to the determination that an event occurred when the event detection unit 121 detected an event. Specifically, the region estimation unit 122 may estimate, as the event area, a region corresponding to a set of pixel information that the event detection unit 121 referenced when calculating a score equal to or greater than a threshold for the image to be processed. For example, the region estimation unit 122 may use, for example, attention map technology, distribution information on the degree of attention for each pixel that the first trained model used by the event detection unit 121 focused on when calculating the score. Therefore, the region estimation unit 122 may estimate, as the event area, a set of pixels (heat map) whose degree of attention is equal to or greater than a threshold in the distribution information obtained from the image to be processed. Note that the region estimation unit 122 may estimate the event area using a second trained model to which attention map technology is applied. The second trained model may be, for example, a model that inputs an image in which an event is detected by the event detection unit 121 and outputs an event area. Alternatively, the second trained model may be, for example, a model that is machine-learned using a pair of an image and an event area as training data.

[0037] The object detection and tracking unit 123 performs object detection processing on each image in the video data 111 and object tracking processing on the image sequence. As a result, the object detection and tracking unit 123 also performs object detection processing on images in which an event has been detected by the event detection unit 121. Specifically, the object detection and tracking unit 123 detects one or more objects from each image by performing object recognition processing or the like using image analysis technology. That is, the object detection and tracking unit 123 outputs a set of position coordinates, an area, etc. of objects detected in an image. Furthermore, the object detection and tracking unit 123 performs object tracking processing on the image sequence for a specific detected object and outputs the position of the specific object in each image, etc.

[0038] The identification unit 124 identifies an event-related object by comparing the area of ​​the event area estimated from the selected image with the area of ​​the object estimated from the same selected image. Specifically, the identification unit 124 identifies the event-related object based on the overlapping range between the position of the object detected by the object detection and tracking unit 123 and the area of ​​the event area estimated by the area estimation unit 122 for the same selected image. At this time, the identification unit 124 generates event-related object information related to the identified event-related object. Then, the identification unit 124 stores the generated event-related object information in the storage unit 110.

[0039] Here, the event detection unit 121 may select two or more images from a sequence of multiple images as images in which an event has been detected. Among the two or more selected images, a first region of the event may be estimated in a first image, but the location of the object may not be estimated. For example, if the event is a car collision, the event may be detected in multiple images taken during the time period in which the collision occurred. However, object detection may not be possible from the images taken at the time of the collision because the shape of the target object (car) may have changed or the cars may be overlapping each other. In other words, in such cases, the region estimation unit 122 may be able to estimate a first region as an event area from the first image, but the object detection and tracking unit 123 may be unable to detect the object (car) from the first image and therefore unable to estimate the location of the car. Even in such cases, the event detection unit 121 may be able to detect the event in both the first image and the second image. The region estimation unit 122 may be able to estimate a second region as an event area from the second image. Therefore, in such a case, the identification unit 124 may identify the event-related object included in the first image based on the second region and the object positions of the event estimated in the second image, and the first region, thereby improving the accuracy of identifying the event-related object.

[0040] Furthermore, the object detection and tracking unit 123 tracks a specific object in two or more images selected from the multiple image sequences, thereby estimating the position of the object in each image. Here, it is possible that a third area of ​​the event is estimated in a third image among the two or more selected images, but the position of the object is not estimated. In such a case, the identification unit 124 may identify an event-related object included in the third image based on the position of the object in each image estimated by tracking by the object detection and tracking unit 123 and the third area. Therefore, even if an object cannot be detected in some of the multiple images in which an event is detected (an event area is estimated), the event-related object in the third image can be identified using the results of estimating the position of the object in the multiple images by object tracking from previous and subsequent images (regardless of whether an event area is estimated). This further improves the accuracy of identifying the event-related object.

[0041] Furthermore, as described above, the object detection and tracking unit 123 further estimates the position of an object included in a fourth image in which no event was detected from the multiple image sequences and which was not selected by the event detection unit 121. In this case, the identification unit 124 may identify an event-related object included in the fourth image based on the position and area of ​​the object estimated from the selected image and the position of the object estimated from the fourth image. In other words, the identification unit 124 can identify an event-related object even in an image in which no event was detected. Therefore, even in images in which no event was detected within the entire video data 111, the position and the like of the event-related object can be displayed on the screen in a manner that makes it identifiable within the image, thereby facilitating support for event analysis.

[0042] The identification unit 124 may also identify each of a plurality of event-related objects related to the same detected event. Identifying a plurality of event-related objects per event allows the plurality of event-related objects to be displayed distinctly within a single image, thereby facilitating support for event analysis.

[0043] The identification unit 124 may also identify a plurality of event-related objects associated with each of the detected events. Since a plurality of events may occur in the video data 111, the event-related objects for each event can be displayed on the screen in a manner that allows them to be identified within the image, thereby facilitating support for event analysis.

[0044] Each of the event detection unit 121, the region estimation unit 122, and the object detection and tracking unit 123 may transmit a processing request including the image to be processed to an external server and receive a processing result from the external server. For example, the event detection unit 121 may transmit an event detection request for each image or including the entire video data 111 to a first server running a first trained model. In response to the received event detection request, the first server inputs each image into the first trained model and returns a score output from the first trained model to the event detection unit 121 (information processing device 100). The event detection unit 121 selects an image in which an event was detected based on the received score. Alternatively, the first server may return an indication of event detection to the event detection unit 121 if the score output from the first trained model is equal to or greater than a threshold. Alternatively, when the first trained model outputs whether an event has been detected, the first server may return information associating whether an event has been detected with each image to the event detection unit 121. In these cases, the event detection unit 121 selects the image to be processed when it receives the notification that an event has been detected as an image in which an event has been detected. Therefore, it can be said that the event detection unit 121 selects an image in which an event has been detected by obtaining whether an event has been detected for each image in the video data 111.

[0045] Furthermore, the area estimation unit 122 may transmit an event area estimation request, including the selected image by the event detection unit 121, to a second server on which a second trained model is running. Then, in response to the received event area estimation request, the second server inputs the image to the second trained model and returns the event area output from the second trained model to the area estimation unit 122 (information processing device 100). In response to this, the area estimation unit 122 estimates the received event area as an area that contributed to the detection of the event within the selected image. Therefore, it can be said that the area estimation unit 122 estimates the event area within the selected image as an area that contributed to the detection of the event by acquiring the event area in the selected image.

[0046] Furthermore, the object detection and tracking unit 123 transmits an object detection and tracking request for each image or including the entire video data 111 to a third server that performs object detection processing and object tracking processing. Then, in response to the received object detection and tracking request, the third server detects the position, etc. of the object detected for each image and returns the detection results to the object detection and tracking unit 123 (information processing device 100). In response to this, the object detection and tracking unit 123 acquires the received detection results as the position, etc. of the object in each image. Therefore, it can be said that the object detection and tracking unit 123 detects the position, etc. of the object by acquiring the detection results, etc. of the object for each image in the video data 111.

[0047] The acquisition unit 125 inputs the event-related object information 1121, etc. into the fourth trained model, acquires the evaluation information 113 from the fourth trained model, and stores the evaluation information 113 in the storage unit 110. Here, the fourth trained model is an AI model that generates and outputs evaluation information, such as explanatory text about an event, from the input event-related object information 1121, etc. Note that the acquisition unit 125 may acquire the evaluation information using a method other than the fourth trained model, i.e., a module in which an evaluation information generation process is implemented. Alternatively, the acquisition unit 125 may transmit an evaluation request including the event-related object information 1121, etc. to a fourth server on which the fourth trained model or the above module is executed. In this case, the fourth server returns the evaluation information output from the event-related object information 1121, etc. included in the evaluation request using the fourth trained model or the above module to the acquisition unit 125 (information processing device 100). The various trained models described above may be stored in the storage unit 110.

[0048] The display control unit 126 outputs to the display terminal 200 a region corresponding to the event-related object in an identifiable manner within an image in which at least the object position and the event area have been estimated, so as to be displayed on the screen 201. The display control unit 126 also outputs to the display terminal 200 a region corresponding to the event-related object in an identifiable manner within an image in which at least the object position and the event area have been estimated, so as to be displayed ... on the screen 201. The display control unit 126 also outputs to the display terminal 200 a region corresponding to the event-related object in an identifiable manner within an event area bar indicative of an event section in which the event was detected during a predetermined period of time. In addition, the display control unit 126 outputs to the display terminal 200 a variety of display information to be displayed, as will be described later. The display control unit 126 may also display a text relating to the detected event and display an edit field on the screen 201 for accepting edits to the text.

[0049] The receiving unit 127 receives various types of information input by the user via the display terminal 200. For example, the receiving unit 127 receives operations on various types of display information displayed on the screen 201 of the display terminal 200, which will be described later.

[0050] 9 is a flowchart showing the flow of the process for identifying an event-related object. First, the event detection unit 121 performs an event detection process for each image in the video data 111 (S11). Then, the event detection unit 121 selects an image in which an event has been detected (S12). Then, the area estimation unit 122 estimates an area (event area) that contributed to the detection of the event within the selected image (S13).

[0051] Furthermore, independently of steps S11 to S13, the object detection and tracking unit 123 performs object detection processing for each image in the video data 111 and object tracking processing for the image sequence (S14). After steps S13 and S14, the identification unit 124 identifies an event-related object by comparing the event area estimated from the selected image with the object area estimated from the same selected image (S15). The identification unit 124 then stores information about the identified event-related object (event-related object information) in the storage unit 110 (S16). After step S16, the acquisition unit 125 acquires the evaluation information 113 generated based on the event-related object information 1121, etc., as described above, and stores it in the storage unit 110.

[0052] (Event Visualization Screen) The user then analyzes the event included in the video data 111 by playing back the video data 111 and referring to the evaluation information. For example, the user may consider a car collision accident as an event and use the video data 111 recorded by an in-vehicle camera of one of the cars involved in the collision to analyze the cause of the accident and determine the degree of fault.

[0053] Specifically, in response to a playback start operation by the user, the display terminal 200 transmits a playback start request for video data to the information processing device 100 via the communication network N. The information processing device 100 generates an event visualization screen in response to the playback start request received from the display terminal 200. Specifically, the display control unit 126 appropriately assigns an area for an event-related object to the image to be played back in the video data 111, generates an event visualization screen including an event area bar, etc., and returns the generated event visualization screen to the display terminal 200 via the communication network N. In response to this, the display terminal 200 displays the received event visualization screen on the screen 201.

[0054] FIG. 10 is a diagram showing an example of an event visualization screen 3. The event visualization screen 3 includes a video display area 31, a visualization support area 32, and an evaluation information display / edit area 33. The video display area 31 is an area for displaying a frame image 51 to be played back. The frame image 51 shows an example in which marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The mark 61 is an example of display information in which a traffic light, which is an example of an event-related object, is enclosed in a rectangle. The color of the border of the rectangle of the mark 61 may match the color of the traffic light. For example, if the traffic light is recognized as a red light, the mark 61 may be a red rectangle. The mark 62 is an example of display information in which a car, which is an example of an event-related object, is enclosed in a rectangle. In other words, the car enclosed by the mark 62 is an oncoming car of the car equipped with an on-board camera that captured the frame image 51. In other words, this example shows that a traffic light and an oncoming car are displayed on the screen 201 in a manner that makes them identifiable as event-related objects in a car collision accident.

[0055] The visualization support area 32 includes a playback loading bar 321, an event area bar 322, a return to the beginning button 41, a return to the previous chapter button 42, a return to one frame button 43, a play button 44, a forward one frame button 45, a forward to the next chapter button 46, a seek bar display switch button 47, a video save button 48, and a marking display switch button 49.

[0056] The playback loading bar 321 is display information that indicates the playback position of the frame image 51 currently being played back during a predetermined period that is the shooting period of the video data 111. In other words, the playback loading bar 321 is display information that indicates, in length, the number of images from the first image to the currently played image relative to the total number of images included in the video data 111. Alternatively, the playback loading bar 321 is display information that indicates, in length from the playback start time, the ratio of the total time that frame images have been played back to the total time of the predetermined period. Note that the playback loading bar 321 shows an example in which the playback time progresses from left to right.

[0057] The event area bar 322 is display information that indicates the time period of a sequence of images in which an event was detected in the video data 111, expressed as a length relative to the total time of a predetermined period and a position within the predetermined period. The playback loading bar 321 and the event area bar 322 both represent time in one dimension in the same direction and are displayed side by side, one above the other. FIG. 10 shows an example in which the playback loading bar 321 has not yet reached the section of the event area bar 322. That is, the frame image 51 is an image in which no event was detected. This makes it easier for the user to visually recognize that the frame image 51 represents a time period before an accident occurred. The direction of progression of the playback time on the playback loading bar 321 and the direction in which the playback loading bar 321 and the event area bar 322 are arranged are not limited to these. For example, the direction of progression of the playback time on the playback loading bar 321 may be from right to left. Alternatively, if the direction of progression of the playback time on the playback loading bar 321 is up and down, the playback loading bar 321 and the event area bar 322 may be displayed side by side.

[0058] The return-to-top button 41 is a button for returning the frame image displayed in the video display area 31 to the first image of the video data 111, i.e., for changing to the frame image at the playback start time. The return-to-previous-chapter button 42 is a button for returning the frame image displayed in the video display area 31 to the frame image assigned to the chapter immediately before the chapter assigned to the frame image displayed in the video display area 31. Note that the "chapter" may correspond to a detected event. The return-to-one-frame button 43 is a button for returning the frame image displayed in the video display area 31 to the frame image at the previous time. The play button 44 is a button for starting playback of the frame images of the video data 111 in the video display area 31. The forward-to-one-frame button 45 is a button for advancing the frame image displayed in the video display area 31 to the frame image at the next later time. The next chapter advance button 46 is a button for advancing the frame image displayed in the video display area 31 to a frame image assigned with the chapter immediately following the chapter assigned to the frame image displayed in the video display area 31.

[0059] The seek bar display switch button 47 is a button for switching between displaying and hiding the seek bar, which will be described later. Note that Fig. 10 shows a state in which the seek bar is not displayed. The video save button 48 is a button for downloading the frame image currently displayed in the video display area 31 to the display terminal 200 as an image file. The marking display switch button 49 is a button for switching between displaying and hiding the "mark," which is display information added to the frame image displayed in the video display area 31. Note that Fig. 10 shows a state in which the marks 61 and 62 are displayed, as described above.

[0060] Here, it is assumed that the user has pressed the seek bar display switch button 47. In this case, the display terminal 200 transmits a notification that the seek bar display switch button 47 has been pressed to the information processing device 100. In response to receiving the notification that the seek bar display switch button 47 has been pressed, the display control unit 126 of the information processing device 100 transmits display information for the seek bar, including the current playback time, playback frame number, etc., to the display terminal 200. The display terminal 200 displays a seek bar on the screen 201 based on the received display information.

[0061] FIG. 11 is a diagram showing an example of a seek bar on the event visualization screen. FIG. 11 shows an example in which a seek bar 3211, a playback position 3212, and a frame movement instruction field 3213 are displayed near the bottom of the above-described return to beginning button 41, etc. The seek bar 3211 is a user interface that assists the user in changing the playback position of the frame image displayed in the video display area 31. The right end of the seek bar 3211 indicates the playback position 3212. In this case, the playback position 3212 of the seek bar 3211 and the right end of the playback loading bar 321 coincide in the vertical direction, making it easy for the user to grasp the playback position. However, the display positions of the seek bar 3211, etc. are not limited to this. Alternatively, the playback loading bar 321 may include the function of a seek bar.

[0062] Playback position 3212 is a reception unit that receives left and right movement by user operation via display terminal 200. When display terminal 200 receives that playback position 3212 has been held and moved by user operation (e.g., mouse), it transmits this information (movement direction and amount) to information processing device 100. Display control unit 126 of information processing device 100 identifies the image of the corresponding playback time according to the received movement direction and amount of movement of playback position 3212, and returns the image to display terminal 200 to display on screen 201. Display terminal 200 displays the received image in video display area 31.

[0063] The frame movement instruction field 3213 is a reception unit that receives, via the display terminal 200, a user's input of the number of frames by which the display should be moved forward or backward, a change in the number using the up and down buttons, and a movement forward or backward by a certain number of frames. After receiving an input or change in the number of frames to be moved by user operation (e.g., a mouse or keyboard), the display terminal 200 transmits a notification (the number of frames and the direction of movement) to the information processing device 100 upon receiving a movement forward or backward by a certain number of frames. The display control unit 126 of the information processing device 100 identifies an image with the corresponding number of frames based on the received number of frames and direction of movement, and returns the image to the display terminal 200 to display on the screen 201. The display terminal 200 displays the received image in the video display area 31. This makes it easier for the user to adjust the playback position.

[0064] The evaluation information display / edit area 33 in Fig. 10 is an area for displaying and editing the above-mentioned evaluation information 113. Fig. 12 is a diagram showing an example of the initial evaluation information 113 displayed in the evaluation information display / edit area 33.

[0065] Fig. 13 is a diagram showing an example of another frame image 52 displayed on the event visualization screen. Similar to the frame image 51 of Fig. 10, the frame image 52 shows an example in which marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the vehicle has progressed further than in Fig. 10 and reached the section of the event area bar 322. In other words, the frame image 52 is an image in which an event has been detected. The mark 62 indicates that the oncoming vehicle is larger than in Fig. 10 and is therefore approaching the vehicle.

[0066] FIG. 14 is a diagram showing an example of another frame image 53 displayed on the event visualization screen. Similar to the frame image 52 of FIG. 13 , the frame image 53 shows an example in which the marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the image has progressed from FIG. 13 and is within the section of the event area bar 322. In other words, the frame image 53 is an image in which an event has been detected. Compared to FIG. 13 , the mark 62 shows the area around the windshield of the oncoming vehicle, but not the area around the hood, indicating that the airbag has been deployed. In other words, the frame image 53 shows an image of a collision between the host vehicle and an oncoming vehicle.

[0067] FIG. 15 is a diagram showing an example of another frame image 54 displayed on the event visualization screen. Similar to the frame image 53 in FIG. 14 , the frame image 54 shows an example in which the marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the image has progressed from FIG. 14 and is within the section of the event area bar 322. In other words, the frame image 54 is an image in which an event has been detected. Similar to FIG. 14 , the mark 62 shows the area around the windshield of the oncoming vehicle, but not the area around the hood, and indicates that the airbag has opened and the wipers are operating. In other words, the frame image 54 continues to show an image of a collision between the vehicle and the oncoming vehicle.

[0068] FIG. 16 is a diagram showing an example of another frame image 55 displayed on the event visualization screen. Similar to the frame image 54 of FIG. 15 , the frame image 55 shows an example in which the marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the image has progressed further than in FIG. 15 and is within the section of the event area bar 322. In other words, the frame image 55 is an image in which an event has been detected. The mark 62 indicates that, compared to FIG. 15 , the oncoming vehicle has been pushed back slightly in reaction to the collision. In other words, the frame image 55 shows an image after the host vehicle and the oncoming vehicle have collided.

[0069] FIG. 17 is a diagram showing an example of another frame image 56 displayed on the event visualization screen. Similar to the frame image 55 of FIG. 16 , the frame image 56 shows an example in which the marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the image has progressed from FIG. 16 and is within the section of the event area bar 322. In other words, the frame image 56 is an image in which an event has been detected. Compared to FIG. 16 , the mark 62 indicates that the oncoming vehicle has been pushed back slightly in reaction to the collision, and that the area around the hood is visible, indicating that the front left portion of the oncoming vehicle is damaged. In other words, the frame image 56 shows an image after the host vehicle and the oncoming vehicle collide.

[0070] FIG. 18 is a diagram showing an example of another frame image 57 displayed on the event visualization screen. Similar to the frame image 56 of FIG. 17 , the frame image 57 shows an example in which the marks 61 and 62 are displayed on the screen 201 in a manner that makes them identifiable within the image. The playback loading bar 321 indicates that the image has progressed from FIG. 17 and is after the section of the event area bar 322. In other words, the frame image 57 is an image in which no event was detected. Compared to FIG. 17 , the mark 62 indicates that the oncoming vehicle has been pushed further back in reaction to the collision, and that the area around the hood is visible, with the left front portion of the oncoming vehicle being damaged. In other words, the frame image 57 shows an image after the host vehicle and the oncoming vehicle collide.

[0071] FIG. 19 is a diagram illustrating the relationship between event detection, object estimation, and event-related object identification in each image. Here, images #1 to #7 correspond to the frame images 51 to 57, respectively. First, the event detection unit 121 selects images #2 to #6 as images in which an event was detected. Meanwhile, the event detection unit 121 assumes that it was unable to detect an event from images #1 and #7. Then, the region estimation unit 122 estimates two event areas corresponding to an oncoming vehicle and a traffic light from images #2 and #6. Furthermore, the region estimation unit 122 estimates two event areas corresponding to the vicinity of the vehicle's windshield and the traffic light from images #3 and #4. Furthermore, the region estimation unit 122 estimates two event areas corresponding to the vicinity of the vehicle's windshield and the hood and the traffic light from image #5.

[0072] Furthermore, the object detection and tracking unit 123 detects the position of an object corresponding to an oncoming vehicle from images #1, #2, and #5 to #7. On the other hand, it is assumed that the object detection and tracking unit 123 was unable to detect an object corresponding to an oncoming vehicle from images #3 and #4. As described above, this is because in frame images 53 and 54, an oncoming vehicle has collided with the host vehicle. The object detection and tracking unit 123 detects the position of an object corresponding to a traffic light (in the direction of travel of the host vehicle) from images #1 to #7.

[0073] Based on this, the identification unit 124 compares the event area #2 estimated from image #2 with the object detection result #2 to identify the oncoming vehicle and the traffic light (in the direction of travel of the host vehicle) (their respective positions) as the event-related object (position). The same applies to images #5 and #6. Furthermore, for image #3, the event area #3 of the oncoming vehicle was estimated, but no object of the oncoming vehicle was detected. Therefore, the identification unit 124 compares the event area #3 in image #3 with the position of the event-related object in image #2 to identify the oncoming vehicle (position) in image #3 as the event-related object (position). For image #4, the identification unit 124 may use the position of the event-related object identified from image #5 to identify the oncoming vehicle (position) in image #4 as the event-related object (position). Alternatively, the object detection and tracking unit 123 may track the position of the oncoming vehicle from the object detection results for images #2 and #6. In this case, the identification unit 124 may identify (the position of) the oncoming vehicle in images #3 and #4 as (the position of) the event-related object based on the tracking results of the oncoming vehicle in images #2 and #6.

[0074] Furthermore, the identification unit 124 may identify (the position of) an oncoming vehicle in image #1 as (the position of) an event-related object by comparing the object detection result #1 for image #1 with the event-related object information #2 for image #2. Similarly, the identification unit 124 may identify (the position of) an oncoming vehicle in image #7 as (the position of) an event-related object by comparing the object detection result #7 for image #7 with the event-related object information #6 for image #6.

[0075] In this way, the technology disclosed herein can further improve the accuracy of identifying an object related to an event in an image. Furthermore, by displaying the area of ​​the identified event-related object on the screen in a manner that makes it identifiable within the image, analysis of the event can be effectively supported when playing back and displaying video data including the events before and after the event.

[0076] Fifth Embodiment Below, examples of variations of the event visualization screen will be described. Note that the configuration of the information processing system according to this fifth embodiment is the same as that shown in Fig. 7 above, and therefore illustrations and descriptions thereof will be omitted. Note that, below, descriptions of functions and processes equivalent to those of the above-mentioned embodiments will be omitted as appropriate.

[0077] FIG. 20 is a block diagram showing the configuration of an information processing device 100a. The information processing device 100a is a modified example of the information processing device 100 and is an example of the information processing devices 1, 1a, and 1b. The information processing device 100a differs from the information processing device 100 in that the display control unit 126 and the reception unit 127 are replaced with a display control unit 126a and a reception unit 127a. The display control unit 126a is an example of the display control unit 15a or 15b. The event information 112 stored in the storage unit 110 is information including event-related object information 1121 to 112m. The event information 112 also includes information other than the event-related object information 1121, etc., and may be referred to as analysis viewpoint information. The event information generation unit 120a is a functional block that combines the event detection unit 121, the area estimation unit 122, the object detection and tracking unit 123, and the identification unit 124.

[0078] The display control unit 126a may have the following functions in addition to the functions of the display control unit 126. The display control unit 126a may display multiple event area bars on the screen 201. Furthermore, the display control unit 126a may display multiple event area bars on the screen 201 so that the positional relationship of each event section can be identified based on the length of a predetermined period. This makes it easier to understand the positional relationship between different event sections. The display control unit 126a may also display multiple event area bars on the screen 201 that indicate event sections corresponding to each of the multiple event-related objects identified by the identification unit 124. This makes it easier to understand the differences between event sections between different event-related objects. The display control unit 126a may also display multiple event area bars on the screen 201 that correspond to each of the multiple events detected by the event detection unit 121. Furthermore, the display control unit 126a may treat each of the multiple states of an event-related object that have changed over a predetermined period as multiple events detected by the event detection unit 121, and display multiple event area bars on the screen 201, with the duration of each state being the event interval.

[0079] In addition to the functions of the receiving unit 127, the receiving unit 127a receives designation of at least some of a plurality of analysis perspectives (event information 112) of an event, including the event-related object information 1121. Then, in response to receiving the designation, the display control unit 126a may control switching between displaying and hiding the area and the event area bar on the screen 201 based on the analysis perspective received by the receiving unit 127a.

[0080] (Event visualization screen) Fig. 21 is a diagram showing an example of an event visualization screen 3a. The event visualization screen 3a includes a video display area 31 and a visualization support area 32a. Note that the event visualization screen 3a does not include the evaluation information display and editing area 33. The visualization support area 32a includes a playback loading bar 321, event area bars 322, 323, 324, and 325, and analysis viewpoint selection fields 3261 to 3264.

[0081] The event area bars 322 and 323 are examples of event area bars corresponding to different event sections. The event area bars 322 and 323 may correspond to different events. For example, the event area bar 322 corresponds to an event of a collision between the host vehicle and an oncoming vehicle (vehicle Y) marked with mark 62. In this case, the event area bar 323 may correspond to an event of a secondary accident between the oncoming vehicle marked with mark 62 and another vehicle. In this case, the event area bar 322 may correspond to a first event section of a first event, and a first chapter may be assigned to the first event section. Similarly, the event area bar 323 may correspond to a second event section of a second event, and a second chapter may be assigned to the second event section. In these cases, the user may switch the image displayed in the video display area 31 to an image belonging to the first or second event section by pressing the previous chapter back button 42 or the next chapter forward button 46 in FIG. 10 above. This allows the user to easily understand the length of the sections, the time difference between events, and the like for each time period of different events within the shooting period of the video data 111. This can more effectively support the analysis of events.

[0082] Furthermore, event area bars 324 and 325 are examples of event area bars corresponding to different event-related objects. For example, event area bar 324 may correspond to a traffic light (signal X) represented by mark 61, and event area bar 325 may correspond to an oncoming vehicle (vehicle Y) represented by mark 62. This allows easy overview understanding of the differences and overlapping of event sections among multiple event-related objects. This can more effectively support event analysis.

[0083] Furthermore, in this case, the event area bar 324 includes sub-event area bars 3241 and 3242. For example, the sub-event area bar 3241 may correspond to an event section where signal X is red, and the sub-event area bar 3242 may correspond to an event section where signal X is green. This allows for an overview of the relationship between the change in signal and the timing of a collision accident, the timing of an oncoming vehicle, and the like. This provides more effective support for event analysis.

[0084] The analysis viewpoint selection fields 3261 to 3264 are fields that accept the selection of an analysis viewpoint included in the event information 112. Each of the analysis viewpoint selection fields 3261 to 3264 accepts a specification for displaying corresponding display information in the frame image 51 displayed in the video display area 31 when selected. As specific examples of analysis viewpoints, the analysis viewpoint selection field 3261 indicates signal X, the analysis viewpoint selection field 3262 indicates (oncoming) vehicle Y, the analysis viewpoint selection field 3263 indicates the traveling direction of the host vehicle, and the analysis viewpoint selection field 3264 indicates the traveling direction of vehicle Y. In FIG. 21 , the analysis viewpoint selection fields 3261 and 3262 are in the selected state, so marks 61 and 62 are displayed in the video display area 31, and event area bars 324 and 325 are displayed in the visualization support area 32 a.

[0085] Here, it is assumed that the user intends to perform analysis by focusing on signal X. In this case, the display terminal 200 accepts a designation in the analysis viewpoint selection field 3262 through a user operation and transmits the designation to the information processing device 100a via the communication network N. At this point, the analysis viewpoint selection field 3262 may be updated to a non-selected state. The reception unit 127a of the information processing device 100a accepts the designation of the analysis viewpoint (vehicle Y) corresponding to the received analysis viewpoint selection field 3262. The display control unit 126a then refers to the event information 112 and identifies the analysis viewpoint whose designation has been accepted. In this case, the display control unit 126a identifies display information related to "vehicle Y," which is part of the event-related object information in the event information 112. The display control unit 126a then returns an instruction to the display terminal 200 to hide the mark 62 and the event area bar 325 corresponding to "vehicle Y." In response to this, the display terminal 200 updates the display content of the screen 201 by hiding the mark 62 from the video display area 31 and hiding the event area bar 325 from the visualization support area 32a.

[0086] 22 is a diagram showing an example of the event visualization screen 3a after the designation of the analytical viewpoint has been changed. Compared to FIG. 21, the event visualization screen 3a shows that the analytical viewpoint selection field 3262b is in a non-selected state, the mark 62 is not displayed in the video display area 31, and the event area bar 325 is not displayed. Note that when another analytical viewpoint is selected, the display information corresponding to the analytical viewpoint is similarly switched between being displayed and not displayed.

[0087] In this way, by displaying multiple event area bars on the screen, the technology disclosed herein allows the user to easily grasp the interval between events, the length of the event, the relationship between event sections in different event-related objects, the timing of transitions between different event sections in the same event-related object, etc. This allows for more effective support for event analysis. Furthermore, by switching between displaying and hiding display information corresponding to multiple different analysis perspectives in real time, the technology can flexibly respond to the user's analysis requests and more effectively support event analysis.

[0088] Sixth Embodiment Hereinafter, a technique will be described in which evaluation information regarding a currently occurring event is generated from information regarding the occurring event detected from video data and information regarding past cases related to the event, and the evaluation information is presented to a user.

[0089] For example, if the event that occurred is a car traffic accident, a user would analyze the accident situation using various data sources, including video data before and after the accident, and create a fault ratio determination result and a report for the accident. In this case, there is a problem that various data sources are required to accurately understand the accident situation. To solve this problem, the technology disclosed herein describes the following method that appropriately utilizes the above technology. Note that, in the following description, an example is described in which the event to be analyzed is a car traffic accident, but events to which the technology disclosed herein can be applied are not limited to this. Note that the configuration of the information processing system according to the sixth embodiment is the same as that shown in FIG. 7 above, and therefore illustrations and descriptions thereof will be omitted. Note that, hereinafter, descriptions of functions and processes equivalent to those of the above-mentioned embodiments will be omitted as appropriate.

[0090] FIG. 23 is a block diagram showing the configuration of an information processing device 100b. The information processing device 100b is a modified example of the information processing device 100 or 100a and is an example of the information processing devices 1, 1a, and 1b. The information processing device 100b differs from the information processing device 100a in that a domain knowledge DB (DataBase) 114 and a trained model 115 are added to the storage unit 110. The information processing device 100b also differs from the information processing device 100a in that the event information generation unit 120a, acquisition unit 125, display control unit 126a, and reception unit 127a are replaced with an event information generation unit 120b, acquisition unit 125b, display control unit 126b, and reception unit 127b. The display control unit 126b is an example of the display control unit 15a or 15b. The information processing device 100b differs from the information processing device 100a in that an extraction unit 128 and a generation unit 129 are added.

[0091] The domain knowledge DB 114 is a database that manages information such as past cases, knowledge, and know-how in a specific field. The domain knowledge DB 114 is an example of a collection of case information that indicates cases related to events. If the specific field is traffic accidents, the domain knowledge DB 114 is a database of precedent information on the fault ratio in traffic accidents. The domain knowledge DB 114 can be said to correspond to a storage area managed by database management software (not shown).

[0092] The trained model 115 is a model that receives input information based on at least a portion of the event information 112 and case information extracted from the domain knowledge DB 114, and outputs evaluation information 113. The input information is text data of instruction sentences for the trained model 115, so-called input prompts. The trained model 115 is, for example, a language model that has been machine-trained to generate and output sentences of evaluation information based on the instruction sentences. The trained model 115 may also be called a machine learning model. For example, the trained model 115 may be an LLM (Large Language Model).

[0093] The event information generation unit 120b includes one or more "analysis units." The analysis units may also be referred to as analysis engines. FIG. 23 illustrates an example in which the event information generation unit 120b includes one or more analysis units 1201 in addition to a functional block that collectively includes the event detection unit 121, the region estimation unit 122, the object detection and tracking unit 123, and the identification unit 124. Therefore, the functional block that collectively includes the event detection unit 121, the region estimation unit 122, the object detection and tracking unit 123, and the identification unit 124 may also be referred to as one analysis unit. Therefore, the event information generation unit 120b can also be considered a second generation means that generates event information including information about an event-related object related to an event identified based on a region that contributed to the detection in an image in which an event was detected among a sequence of multiple images captured during a predetermined period.

[0094] The analysis unit 1201 analyzes the video data 111 for a specific field using image recognition or the like, generates a portion of the event information 112 as the analysis result, and stores the generated portion in the storage unit 110. For example, the identification unit 124, which is part of the analysis unit, includes event-related object information 1121 as the analysis result in the event information 112 and stores the event information 112 in the storage unit 110. Examples of the analysis engine include, but are not limited to, the following: When the specific field is accident investigation, the analysis engine may be an accident scene extraction engine, a traffic light color detection engine, a vehicle motion recognition engine, etc. When the specific field is crime investigation, the analysis engine may be an abandoned object detection engine, a person attribute detection engine, etc. When the specific field is factory work, the analysis engine may be a posture detection engine, a visual inspection engine, etc.

[0095] The extraction unit 128 extracts case information from the domain knowledge DB 114, which is a collection of multiple case information, by narrowing down the data based on the event information 112. When a user selects case information to be analyzed from the extraction results (multiple case information) by the extraction unit 128, the extraction unit 128 extracts one or more case information candidates from the domain knowledge DB 114 based on the event information 112.

[0096] The generation unit 129 uses the event information and the case information extracted by the extraction unit 128 to generate, as input information, an instruction sentence for the trained model 115. For example, the generation unit 129 may apply the event information and the case information to a template of an instruction sentence to generate the instruction sentence.

[0097] The acquisition unit 125b inputs input information to the trained model 115 to acquire the evaluation information 113 output from the trained model 115. The trained model 115 may run on an external server. In this case, the acquisition unit 125b may transmit an evaluation request including the input information to a fifth server on which the trained model 115 runs. In this case, the fifth server inputs the input information included in the evaluation request to the trained model 115 and returns the evaluation information output from the trained model 115 to the acquisition unit 125b (information processing device 100b). In response to this, the acquisition unit 125b acquires the received evaluation information. The generation and acquisition of the evaluation information do not need to use the trained model 115. For example, the acquisition unit 125b may acquire the evaluation information from a predetermined evaluation information generation module without using the trained model 115. The evaluation information generation module may be a computer program in which a predetermined algorithm for generating evaluation information from event information and case information extracted by the extraction unit 128 is implemented. Therefore, the generation unit 129 may have the function of an evaluation information generation module. In this case, instead of the acquisition unit 125b, the generation unit 129 may generate evaluation information from the event information and the case information extracted by the extraction unit 128.

[0098] The display control unit 126b displays, on the screen 201, information to be analyzed from the event information 112, (candidates for) case information extracted from the domain knowledge DB 114, evaluation information acquired by the acquisition unit 125b, etc. The display control unit 126b also displays, on the screen 201, input information generated by the generation unit 129. The display control unit 126b also displays, on the screen 201, a reception field for receiving input for correcting the event information. When the evaluation information 113 is updated by the trained model 115 in response to an update of at least one of the event information and the case information, the display control unit 126b also displays the updated evaluation information on the screen 201.

[0099] The receiving unit 127b receives input for modifying the event information via the reception field. The receiving unit 127b also receives a user's selection from among candidate case information. The receiving unit 127b receives from the user instructions for generating and displaying evaluation information, instructions for making the evaluation information editable, and instructions for displaying input information.

[0100] 24 is a flowchart showing the flow of the initial evaluation information generation and display process. First, the event information generation unit 120b generates event information 112 from the video data 111 (S31). At this time, as described above, the event information generation unit 120b uses multiple analysis engines (analysis units) to store the analysis results of each image in the video data 111 as event information 112 in the storage unit 110 (S32). At this time, the event information generation unit 120b generates and stores the event information 112 including event-related object information 1121, etc.

[0101] Next, the extraction unit 128 extracts case information based on the event information 112 (S33). Specifically, the extraction unit 128 searches the domain knowledge DB 114 using keywords included in the event information 112 and extracts case information candidates as search results. At this point, since this is the initial evaluation information generation process, the extraction unit 128 extracts some case information from multiple case information candidates based on predetermined criteria. This is because the generated event information 112 may lack some information necessary for analysis, making it difficult to sufficiently narrow down the case information. For example, the generated event information 112 may include the vehicle's traveling direction and the color of its traffic light, but may not include the traveling direction of oncoming vehicles or the color of the traffic light in the oncoming vehicle's traveling direction. Therefore, for example, the extraction unit 128 may further extract, as case information, the case with the highest number of accidents from the search results (case information candidates) using keywords included in the event information 112.

[0102] The generation unit 129 then generates an input prompt from the event information 112 generated in step S31 and the case information extracted in step S32 (S34). The acquisition unit 125b then inputs the input prompt to the trained model 115 (S35). The acquisition unit 125b then acquires output of evaluation information from the trained model 115 (S36). The reception unit 127b then receives a display operation for the evaluation information from the display terminal 200 in response to a user operation (S37). In response to this, the display control unit 126b displays an event visualization screen including event information, case information, evaluation information, and the like on the screen 201 of the display terminal 200 (S38).

[0103] 25 is a diagram showing an example of the layout of the event visualization screen 3b. The event visualization screen 3b includes a video display area 31, a visualization support area 32, an evaluation information display / edit area 33, a user operation reception area 34, an evaluation information generation result display area 35, an event information display / edit area 36, ​​and a case information candidate display / selection area 37. Note that the layout of the event visualization screen 3b is not limited to this.

[0104] The video display area 31, visualization support area 32, and evaluation information display / editing area 33 are the same as those in FIG. 10 above. The user operation reception area 34 is an area for receiving user operations related to the generation, display instructions, and editing of evaluation information. The evaluation information generation result display area 35 is an area for displaying the results of the generation of evaluation information. The event information display / editing area 36 is an area for displaying and accepting editing of event information. The case information candidate display / selection area 37 is an area for displaying and accepting selection of case information candidates.

[0105] FIG. 26 is a diagram showing an example of initial evaluation information displayed in the user operation reception area 34 and the evaluation information generation result display area 35. The user operation reception area 34 includes a generate evaluation information button 341, a copy evaluation information button 342, and a confirm previous prompt button 343. The generate evaluation information button 341 is a button for accepting the generation and display of evaluation information from accident information (analyzed event information) and negligence conformance judgment (selected case information). The copy evaluation information button 342 is a button for accepting the copying of evaluation information displayed in the evaluation information generation result display area 35 to the evaluation information display / edit area 33, making the evaluation information editable. The confirm previous prompt button 343 is a button for accepting the display of the immediately preceding (latest) input prompt used when generating the evaluation information displayed in the evaluation information generation result display area 35. The evaluation information generation result display area 35 is an area for displaying evaluation information generated in response to pressing the generate evaluation information button 341.

[0106] When the user presses any of the evaluation information generation button 341 to the last prompt confirmation button 343, the display terminal 200 transmits a message to that effect to the information processing device 100b. The information processing device 100b performs processing according to the pressed button and returns the processing result to the display terminal 200. The display terminal 200 displays the received processing result on the screen 201.

[0107] FIG. 27 is a diagram illustrating an example of automatically generated event information displayed in the event information display / editing area 36. Here, the event information is referred to as "accident information." The event information display / editing area 36 includes multiple fields for display, selection, and input. The event information display / editing area 36 displays some of the information (e.g., keywords) included in the event information in each corresponding field. For example, the date and time of occurrence are assumed to be information specified by the user or extracted from the event information by the display control unit 126b or the like. The event information display / editing area 36 also includes event information input fields 361 and 362. The event information input field 361 illustrates an example in which an analysis result of the accident situation extracted from the event information by the display control unit 126b or the like is displayed. The event information input field 361 indicates, for example, that the location type of the accident is analyzed as "intersection with traffic lights," the vehicle type is "four-wheeled vehicle," the vehicle traveling direction is "straight," and the vehicle's signal is "green." In other words, this indicates that the event information generation unit 120b has analyzed this information from the video data 111 and saved it by including it in the event information 112. Meanwhile, the event information input acceptance field 362 indicates that the other vehicle type, other vehicle position, other vehicle traveling direction, and other vehicle signal have not yet been input at this point. Specifically, this indicates that the event information generation unit 120b was unable to analyze the other vehicle type, other vehicle position, other vehicle traveling direction, and other vehicle signal from the video data 111. The event information display / edit area 36 may also include a button for "resetting the judgment conditions."

[0108] FIG. 28 is a diagram showing an example of case information candidates initially extracted and displayed in the case information candidate display selection area 37. The case information candidate display selection area 37 is an area that displays, for each case information candidate, items such as number, image, accident type, subject vehicle type, subject vehicle direction of travel, subject vehicle signal, other vehicle type, other vehicle direction of travel, other vehicle signal, remarks, subject vehicle fault ratio, and other vehicle fault ratio. The "image" may be a thumbnail image of an accident situation diagram. However, the case information candidate display selection area 37 does not necessarily require the display of an "image." Furthermore, when a row of any case information candidate in the case information candidate display selection area 37 is pressed (e.g., double-clicked), the display control unit 126b may display a detailed display screen of the corresponding case information candidate on the screen 201. In this case, the detailed display screen may display, for example, an accident situation diagram as detailed information about the case information. Furthermore, when a selection operation for any case information candidate is received in the case information candidate display selection area 37, the check box of the corresponding case information candidate is selected. In this example, the number 98 is in the selected state.

[0109] FIG. 29 is a diagram showing an example of an input prompt when generating initial evaluation information. The previous prompt display field 38 displays an instruction sentence 381. The instruction sentence 381 is an example of an input prompt generated in step S34 of FIG. 24 and input to the trained model 115 in step S35. As described above, when generating initial evaluation information, the event information input acceptance field 362 of FIG. 27 is empty. Therefore, as described above, for example, the extraction unit 128 complements information that was not included in the event information 112 (that was not analyzed from the video data 111) according to a predetermined standard. Then, the generation unit 129 generates the instruction sentence 381 using the case information extracted by the complementation. In this example, the other vehicle type, other vehicle position, other vehicle traveling direction, and other vehicle signal are unknown. Therefore, the extraction unit 128 selects, for example, the case with the highest number of accidents from the case information candidates extracted based on the display content of the event information input acceptance field 361, and complements it with the other vehicle type "four-wheeled vehicle," the other vehicle position "left," the other vehicle's traveling direction "straight," and the other vehicle's traffic light "red." The case information candidate selected by the extraction unit 128 is a case in which the fault ratio of the subject vehicle is "0%" and the fault ratio of the other vehicle is "100%." ​​The generation unit 129 then generates the instruction sentence 381 based on the display content of the event information input acceptance field 361, the complemented information, and the case information candidate selected as the case with the highest number of accidents.

[0110] FIG. 30 is a flowchart showing the flow of the evaluation information update display process. For example, the user checks the initial evaluation information in the evaluation information generation result display area 35 in FIG. 26 , the event information input acceptance field 362 in FIG. 27 , the case candidate selected in the case information candidate display selection area 37 in FIG. 28 , and the instruction sentence 381 in FIG. 29 , and realizes that the information about the other vehicle differs from the accident situation confirmed in the actual video data 111. Therefore, the user inputs information about the other vehicle into the event information input acceptance field 362. For example, the user inputs the other vehicle type "four-wheeled vehicle," the other vehicle position "oncoming," the other vehicle's direction of travel "right turn," and the other vehicle's traffic light "green" into the event information input acceptance field 362. In response, the accepting unit 127b accepts input for modifying the event information from the display terminal 200 (S41). Specifically, the accepting unit 127b accepts the input content into the event information input acceptance field 362.

[0111] The extraction unit 128 then re-extracts the relevant case information candidates based on the corrected event information (S42), and the display control unit 126b displays the re-extracted case information candidates (S43).

[0112] 31 is a diagram showing an example in which case information candidates are re-extracted following a correction to event information. As described above, the event information input acceptance field 362 indicates that the user has input the other vehicle type "four-wheeled vehicle," the other vehicle position "oncoming," the other vehicle's direction of travel "right turn," and the other vehicle's traffic light "green." Case information candidates 371 indicate case information candidates re-searched by the extraction unit 128 from the domain knowledge DB 114 based on the display contents of the event information display / edit area 36.

[0113] Thereafter, the accepting unit 127b accepts a selection of a case information candidate from the display terminal 200 through a user operation (S44). Specifically, it is assumed that the number 107 of the case information candidate 371 in Fig. 31 is selected.

[0114] The receiving unit 127b then receives a user operation to press the evaluation information generation button 341 from the display terminal 200 (S45). In response, the generation unit 129 generates an input prompt from the corrected event information and the selected case information candidate (S46). That is, the generation unit 129 generates an instruction sentence based on the display content of the event information display / editing area 36 in FIG. 31 and the case information candidate 371 selected in the case information candidate display / selection area 37. The acquisition unit 125b then inputs the generated instruction sentence, which is the input prompt, to the trained model 115 (S47). The acquisition unit 125b then acquires the evaluation information output from the trained model 115 (S48). The display control unit 126b then displays the acquired evaluation information on the display terminal 200 (S49). Specifically, the display control unit 126b sends a response to the display terminal 200 to display the acquired evaluation information in the evaluation information generation result display area 35. The display terminal 200 displays the received evaluation information in the evaluation information generation result display area 35 on the screen 201 .

[0115] 32 is a diagram showing an example of an updated input prompt. The previous prompt display field 38 displays an instruction sentence 382. Compared to the instruction sentence 381, the instruction sentence 382 has been changed in the section "Description of the accident situation:" to "Accident between vehicles with a green light and another vehicle with a green light." Furthermore, compared to the instruction sentence 381, the instruction sentence 382 has been changed to "Other vehicle type: four-wheeled vehicle," "Other vehicle position: oncoming vehicle," "Other vehicle direction: right turn," and "Other vehicle signal: green." Furthermore, compared to the instruction sentence 381, the instruction sentence 382 has been changed to "Own vehicle's fault percentage: 20%" and "Other vehicle's fault percentage: 80%."

[0116] FIG. 33 is a diagram showing an example of updated evaluation information. The evaluation information generation result display area 35 displays evaluation information 113b. Compared to the evaluation information generation result display area 35 in FIG. 26, the following changes have been made to the evaluation information 113b, similar to the instruction sentence 382. Specifically, in the "accident type" section, "an accident between a vehicle with a green light and a vehicle with a red light" in the evaluation information 113 has been changed to "an accident between two vehicles with a green light and another vehicle with a green light" in the evaluation information 113b. Furthermore, in the "accident fault ratio determination result," "the fault ratio of the vehicle itself is 0%, while the fault ratio of the other vehicle is 100%" in the evaluation information 113 has been changed to "the fault ratio of the vehicle itself is 20%, while the fault ratio of the other vehicle is 80%" in the evaluation information 113b. In addition, the "accident fault ratio determination result" was changed from "an accident in which a four-wheeled vehicle traveling straight on a green light collided with another four-wheeled vehicle attempting to travel straight on a red light" in the evaluation information 113 to "an accident in which a four-wheeled vehicle traveling straight on a green light collided with another four-wheeled vehicle attempting to turn right on a green light" in the evaluation information 113b. In addition, the "accident fault ratio determination result" was changed from "both vehicles were traveling straight at the intersection, so the basic fault ratio is determined to be 0:100" in the evaluation information 113 to "both vehicles were entering on green lights, so the basic fault ratio is determined to be 20:80" in the evaluation information 113b. In addition, the "accident fault ratio determination result" was changed from "the fault ratio of the vehicle is determined to be 0%" in the evaluation information 113 to "the fault ratio of the vehicle is determined to be 20%" in the evaluation information 113b.

[0117] Until now, when users analyzing the causes of accidents wanted to create evaluation information (such as a statement explaining the degree of fault) based on the results of their analysis of the accident causes, they had to personally review the video footage and manually search for related knowledge to create the resulting text (evaluation information, judgment results), which tended to be time-consuming due to the large number of searches required. Furthermore, differences in the know-how of the analyzing users could result in differences in the quality of the results of the work (created evaluation information). Furthermore, various data sources were required to accurately grasp the accident situation.

[0118] In response to this, the technology disclosed herein can generate evaluation information about a currently occurring event based on information about the event detected from video data and past case information about the event, and present the evaluation information to the user. Therefore, even if the data source is limited to, for example, video data captured by an in-vehicle camera on the user's own vehicle, the technology can assist in understanding the situation of the event by presenting highly accurate assessment results (evaluation information) about the currently occurring event. In particular, the technology disclosed herein can generate evaluation information using event information detected from video data and case information searched from the domain knowledge DB 114 using the event information. Specifically, the technology disclosed herein can automatically generate instruction sentences, which are input information to the LLM, using the event information and case information. Therefore, by inputting instruction sentences into the LLM, highly accurate evaluation information that takes into account the currently occurring event and past cases can be generated and presented. Furthermore, the user can make partial corrections to the event information by checking the presented evaluation information and the instruction sentences used to generate the evaluation information. Then, by making some corrections to the event information and searching again from the domain knowledge DB 114, that is, by using the filtered case information, it is possible to generate and present more accurate evaluation information. In this way, the technology according to the present disclosure can assist in understanding the situation of an event.

[0119] (Other Embodiments) Note that the user interface of embodiment 5 or 6 may be used for the screen of embodiment 4. Similarly, the user interface of embodiment 6 may be used for the screen of embodiment 5. Furthermore, in embodiments 2, 3, 5, and 6, the process of identifying the area of ​​the event-related object may be performed by a method other than that of embodiment 1 or 4. Furthermore, in the above description, a vehicle collision accident has been used as the subject of the event (phenomenon), and therefore vehicles and traffic lights have been given as examples of event-related objects. However, if the event is an accident resulting in injury or death, the event-related object may include a person. Furthermore, the event is not limited to the above. In such a case, the event-related object may also be changed as appropriate depending on the event.

[0120] The information processing devices 1, 1a, and 1b each include a processor, a memory, and a storage device (not shown). The storage device stores a computer program that implements the processing of the information processing method shown in FIG. 2, 4, or 6, for example. The processor then loads the computer program from the storage device into the memory and executes the computer program. This allows the processor to implement the functions of the selection unit 11, the area estimation unit 12, the object estimation unit 13, and the identification unit 14. Alternatively, the processor also implements the functions of the display control unit 15a or 15b.

[0121] Alternatively, each component of the information processing device 1 may be realized by dedicated hardware. Furthermore, some or all of the components of each device may be realized by general-purpose or dedicated circuits, processors, etc., or a combination thereof. These may be configured by a single chip, or by multiple chips connected via a bus. Some or all of the components of each device may be realized by a combination of the above-mentioned circuits, etc., and a program. Furthermore, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field-Programmable Gate Array), a quantum processor (quantum computer control chip), etc., may be used as the processor.

[0122] Furthermore, when some or all of the components of the information processing device 1 or the like are realized by multiple information processing devices, circuits, etc., the multiple information processing devices, circuits, etc. may be centrally or distributedly arranged. For example, the information processing devices, circuits, etc. may be realized as a client-server system, a cloud computing system, or the like, in a form in which each is connected via a communication network. Furthermore, the functions of the information processing device 1 or the like may be provided in a SaaS (Software as a Service) format.

[0123] Fig. 34 is a block diagram showing the hardware configuration of the information processing device 100. The hardware configurations of the information processing devices 100a and 100b are also the same as that shown in Fig. 34. The information processing device 100 etc. includes a memory 101, a processor 102, and a network interface 103.

[0124] The memory 101 is configured by a combination of volatile memory and non-volatile memory. The volatile memory is, for example, a volatile storage device such as RAM, and is a storage area for temporarily holding information when the processor 102 is operating. The non-volatile memory is, for example, a non-volatile storage device such as flash memory. The memory 101 stores at least a computer program that implements at least part of the processing of the information processing method of the information processing device 100, etc., according to the present disclosure. The memory 101 may store the above-mentioned video data 111, event-related object information 1121, etc., evaluation information 113, domain knowledge DB 114, and trained model 115.

[0125] The processor 102 is a control device that controls each component of the information processing device 100, etc. The processor 102 reads and executes software (computer programs) from the memory 101. As a result, the processor 102 realizes the functions of the event detection unit 121, the area estimation unit 122, the object detection and tracking unit 123, the identification unit 124, the acquisition unit 125, etc., the display control unit 126, etc., and the reception unit 127, etc. (the event information generation unit 120a, etc., the analysis unit 1201, the extraction unit 128, and the generation unit 129). In other words, the processor 102 performs at least a part of the processing of the information processing method in the information processing device 100, etc. according to the present disclosure. The processor 102 may be, for example, a microprocessor, a multi-processing unit (MPU), or a central processing unit (CPU). The processor 102 may also include multiple processors.

[0126] The network interface 103 may be used to communicate with network nodes. The network interface 103 may include, for example, a network interface card (NIC) conforming to the IEEE 802.3 series. IEEE stands for Institute of Electrical and Electronics Engineers. The network interface 103 may also include a wireless local area network (LAN), a wired LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0127] The program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.

[0128] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0129] Each drawing is merely an example for describing one or more embodiments. Each drawing may not relate to only one particular embodiment, but may also relate to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.

[0130] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes: (Supplementary Note A1) An information processing device comprising: a selection means for selecting images in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period; a region estimation means for estimating a region in the selected images that contributed to the detection of the event; an object estimation means for estimating a position of an object included in at least the selected images from the sequence of multiple images; and an identification means for identifying an event-related object related to the event based on the position and region of the object estimated from the selected images. (Supplementary Note A2) The information processing device according to Supplementary Note A1, wherein the identification means identifies the event-related object based on an overlapping range between the position of the object and the region. (Supplementary Note A3) The information processing device according to Supplementary Note A1 or A2, wherein the identification means, when a first region of the event is estimated in a first image among the two or more images selected from the plurality of image sequences, but a position of the object is not estimated, identifies the event-related object included in the first image based on a second region of the event and a position of the object estimated in a second image and the first region. (Supplementary Note A4) The information processing device according to Supplementary Note A1 or A2, wherein the object estimation means estimates the position of the object in each image by tracking a specific object for the two or more images selected from the plurality of image sequences, and when a third region of the event is estimated in a third image among the two or more selected images, but a position of the object is not estimated, identifies the event-related object included in the third image based on the position of the object in each image estimated by the tracking and the third region. (Appendix A5) The information processing device described in Appendix A1 or A2, wherein the object estimation means further estimates the position of an object included in a fourth image that was not selected from the plurality of image sequences, and the identification means identifies the event-related object included in the fourth image based on the position and area of ​​the object estimated from the selected image and the position of the object estimated from the fourth image.(Appendix A6) The information processing device according to any one of Appendices A1 to A5, wherein the identification means identifies each of the plurality of event-related objects related to the same detected event. (Appendix A7) The information processing device according to any one of Appendices A1 to A6, wherein the identification means identifies each of the plurality of event-related objects related to each of the detected events. (Appendix A8) The information processing device according to any one of Appendices A1 to A7, further comprising display control means for outputting to a display device so as to display on a screen in an identifiable manner an area corresponding to the event-related object within an image in which at least the position and the area of ​​the object have been estimated. (Appendix B1) An information processing method, comprising: a computer selecting images in which a predetermined event has been detected from a plurality of image sequences taken within a predetermined period; estimating an area within the selected images that contributed to the detection of the event; estimating a position of an object included in at least the selected images from the plurality of image sequences; and identifying an event-related object related to the event based on the position and the area of ​​the object estimated from the selected images. (Supplementary Note C1) An information processing program that causes a computer to execute the following steps: a selection process that selects images in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period, an event estimation process that estimates an area in the selected images that contributed to the detection of the event, an object estimation process that estimates a position of an object included in at least the selected images from the sequence of multiple images, and an identification process that identifies an event-related object related to the event based on the position and area of ​​the object estimated from the selected images. (Supplementary Note D1) An information processing device comprising: display control means that displays, in images in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period, on a predetermined screen an area corresponding to an event-related object related to the event that has been identified based on the area that contributed to the detection, in a manner that is identifiable within the image, and displays on the screen an event area bar that indicates an event section in which the event was detected during the predetermined period.(Appendix D2) The information processing device according to Appendix D1, wherein the display control means displays a plurality of the event area bars on the screen. (Appendix D3) The information processing device according to Appendix D2, wherein the display control means displays a plurality of the event area bars on the screen so that the positional relationship of each event section is distinguishable based on the length of the predetermined period. (Appendix D4) The information processing device according to Appendix D2 or D3, wherein the display control means displays a plurality of the event area bars on the screen indicating the event sections corresponding to each of the identified event-related objects. (Appendix D5) The information processing device according to any one of Appendix D1 to Appendix D4, wherein the display control means displays a plurality of the event area bars on the screen corresponding to each of the detected events. (Appendix D6) The information processing device according to Appendix D5, wherein the display control means displays a plurality of the event area bars on the screen, wherein each of a plurality of states of the event-related object that changed during the predetermined period is defined as the detected events, and the duration of each state is defined as the event section. (Appendix D7) The information processing device according to any one of Appendices D1 to D6, wherein the display control means displays a text related to the detected event and displays a display / edit field on the screen for accepting edits to the text. (Appendix D8) The information processing device according to any one of Appendices D1 to D6, wherein the display control means accepts designation of at least some of a plurality of analysis perspectives of the event including the event-related object, and the display control means, in response to accepting the designation, controls switching between displaying and hiding on the screen for each of the area and the event area bar based on the accepted analysis perspective.(Appendix E1) An information processing system comprising an information processing device and a display device, wherein the information processing device displays, in an image in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period, an area corresponding to an event-related object related to the event, identified based on an area that contributed to the detection, in an identifiable manner within the image, on a screen of the display device, and displays on the screen an event area bar indicating an event section in which the event was detected during the predetermined period. (Appendix F1) An information processing method, wherein a computer displays, in an image in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period, an area corresponding to an event-related object related to the event, identified based on an area that contributed to the detection, in an identifiable manner within the image, and displays on the screen an event area bar indicating an event section in which the event was detected during the predetermined period. (Appendix G1) An information processing program that causes a computer to execute a display control process to display, on a predetermined screen, an area corresponding to an event-related object related to a predetermined event, identified based on an area that contributed to the detection, in an image in which a predetermined event has been detected from a sequence of multiple images captured during a predetermined period, in an identifiable manner within the image, and to display on the screen an event area bar indicating an event section in which the event was detected during the predetermined period. (Appendix H1) An information processing device comprising: display control means that displays, on a predetermined screen, event information including an event-related object related to the event, identified based on an area that contributed to the detection of the predetermined event in the image, case information indicating cases related to the event, and evaluation information related to the event output from a predetermined trained model using the event information and input information based on the case information. (Appendix H2) The information processing device according to Appendix H1, wherein the case information is information narrowed down based on the event information from a set of multiple pieces of case information. (Appendix H3) The information processing device according to Appendix H1 or H2, wherein the display control means displays the input information on the screen.(Appendix H4) The information processing device according to any one of Appendices H1 to H3, further comprising an acquisition means for acquiring the evaluation information output from the trained model by inputting the input information to the trained model. (Appendix H5) The information processing device according to any one of Appendices H1 to H4, further comprising a first generation means for generating, as the input information, an instruction sentence for the trained model using the event information and the case information. (Appendix H6) The information processing device according to any one of Appendices H1 to H5, wherein the display control means displays, on the screen, a reception field for receiving input for correcting the event information. (Appendix H7) The information processing device according to any one of Appendices H1 to H6, wherein, when the evaluation information is updated by the trained model in response to an update of at least one of the event information or the case information, the display control means displays the updated evaluation information on the screen. (Appendix H8) The information processing device of any one of Appendices H1 to H7, further comprising second generation means for generating the event information including information on an event-related object related to the event identified based on an area that contributed to the detection in the image in which the event was detected from a sequence of multiple images captured during a predetermined period. (Appendix I1) An information processing system comprising an information processing device and a display device, wherein the information processing device causes the display device to display event information including an event-related object related to the event identified based on an area that contributed to the detection of the predetermined event in the image, case information indicating examples related to the event, and evaluation information related to the event output from a predetermined trained model using the event information and input information based on the event information and the case information. (Appendix J1) An information processing method, wherein a computer causes a predetermined screen to display event information including an event-related object related to the event identified based on an area that contributed to the detection of the predetermined event in the image, case information indicating examples related to the event, and evaluation information related to the event output from a predetermined trained model using the event information and input information based on the case information.(Appendix K1) An information processing program that causes a computer to execute a display control process that displays on a specified screen event information including event-related objects related to a specified event identified based on an area that contributed to the detection of the event in an image, case information indicating examples related to the event, and evaluation information related to the event output from a specified trained model using the event information and input information based on the case information.

[0131] Some or all of the elements (e.g., configurations and functions) described in Appendix A2 to Appendix A8 that are dependent on Appendix A1 (e.g., apparatus) may also be dependent on Appendix B1 (e.g., method) and Appendix C1 (e.g., program) in the same dependency relationship as Appendix A2 to Appendix A8. Some or all of the elements (e.g., configurations and functions) described in Appendix D2 to Appendix D8 that are dependent on Appendix D1 (e.g., apparatus) may also be dependent on Appendix E1 (e.g., system), Appendix F1 (e.g., method), and Appendix G1 (e.g., program) in the same dependency relationship as Appendix D2 to Appendix D8. Some or all of the elements (e.g., configurations and functions) described in Appendix H2 to Appendix H8 that are dependent on Appendix H1 (e.g., apparatus) may also be dependent on Appendix I1 (e.g., system), Appendix J1 (e.g., method), and Appendix K1 (e.g., program) in the same dependency relationship as Appendix H2 to Appendix H8. Some or all of the elements described in any appendix may be applied to a variety of hardware, software, recording means for recording software, systems, and methods.

[0132] This application claims priority based on Japanese Patent Application No. 2024-135298, filed August 14, 2024, the disclosure of which is incorporated herein in its entirety by reference.

[0133] 1 Information processing device, 1a Information processing device, 1b Information processing device, 11 Selection unit, 12 Area estimation unit, 13 Object estimation unit, 14 Identification unit, 15a Display control unit, 15b Display control unit, 1000 Information processing system, 100 Information processing device, 100a Information processing device, 100b Information processing device, 200 Display terminal, 201 Screen, N Communication network, 110 Storage unit, 111 Video data, 1111 Image, 111n Image, 112 Event information, 1121 Event-related object information, 112m Event-related object information, 113 Rating information, 113b Rating information, 114 Domain knowledge DB, 115 Trained model, 120a Event information generation unit, 120b Event information generation unit, 121 Event detection unit, 122 Area estimation unit, 123 Object detection and tracking unit, 124 Identification unit, 125 Acquisition unit, 125b Acquisition unit, 126 Display control unit, 126a Display control unit, 126b Display control unit, 127 Reception unit, 127a Reception unit, 127b Reception unit, 1201 Analysis unit, 128 Extraction unit, 129 Generation unit, 101 Memory, 102 Processor, 103 Network interface, 3 Event visualization screen, 3a Event visualization screen, 3b Event visualization screen, 31 Video display area, 32 Visualization support area, 32a Visualization support area, 33 Evaluation information display and editing area, 34 User operation reception area, 35 Evaluation information generation result display area, 36 Event information display and editing area, 37 Case information candidate display and selection area, 321 Playback loading bar, 3211 Seek bar, 3212 Playback position, 3213 Frame movement instruction field, 322 Event area bar, 323 Event area bar, 324 Event area bar, 3241 Sub-event area bar, 3242 sub-event area bar, 325 event area bar, 3261 analysis viewpoint selection field, 3262 analysis viewpoint selection field, 3262b analysis viewpoint selection field, 3263 analysis viewpoint selection field, 3264 analysis viewpoint selection field, 341 evaluation information generation button, 342 evaluation information copy button, 343 previous prompt confirmation button, 361 event information input acceptance field, 362 event information input acceptance field, 371 case information candidate, 38 previous prompt display field, 381 instruction sentence, 382 instruction sentence, 41 return to top button, 42 return to previous chapter button,43 Back one frame button, 44 Play button, 45 Forward one frame button, 46 Forward next chapter button, 47 Seek bar display switch button, 48 Video save button, 49 Marking display switch button, 51 Frame image, 52 Frame image, 53 Frame image, 54 Frame image, 55 Frame image, 56 Frame image, 57 Frame image, 61 Mark, 62 Mark,

Claims

1. An information processing device comprising: a selection means for selecting images in which a predetermined event has been detected from a sequence of multiple images taken during a predetermined period; a region estimation means for estimating a region within the selected images that contributed to the detection of the event; an object estimation means for estimating a position of an object included in at least the selected images from the sequence of multiple images; and an identification means for identifying an event-related object related to the event based on the position and region of the object estimated from the selected images.

2. The information processing device according to claim 1, wherein the identifying means identifies the event-related object based on the range in which the position of the object overlaps with the area.

3. The information processing device of claim 1 or 2, wherein, when a first region of the event is estimated in a first image of the two or more images selected from the sequence of images, but the position of the object is not estimated, the identification means identifies the event-related object contained in the first image based on the second region of the event and the position of the object estimated in the second image and the first region.

4. The information processing device of claim 1 or 2, wherein the object estimation means estimates the position of a specific object in each of the two or more images selected from the sequence of images by tracking the object, and wherein the identification means, when a third region of the event is estimated in a third image among the two or more selected images but the position of the object is not estimated, identifies the event-related object contained in the third image based on the position of the object in each image estimated by the tracking and the third region.

5. An information processing device as described in claim 1 or 2, wherein the object estimation means further estimates the position of an object included in a fourth image that was not selected from the plurality of image sequences, and the identification means identifies the event-related object included in the fourth image based on the position and area of ​​the object estimated from the selected image and the position of the object estimated from the fourth image.

6. The information processing device according to claim 1 or 2, wherein the identification means identifies each of the plurality of event-related objects related to the same detected event.

7. The information processing device according to claim 1 or 2, wherein the identification means identifies each of the plurality of event-related objects associated with each of the plurality of detected events.

8. An information processing device according to claim 1 or 2, further comprising a display control means for outputting to a display device so as to identifiably display an area corresponding to the event-related object within an image in which at least the position and area of ​​the object have been estimated.

9. An information processing method in which a computer selects images in which a specified event has been detected from a sequence of multiple images taken over a specified period of time, estimates an area in the selected images that contributed to the detection of the event, estimates a position of an object included in at least the selected images from the sequence of multiple images, and identifies an event-related object related to the event based on the estimated position and area of ​​the object from the selected images.

10. The information processing method according to claim 9, wherein the computer, in the step of identifying the event-related object, identifies the event-related object based on the range in which the position of the object overlaps with the area.

11. An information processing method as described in claim 9 or 10, wherein, in the step of identifying the event-related object, if a first region of the event is estimated in a first image of the two or more images selected from the plurality of image sequences but the position of the object is not estimated, the computer identifies the event-related object contained in the first image based on the second region of the event and the position of the object estimated in the second image and the first region.

12. An information processing method as described in claim 9 or 10, wherein the computer, in the step of estimating the position of the object, estimates the position of the object in each image by tracking a specific object in two or more images selected from the plurality of image sequences, and, in the step of identifying the event-related object, if a third region of the event is estimated in a third image among the two or more selected images but the position of the object is not estimated, identifies the event-related object included in the third image based on the position of the object in each image estimated by the tracking and the third region.

13. An information processing method as described in claim 9 or 10, wherein the computer, in the step of estimating the position of the object, further estimates the position of the object included in the fourth image that was not selected from the plurality of image sequences, and, in the step of identifying the event-related object, identifies the event-related object included in the fourth image based on the position and area of ​​the object estimated from the selected image and the position of the object estimated from the fourth image.

14. The information processing method according to claim 9 or 10, wherein the computer, in the step of identifying the event-related objects, identifies each of the plurality of event-related objects related to the same detected event.

15. An information processing method according to claim 9 or 10, wherein the computer, in the step of identifying the event-related objects, identifies each of the plurality of event-related objects associated with each of the plurality of detected events.

16. An information processing method according to claim 9 or 10, wherein the computer outputs to a display device so as to identifiably display an area corresponding to the event-related object within an image in which at least the position and area of ​​the object have been estimated.

17. An information processing program that causes a computer to execute the following steps: a selection process that selects images in which a specified event has been detected from a sequence of multiple images taken over a specified period; an event estimation process that estimates an area within the selected images that contributed to the detection of the event; an object estimation process that estimates the position of an object included in at least the selected images from the sequence of multiple images; and an identification process that identifies an event-related object related to the event based on the position and area of ​​the object estimated from the selected images.

18. An information processing program according to claim 17, wherein the identification process identifies the event-related object based on the range in which the position of the object overlaps with the area.

19. An information processing program as described in claim 17 or 18, wherein the identification process, when a first region of the event is estimated in a first image of the two or more images selected from the sequence of images but the position of the object is not estimated, identifies the event-related object contained in the first image based on the second region of the event and the position of the object estimated in the second image and the first region.

20. The information processing program of claim 17 or 18, wherein the object estimation process estimates the position of a specific object in each of the two or more images selected from the sequence of images by tracking the specific object, and wherein the identification process, when a third region of the event is estimated in a third image among the two or more selected images but the position of the object is not estimated, identifies the event-related object contained in the third image based on the position of the object in each image estimated by the tracking and the third region.

Citation Information

Patent Citations

  • Information processing apparatus and image processing method

    JP2022182149A

  • Information processing device, information processing method, and program

    JP2023065855A

  • Information processing device, information processing method, and information processing program

    WO2014050518A1

  • Recording control device, recording control method, and program

    WO2021152985A1