Video processing system, trained model, video processing method, and video processing program
Patent Information
- Application Number
- PCT/JP2025/045878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2025-12-26
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025045878_01102026_PF_FP_ABST
Abstract
Description
Video processing system, trained model, video processing method and video processing program
[0001] The present disclosure relates to a video processing system, a trained model, a video processing method and a video processing program.
[0002] Techniques for storing video of a scene including a predetermined event from video captured by an imaging means have been developed (for example, Patent Document 1). The system disclosed in Patent Document 1 detects impacts, sudden starts, sudden braking, etc. of a moving body as predetermined events. Then, the system stores video for a predetermined period including the time point at which the detected event occurs.
[0003] Japanese Unexamined Patent Publication No. 2024-056960
[0004] However, although the system disclosed in Patent Document 1 can identify a predetermined event, it cannot identify the cause of the occurrence of the event. Therefore, a user of the system needs to visually check the video to identify the cause of the event. However, the work of checking a large amount of video requires enormous time and effort, placing a large burden on the user. Against this background, there is a demand for a technique that reduces the burden of video checking work.
[0005] The present disclosure has been made to solve such problems, and an object thereof is to provide a video processing system, a video processing method, and a video processing program that can efficiently extract video requiring visual confirmation.
[0006] The video processing system according to the present disclosure includes: an imperative sentence generation unit that generates an imperative sentence for detecting a scene corresponding to a cause label based on one or more cause labels including information indicating a cause of occurrence of a dangerous situation event; and a detection unit that executes detection of the scene on video based on the imperative sentence and outputs a detection result.
[0007] The trained model relating to this disclosure is a trained model that takes video footage and one or more cause labels containing information indicating the cause of an event in a dangerous situation as input, and is trained to machine learn to detect scenes corresponding to the cause labels in the video footage, and takes the video footage and one or more of the cause labels as input, and outputs a detection result and an explanatory text containing an explanation of the detection result.
[0008] The video processing method relating to this disclosure comprises the steps of: acquiring video; acquiring one or more cause labels including information indicating the cause of an event in a dangerous situation; generating a command statement for detecting a scene corresponding to the cause label based on the cause label; performing scene detection on the video based on the command statement; and outputting the detection result obtained from the detection.
[0009] The video processing program relating to this disclosure causes a computer to perform the following steps: acquire video; acquire one or more cause labels including information indicating the cause of an event in a dangerous situation; generate a command statement to detect a scene corresponding to the cause label based on the cause label; perform the detection of the scene on the video based on the command statement; and output the detection result obtained from the detection.
[0010] This disclosure provides a video processing system, a video processing method, and a video processing program that enable efficient video verification.
[0011] This figure shows an example of the configuration of the video processing system related to this disclosure. This figure shows another example of the configuration of the video processing system related to this disclosure. This figure shows an example of the configuration of the cause label list related to this disclosure. This figure shows an example of the instruction statement related to this disclosure. This figure shows an example of the configuration of the user interface that displays the detection results and explanatory text related to this disclosure. This flowchart shows an example of the video processing method related to this disclosure. This figure shows another example of the configuration of the video processing system related to this disclosure. This figure shows another example of the configuration of the user interface that displays the detection results and explanatory text related to this disclosure. This flowchart shows another example of the video processing method related to this disclosure. This block diagram shows an example of the hardware configuration that realizes the video processing of the video processing system related to this disclosure.
[0012] (Embodiment 1) <Configuration of the video processing system> An example of a video processing system will be described below with reference to Figure 1. Figure 1 is a diagram showing an example of the configuration of a video processing system according to the present disclosure. As shown in Figure 1, the video processing system 1 comprises an instruction generation unit 32 and a detection unit 33. The video processing system 1 is a system that takes video and cause labels as input. The video processing system 1 then detects scenes corresponding to the input cause labels in the input video.
[0013] A cause label is a label that contains information about the cause of a hazardous event. Hazardous events are, for example, accidents that occur in situations such as driving a car or working in a factory, or events that are factors that cause accidents. More specifically, in driving a car, examples of hazardous events include sudden stops, sudden braking, sudden steering, and collisions. In a factory work environment, examples of hazardous events include workers falling or workers coming into contact with obstacles.
[0014] The command generation unit 32 generates a command that causes the detection unit 33 to detect a scene corresponding to one or more cause labels, based on the cause labels. The command includes multiple options. These options include at least an option corresponding to a cause label and an option for when no cause label is detected in the video. The command generated by the command generation unit 32 may be displayed on a display means (not shown), such as a monitor, for presentation to the user. The user may then modify the content of the command and add or change options. The command generation unit 32 inputs the generated command to the detection unit 33.
[0015] The detection unit 33 executes a process to detect scenes corresponding to the cause label in the video based on the command input from the command generation unit 32. The detection unit 33 then outputs the detection results. If the detection unit 33 detects a scene that corresponds to an option corresponding to the cause label, it outputs the detected scene and the option corresponding to that scene, i.e., the cause label, as the detection results. In this way, the detection unit 33 can filter scenes corresponding to the cause label, i.e., scenes that require user confirmation, from the video.
[0016] On the other hand, if the detection unit 33 does not detect a scene that corresponds to an option that matches the cause label, it outputs the options for when the cause label is not detected in the video as the detection result.
[0017] The detection unit 33 may also use a pre-trained model that has been trained to take video and command text as input and output detection results. For example, the detection unit 33 may use a Vision-Language Model (VLM) as the pre-trained model.
[0018] As described above, the video processing system 1 acquires video footage and one or more cause labels containing information indicating the cause of the dangerous event. The video processing system 1 then detects scenes in the video that correspond to the cause labels. When a scene corresponding to a cause label is detected, the video processing system 1 outputs the detected scene and the corresponding cause label as the detection result. In this way, by outputting scenes corresponding to cause labels, i.e., scenes that require user verification, the video processing system 1 enables the user to efficiently perform the video verification task.
[0019] The video processing system 1 may also use a pre-trained model that has been trained through machine learning to input video and one or more cause labels containing information indicating the cause of the dangerous situation event, detect scenes corresponding to the cause labels in the video, and output the detection results.
[0020] (Embodiment 2) <Configuration of the video processing system> Another example of the video processing system will be described below with reference to Figure 2. Figure 2 is a diagram showing another example of the configuration of the video processing system according to this disclosure. The video processing system 1 can be applied to the analysis of hazardous situations, such as driving a car or working in a factory. In this embodiment 2, an example of applying the video processing system 1 to driving a car will be described.
[0021] As shown in Figure 2, the video processing system 1 is a processing device comprising a video input UI unit 10, a cause label input UI unit 20, a video processing unit 30, and a detection result display UI unit 40. The video processing system 1 according to this disclosure may be implemented by a single device or by distributed processing using multiple devices. For example, the video processing system 1 may consist of a user terminal operated by a user and a processing device that performs video processing. In this case, for example, the video input UI unit 10, the cause label input UI unit 20, and the detection result display UI unit 40 are provided on the user terminal, and the video processing unit 30 is provided on the processing device.
[0022] The video input UI unit 10 is a user interface for the user to input video and supplementary video information. The video input UI unit 10 is displayed on a display means (not shown), such as a monitor. The video input UI unit 10 comprises a video input unit 11 and a supplementary information input unit 12.
[0023] The video input unit 11 is a UI (User Interface) element that accepts video input. In this second embodiment, the video input to the video input unit 11 includes at least video footage of the interior of the car. This video is captured, for example, by a drive recorder installed in the car. In addition to the video footage of the interior of the car, the video input unit 11 may also receive video footage of the exterior of the car captured by the drive recorder. The following describes an example in which both video footage of the interior of the car and video footage of the exterior of the car are input.
[0024] The supplementary information input unit 12 is a UI element that accepts input of supplementary video information to supplement the video input to the video input unit 11, such as vehicle position data and vehicle acceleration data. Position data is acquired, for example, by a drive recorder or a position sensor installed in the vehicle itself. Acceleration data is acquired, for example, by an acceleration sensor in the drive recorder. Preferably, the video, position data, and acceleration data are each assigned a timestamp indicating the time they were captured or acquired. This allows the video processing unit 30 to synchronize the video, position data, and acceleration data by time. The information received by the video input unit 11 and the supplementary information input unit 12 is transmitted to the video processing unit 30.
[0025] The cause label input UI unit 20 is a user interface for the user to input cause labels. The cause label input UI unit 20 is displayed on a display means such as a monitor. The cause label input UI unit 20 includes a cause label addition unit 21 and a cause label selection unit 22.
[0026] A cause label is a label that contains information about the cause of a dangerous event. In this second embodiment, examples of dangerous events include sudden stops, sudden braking, sudden steering, and collisions. Examples of cause labels include driver actions such as "using a smartphone," "talking to a passenger," "smoking," "eating or drinking," and "falling asleep."
[0027] The cause label addition unit 21 is a UI element for the user to add a new cause label to the cause label list. Specifically, the cause label addition unit 21 is a text box UI element. Figure 3 is a diagram showing an example of the configuration of the cause label list according to this disclosure. As shown in Figure 3, the cause label list stores one or more cause labels. The cause label entered into the cause label addition unit 21 is registered in the cause label list shown in Figure 3.
[0028] The cause label selection unit 22 is a UI element for the user to select a cause label registered in the cause label list shown in Figure 3. Specifically, the cause label selection unit 22 is a UI element such as a list box, a dropdown list, a checkbox, or radio buttons. The user selects one or more cause labels in the cause label selection unit 22. The cause labels selected in the cause label selection unit 22 are then transmitted to the video processing unit 30.
[0029] The video processing unit 30 detects scenes in the video that indicate the cause of a dangerous event. As shown in Figure 2, the video processing unit 30 comprises an extraction unit 31, a command generation unit 32, a detection unit 33, and a display control unit 34.
[0030] The extraction unit 31 extracts video footage for a predetermined period from the video input to the video input UI unit 10, based on the time of occurrence of a dangerous situation event. More specifically, the extraction unit 31 first detects the occurrence of a dangerous situation event in the video input to the video input UI unit 10, for example, based on acceleration data input to the video input UI unit 10. The extraction unit 31 detects the occurrence of a dangerous situation event, for example, from the magnitude of acceleration or the acceleration pattern. If the occurrence of a dangerous situation event is detected, the extraction unit 31 extracts video footage for a predetermined period based on the time of occurrence of the event. The extraction unit 31 may also detect the end time of the event based on supplementary information such as the acceleration data of the vehicle, and extract video footage up to that end time. The extraction unit 31 then inputs the extracted video footage to the detection unit 33. If the extraction unit 31 detects multiple dangerous situation events in the video input to the video input UI unit 10, it extracts video footage for each detected event. The extraction unit 31 then inputs the extracted video data to the detection unit 33.
[0031] The command generation unit 32 generates a command to cause the detection unit 33 to detect a scene corresponding to a cause label, based on one or more cause labels selected in the cause label input UI unit 20. The command includes multiple options. These options include at least an option corresponding to a cause label and an option for when no cause label is detected in the video. The command generated by the command generation unit 32 may be displayed on a display means such as a monitor for presentation to the user. The user may then modify the content of the command by, for example, adding or changing options. The command generation unit 32 inputs the generated command to the detection unit 33.
[0032] Figure 4 shows an example of an instruction statement related to this disclosure. More specifically, Figure 4 is an example of an instruction statement generated by the instruction statement generation unit 32 when the cause labels "using a smartphone," "smoking a cigarette," and "talking to a passenger" are selected in the cause label input UI unit 20. The instruction statement shown in Figure 4 first detects the driver and then detects whether the driver performed an action corresponding to the cause label. The options "1. Using a smartphone," "2. Smoking a cigarette," and "3. Talking to a passenger" shown in Figure 4 are options corresponding to the cause label selected in the cause label input UI unit 20. The option "4. Camera malfunction" is an option used when no driver is detected. The option "5. No applicable action" is an option used when a driver is detected, but the action corresponding to the cause label is not detected.
[0033] The detection unit 33, based on the command input from the command generation unit 32, executes a process to detect scenes corresponding to the cause label in the video input from the extraction unit 31. The detection unit 33 then outputs the detection results. If the detection unit 33 detects a scene that corresponds to an option corresponding to the cause label, it outputs the detected scene and the option corresponding to that scene, i.e., the cause label, as the detection results. In this way, the detection unit 33 can filter scenes corresponding to the cause label, i.e., scenes that require user confirmation, from the video.
[0034] On the other hand, if the detection unit 33 does not detect a scene that corresponds to an option that matches the cause label, it outputs as a detection result the options for when the cause label is not detected in the video, or the options for when the driver is not detected.
[0035] Furthermore, the detection unit 33 outputs an explanatory text containing an explanation of the detection result along with the detection result. More specifically, the detection unit 33 outputs an explanatory text in natural language text based on the selection of the detection result. The explanatory text may also include a sentence indicating the basis for obtaining the detection result. The detection unit 33 transmits the outputted detection result and explanatory text to the detection result display UI unit 40.
[0036] The detection unit 33 may use a pre-trained model that takes video and a command as input and outputs a detection result and an explanatory text containing an explanation of the detection result. The detection unit 33 may use, for example, a Vision-Language Model (VLM) as the pre-trained model.
[0037] Furthermore, when the detection unit 33 receives multiple video inputs from the extraction unit 31, that is, when the extraction unit 31 detects multiple dangerous events within the video input to the video input UI unit 10, the detection unit 33 performs detection processing on each video input from the extraction unit 31. The detection unit 33 then outputs a detection result and an explanatory text for each video. The detection unit 33 transmits the outputted multiple detection results and explanatory texts to the detection result display UI unit 40.
[0038] The display control unit 34 controls the screen displayed on a display means such as a monitor. More specifically, the display control unit 34 first displays a video input UI unit 10 on the display means to accept video input. After accepting video input, the display control unit 34 displays a cause label input UI unit 20 on the display means to accept the selection of a cause label. Then, after the detection unit 33 has performed the detection process, it displays a detection result display UI unit 40, which will be described later, on the display means to present the detection result and explanatory text to the user.
[0039] The detection result display UI unit 40 is a user interface for presenting detection results and explanatory text to the user. The detection result display UI unit 40 is displayed on a display means such as a monitor. Figure 5 is a diagram showing an example of the configuration of a user interface for displaying detection results and explanatory text according to the present disclosure. As shown in Figure 5, the detection result display UI unit 40 includes a first display area 41, a second display area 42, a third display area 43, a fourth display area 44, a fifth display area 45, and a sixth display area 46.
[0040] The first display area 41 is an area that displays video footage of the outside of the car, which has been input to the video input unit 11. The first display area 41 may also display UI elements such as a button to stop the video displayed in the area, or a button to play the video. Furthermore, when the user operates the display button 420 for another event in the second display area 42, which will be described later, the video displayed in the first display area 41 switches to the video corresponding to that event.
[0041] The second display area 42 is an area that displays a list of dangerous events detected by the extraction unit 31. The second display area 42 displays event identification information, time and location information, and operation buttons. The event identification information indicates the type of dangerous event, such as sudden stop, sudden braking, sudden steering, or collision. The time and location information is the timestamps of the start and end points of the video extracted by the extraction unit 31 from the video input UI unit 10.
[0042] The operation buttons include a display button 420 and a save button 421. The display button 420 is a UI element that instructs the display of an event in the first display area 41, third display area 43, fourth display area 44, fifth display area 45, and sixth display area 46 to switch to the display of the event corresponding to the display button. The save button 421 is a UI element that instructs the display of an event corresponding to the save button 421 in the first display area 41, third display area 43, fourth display area 44, fifth display area 45, and sixth display area 46 to save it.
[0043] The third display area 43 is an area for displaying a map showing the occurrence position of an event in a dangerous situation. On the map, a marker indicating the occurrence point of the event is displayed. The position of the marker is determined based on, for example, vehicle position data input to the supplementary information input unit 12. Further, in the third display area 43, latitude and longitude indicating the position of the marker may be displayed together with the marker. In addition, when the user operates the display button 420 for another event in the second display area 42, the position of the marker displayed in the third display area 43 is switched to the position corresponding to that event.
[0044] The fourth display area 44 is an area for displaying the scene output by the detection unit 33, that is, the scene corresponding to the option matching the cause label. In this way, the video processing system 1 presents to the user the scene corresponding to the cause label, that is, the scene that requires confirmation by the user, so that the user can efficiently perform the video confirmation work. In the case where the detection unit 33 does not detect a scene corresponding to the option matching the cause label, nothing is displayed in the fourth display area 44, or a message such as "No corresponding scene was detected" is displayed. In addition, when the user operates the display button 420 for another event in the second display area 42, the scene displayed in the fourth display area 44 is switched to the scene corresponding to that event.
[0045] The fifth display area 45 is an area that displays the options included in the command input to the detection unit 33. The display mode of the options displayed in the fifth display area 45 is determined based on the display in the fourth display area 44. For example, if a scene corresponding to an option that corresponds to a cause label is displayed in the fourth display area 44, the fifth display area 45 will display only the option that corresponds to that scene, or the option will be displayed in a way that makes it stand out compared to other options. If there are multiple options that correspond to one scene, the fifth display area 45 may display multiple options that correspond to that scene, or multiple options that correspond to that scene may be displayed in a way that makes them stand out compared to other options. Furthermore, if a scene corresponding to an option that corresponds to a cause label is not displayed in the fourth display area 44, only the options for when the cause label is not detected in the video, or the options for when the target person, the driver, is not detected, will be displayed, or one of these options will be displayed in a way that makes it stand out compared to other options.
[0046] Methods for making options stand out include increasing the font size, using accent colors, thickening the border, and adding animation effects. In the example shown in Figure 5, the scene corresponding to the cause label "using a smartphone" is displayed in the fourth display area 44. Then, in the fifth display area 45, the icon for "using a smartphone" is displayed in a different color from the icons of the other options.
[0047] Furthermore, the display manner of the options corresponding to the scene displayed in the fourth display area 44 may be determined based on the risk level of the cause label, that is, the risk level of the cause of the occurrence. For example, the color of the icon and the size of the font of the options corresponding to the cause label may be determined based on the risk level of the cause label. This makes it easier for the user to intuitively identify the options corresponding to cause labels with a high risk level.
[0048] When the user operates the display button 420 of another event in the second display area 42, the scene displayed in the fourth display area 44 is switched to the scene corresponding to said event. Then, in accordance with the switching of the scene displayed in the fourth display area 44, the display mode of the options displayed in the fifth display area 45 is switched to a display mode corresponding to the display of the fourth display area 44.
[0049] The sixth display area 46 is an area for displaying the explanatory text output by the detection unit 33, that is, explanatory text containing a description relating to the scene displayed in the fourth display area 46. In the example shown in Fig. 5, a scene corresponding to the cause label "using a smartphone" is displayed in the fourth display area 44. Then, in the sixth display area 46, the explanatory text "The driver is using a smartphone" is displayed. In addition to the content of the cause label, the explanatory text output by the detection unit 33 may also include detailed content of the scene, such as the surrounding situation or temporal changes in behavior, for example. For example, in the sixth display area 46, the explanatory text "The driver takes a smartphone out of the bag on the adjacent seat and is using the smartphone while driving." may be displayed. As described above, the video processing system 1 presents to the user the explanatory text of the scene corresponding to the cause label, that is, the scene that requires confirmation by the user, thereby making it easy for the user to understand the content of the detection result. Therefore, the user can efficiently perform the video confirmation work.
[0050] In addition, when a scene corresponding to the option for the cause label is not displayed in the fourth display area 44, the sixth display area 46 displays explanatory text explaining that the cause label was not detected in the video, or explanatory text explaining that the driver who is the target person was not detected. Further, similarly to the fifth display area 45, in accordance with the switching of the scene displayed in the fourth display area 44, the explanatory text displayed in the sixth display area 46 is also switched to explanatory text corresponding to the display of the fourth display area 44.
[0051] <Video Processing Method> Next, an example of the video processing method will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the video processing method according to the present disclosure.
[0052] First, the video processing system 1 acquires video (step S101). More specifically, the display control unit 34 displays a video input UI unit 10 on the display means to receive video input from the user. The video input UI unit 10 also accepts video supplementary information to supplement the video, such as car position data and car acceleration data, along with the video. The video and acceleration data received by the video input UI unit 10 are transmitted to the extraction unit 31.
[0053] Next, the video processing system 1 extracts video containing dangerous events (step S102). More specifically, first, the extraction unit 31 detects the occurrence of dangerous events in the video input to the video input UI unit 10 based on acceleration data input to the video input UI unit 10. Then, the extraction unit 31 extracts video for a predetermined period based on the time of occurrence of the event. The extraction unit 31 may also detect the end time of the event based on supplementary information such as the acceleration data of the vehicle and extract video up to that end time. The extraction unit 31 inputs the extracted video to the detection unit 33. If the extraction unit 31 detects multiple dangerous events in the video input to the video input UI unit 10, it extracts video for each detected event. Then, the extraction unit 31 inputs the multiple extracted videos to the detection unit 33.
[0054] Next, the video processing system 1 acquires one or more cause labels (step S103). More specifically, first, the display control unit 34 displays a cause label input UI unit 20 on the display means to accept the user's selection of cause labels. The one or more cause labels selected in the cause label input UI unit 20 are transmitted to the command statement generation unit 32.
[0055] Next, the video processing system 1 generates an instruction statement (step S104) that causes the detection unit 33 to detect a scene corresponding to a cause label, based on the one or more cause labels acquired in step S103. More specifically, the instruction statement generation unit 32 generates an instruction statement that includes multiple options. The multiple options include at least an option corresponding to a cause label and an option for when no cause label is detected in the video. The instruction statement generated by the instruction statement generation unit 32 may be displayed on a display means such as a monitor for presentation to the user. The user may then modify the instruction statement by adding or changing options. The instruction statement generation unit 32 inputs the generated instruction statement to the detection unit 33.
[0056] Next, the video processing system 1 detects scenes corresponding to the cause label in the video extracted in step S102 based on the command statement generated in step S104 (step S105). More specifically, the detection unit 33 executes a process to detect scenes corresponding to the cause label in the video input from the extraction unit 31 based on the command statement input from the command statement generation unit 32. The detection unit 33 then outputs the detection result and an explanatory text containing an explanation of the detection result. The detection unit 33 transmits the outputted detection result and explanatory text to the detection result display UI unit 40.
[0057] Furthermore, if the detection unit 33 detects a scene that corresponds to an option corresponding to the cause label, it outputs the detected scene and the option corresponding to that scene, i.e., the cause label, as detection results. On the other hand, if it does not detect a scene that corresponds to an option corresponding to the cause label, the detection unit 33 outputs the options that would be available if the cause label were not detected in the video, or the options that would be available if the driver were not detected, as detection results.
[0058] Furthermore, when the detection unit 33 receives multiple video feeds from the extraction unit 31, that is, when multiple dangerous events are detected in S102, it performs detection processing on each video feed received from the extraction unit 31. The detection unit 33 then outputs a detection result and an explanatory text for each video feed. The detection unit 33 transmits the outputted multiple detection results and explanatory texts to the detection result display UI unit 40.
[0059] Finally, the video processing system 1 displays the detection result output in step S105 and the explanatory text on the display means (step S106). More specifically, the display control unit 34 displays the detection result display UI unit 40 on the display means as shown in Figure 5 in order to present the detection result and explanatory text to the user.
[0060] As described above, the video processing system 1 acquires video footage and one or more cause labels containing information indicating the cause of the dangerous event. The video processing system 1 then detects scenes in the video that correspond to the cause labels. When a scene corresponding to a cause label is detected, the video processing system 1 presents the detected scene and the corresponding cause label to the user as detection results. In this way, the video processing system 1 presents the user with scenes corresponding to cause labels, i.e., scenes that require user confirmation, allowing the user to efficiently review the video. Furthermore, the video processing system 1 presents the user with an explanation of the scenes that require user confirmation. This makes it easier for the user to understand the content of the detection results. Therefore, the user can review the video more efficiently.
[0061] (Embodiment 3) <Configuration of the video processing system> Another example of the video processing system will be described below with reference to Figure 7. Figure 7 is a diagram showing another example of the configuration of the video processing system according to the present disclosure. As shown in Figure 7, the video processing system 1 according to Embodiment 3 is a processing device comprising a video input UI unit 10, a cause label input UI unit 20, a video processing unit 30, and a detection result display UI unit 40. The video processing system 1 according to Embodiment 3 differs from the video processing system 1 according to Embodiment 2 in that the video processing unit 30 includes a model update unit 35 and the configuration of the detection result display UI unit 40. The other configurations are the same as those of the video processing system 1 described in Embodiment 2, so redundant explanations will be omitted.
[0062] The detection unit 33, based on the command input from the command generation unit 32, executes a process to detect scenes corresponding to the cause label in the video input from the extraction unit 31. The detection unit 33 then outputs the detection result and an explanatory text containing an explanation of the detection result. In this embodiment 3, the detection unit 33 uses a pre-trained model that has been trained to take video and a command input as input and output a detection result and an explanatory text containing an explanation of the detection result.
[0063] When the model update unit 35 receives a changed cause label or a changed explanation from the detection result display UI unit 40 (described later), it updates the trained model stored in the detection unit 33 based on the said cause label or explanation. More specifically, it retrains the already trained model using a training dataset that includes the changed content. The model update unit 35 then stores the retrained trained model in the detection unit 33. By using the trained model updated based on the changed cause label or explanation, the detection unit 33 can output detection results and explanations with greater accuracy.
[0064] Next, the UI unit 40 for displaying detection results in this third embodiment will be described. Figure 8 is a diagram showing another example of the configuration of a user interface for displaying detection results and explanatory text according to this disclosure. As shown in Figure 8, the UI unit 40 for displaying detection results includes a first display area 41, a second display area 42, a third display area 43, a fourth display area 44, a fifth display area 45, a sixth display area 46, and a change input unit 47.
[0065] In this embodiment 3, the fifth display area 45 is configured not only to display the options included in the command input to the detection unit 33, but also to allow the user to edit the options as needed. For example, the fifth display area 45 contains multiple text box UI elements, each displaying one option. The user can directly edit the options in each text box UI element.
[0066] In this embodiment 3, the sixth display area 46 is configured not only to display the explanatory text output by the detection unit 33, but also to allow the user to edit the explanatory text as needed. For example, the sixth display area 46 displays the explanatory text output by the detection unit 33 in the UI element of a text box. The user can then edit the explanatory text by directly modifying or adding to it in the UI element of the text box.
[0067] The change input unit 47 is a button UI element that instructs the model update unit 35 to input the cause label or explanation text that has been changed by user editing. For example, after the user has edited a selection, i.e., the cause label, in the fifth display area 45, or edited the explanation text in the sixth display area 46, the user can operate the change input unit 47 to send the changed cause label or changed explanation text to the model update unit 35.
[0068] <Image Processing Method> Next, another example of the image processing method will be explained using Figure 9. Figure 9 is a flowchart of another example of the image processing method according to the present disclosure. As shown in Figure 9, the image processing method according to Embodiment 3 differs from the image processing method according to Embodiment 2 in that it further includes steps S107, S108, and S109. The other steps are the same as those described in Embodiment 2, so a redundant explanation will be omitted.
[0069] The video processing system 1 displays the detection result output in step S105 and the explanatory text on the display means (step S106). More specifically, the display control unit 34 displays a detection result display UI unit 40 on the display means as shown in Figure 8 in order to present the detection result and explanatory text to the user.
[0070] Then, if the video processing system 1 receives input for a change in the cause label or a change in the explanatory text (step S107: YES), it updates the trained model (step S108). More specifically, when the user edits a selection, i.e., the cause label, in the fifth display area 45, or edits the explanatory text in the sixth display area 46, and operates the change input unit 47, the detection result display UI unit 40 transmits the changed cause label or changed explanatory text to the model update unit 35. The model update unit 35 then updates the trained model stored in the detection unit 33 based on the changed cause label or changed explanatory text.
[0071] If the system does not accept input for a change in the cause label or explanation (step S107: NO), the video processing system 1 terminates the flowchart shown in Figure 9.
[0072] After step S108, if the video processing system 1 determines that detection needs to be performed again (step S109: YES), it repeats step S105. The determination of whether detection needs to be performed again is made, for example, based on user instructions or the settings of the detection unit 33. For example, if the cause label is changed, if the changed description contains modifications exceeding a certain threshold, or if new information is added to the changed description, the detection unit 33 determines that detection needs to be performed again. In this case, the detection unit 33 repeats the detection process using the updated trained model and outputs the detection result and description.
[0073] If further detection is not required (step S109: NO), the video processing system 1 terminates the flowchart shown in Figure 9.
[0074] As described above, the video processing system 1 is configured so that the user can edit the options included in the command statement, namely the cause label and the explanatory text containing an explanation of the detection result. When the cause label or explanatory text is changed by the user's editing, the video processing system 1 updates the trained model based on the changed cause label or explanatory text. By using the updated trained model in the detection process, the video processing system 1 can output detection results and explanatory text with greater accuracy.
[0075] <Hardware Configuration for Realizing the Video Processing Functions of the Video Processing System> Some or all of the video processing realized by the video processing system 1 can be realized by a general-purpose computer system. This will be briefly explained below using Figure 10.
[0076] Figure 10 is a block diagram showing an example of a hardware configuration for realizing video processing in the video processing system according to this disclosure. The computer 60 includes, for example, a CPU (Central Processing Unit) 61 which is a control device, RAM (Random Access Memory) 62 and ROM (Read Only Memory) 63. The computer 60 further includes an IF (Interface) 64 which is an interface to the outside and an HDD (Hard Disk Drive) 65 which is an example of a non-volatile storage device. Furthermore, the computer 60 may also include other configurations not shown, such as input devices like a keyboard or mouse and display devices like a display.
[0077] The HDD 65 stores an OS (Operating System) and a control program 66, which are not shown in the diagram. The control program 66 is a computer program (image processing program) that implements the image processing of the image processing system 1.
[0078] The CPU 61 controls various processes in the computer 60, including access to RAM 62, ROM 63, IF 64, and HDD 65. The computer 60 reads and executes the OS and control program 66 stored in HDD 65 by the CPU 61. As a result, the computer 60 realizes the video processing of the video processing system 1.
[0079] The program described above includes a set of instructions (or software code) that, when loaded into a computer, causes the computer to perform one or more of the functions described in this disclosure. The program may be stored in a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include RAM, ROM, flash memory, SSD (Solid-State Drive), or other memory technologies, CD-ROM, DVD (Digital Versatile Disc), Blu-ray® disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include, a temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically, or otherwise propagating signals.
[0080] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0081] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.
[0082] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0083] (Note 1) A video processing system comprising: an instruction generation unit that generates an instruction statement for detecting a scene corresponding to a cause label based on one or more cause labels that include information indicating the cause of an event in a dangerous situation; and a detection unit that performs scene detection on a video based on the instruction statement and outputs the detection result.
[0084] (Note 2) The video processing system according to Note 1, wherein the detection result includes the detected scene and the cause label corresponding to the scene.
[0085] (Note 3) The video processing system according to Note 1 or 2, wherein the detection unit further outputs an explanatory text that includes an explanation of the detection result.
[0086] (Appendix 4) The video processing system according to Appendix 3, further comprising a display control unit that displays the detection result and the explanatory text on a display means.
[0087] (Note 5) The image processing system according to Note 4, wherein the display control unit determines the display mode of the cause label displayed on the display means based on the degree of risk of the cause of occurrence.
[0088] (Note 6) The video processing system according to Note 4 or 5, wherein the detection unit takes the video and the command statement as input and outputs the detection result and the explanatory text using a trained model that has been trained to output the detection result and the explanatory text.
[0089] (Note 7) The video processing system according to Note 6, further comprising a model update unit for updating the learned model, wherein the display control unit causes a change input unit for receiving input of a change in the cause label or a change in the explanatory text to be displayed on the display means, and the model update unit updates the learned model based on the change in the cause label or the change in the explanatory text.
[0090] (Appendix 8) A trained model that takes video footage and one or more cause labels containing information indicating the cause of an event in a dangerous situation as input, and is trained to detect scenes corresponding to the cause labels in the video footage, wherein the trained model takes video footage and one or more of the cause labels as input, and outputs a detection result and an explanatory text containing an explanation of the detection result.
[0091] (Note 9) A trained model that takes video footage and one or more command statements generated based on cause labels that include information indicating the cause of an event in a dangerous situation as input, and has been trained to detect scenes corresponding to the cause labels in the video footage based on the command statements, and takes the video footage and the command statements as input, and outputs a detection result and an explanatory text that includes an explanation of the detection result as content.
[0092] (Note 10) A video processing method comprising: acquiring video; acquiring one or more cause labels including information indicating the cause of an event in a dangerous situation; generating a command statement for detecting a scene corresponding to the cause label based on the cause label; performing the detection of the scene on the video based on the command statement; and outputting the detection result obtained from the detection.
[0093] (Note 11) A video processing program that causes a computer to perform the following: a process of acquiring video; a process of acquiring one or more cause labels including information indicating the cause of an event in a dangerous situation; a process of generating a command statement to detect a scene corresponding to the cause label based on the cause label; a process of performing the detection of the scene on the video based on the command statement; and a process of outputting the detection result obtained from the detection.
[0094] Some or all of the elements described in Appendices 2 to 7, which are subordinate to Appendice 1, may also be subordinate to Appendices 10 and 11 in the same manner as those described in Appendices 2 to 7. Some or all of the elements described in any appendice may be applied to various hardware, software, recording means, systems, and methods for recording software.
[0095] This application claims priority based on Japanese Patent Application No. 2025-048341, filed on 24 March 2025, and incorporates all of its disclosures herein.
[0096] 1 Video Processing System 10 UI for Video Input 11 Video Input Unit 12 Supplementary Information Input Unit 20 UI for Cause Label Input 21 Cause Label Addition Unit 22 Cause Label Selection Unit 30 Video Processing Unit 31 Extraction Unit 32 Command Generation Unit 33 Detection Unit 34 Display Control Unit 35 Model Update Unit 40 UI for Displaying Detection Results 41 First Display Area 42 Second Display Area 43 Third Display Area 44 Fourth Display Area 45 Fifth Display Area 46 Sixth Display Area 47 Change Input Unit 60 Computer 61 CPU 62 RAM 63 ROM 64 IF 65 HDD 66 Control Program 420 Display Button 421 Save Button
Claims
1. A video processing system comprising: an instruction generation unit that generates an instruction statement for detecting a scene corresponding to a cause label based on one or more cause labels that include information indicating the cause of an event in a dangerous situation; and a detection unit that performs scene detection on a video based on the instruction statement and outputs the detection result.
2. The video processing system according to claim 1, wherein the detection result includes the detected scene and the cause label corresponding to the scene.
3. The video processing system according to claim 2, wherein the detection unit further outputs an explanatory text that includes an explanation of the detection result.
4. The video processing system according to claim 3, further comprising a display control unit that displays the detection result and the explanatory text on a display means.
5. The image processing system according to claim 4, wherein the display control unit determines the display mode of the cause label displayed on the display means based on the degree of risk of the cause of occurrence.
6. The video processing system according to claim 4 or 5, wherein the detection unit takes the video and the command statement as input and outputs the detection result and the explanatory text using a trained model that has been trained to output the detection result and the explanatory text.
7. The video processing system according to claim 6, further comprising a model update unit for updating the learned model, wherein the display control unit causes a change input unit for receiving input of a change in the cause label or a change in the explanatory text to be displayed on the display means, and the model update unit updates the learned model based on the change in the cause label or the change in the explanatory text.
8. A trained model that takes video footage and one or more cause labels containing information indicating the cause of an event in a dangerous situation as input, and is trained to detect scenes corresponding to the cause labels in the video footage, wherein the trained model takes video footage and one or more of the cause labels as input, and outputs a detection result and an explanatory text containing an explanation of the detection result.
9. A trained model that takes video footage and a command statement generated based on one or more cause labels containing information indicating the cause of an event in a dangerous situation as input, and has been trained to machine-learn to detect scenes corresponding to the cause labels in the video footage based on the command statement, wherein the trained model takes the video footage and the command statement as input and outputs a detection result and an explanatory statement containing an explanation of the detection result.
10. A video processing method comprising: acquiring video; acquiring one or more cause labels including information indicating the cause of an event in a dangerous situation; generating a command statement for detecting a scene corresponding to the cause label based on the cause label; performing scene detection on the video based on the command statement; and outputting the detection result obtained from the detection.
11. A video processing program that causes a computer to perform the following steps: acquire video; acquire one or more cause labels including information indicating the cause of an event in a dangerous situation; generate a command statement to detect a scene corresponding to the cause label based on the cause label; perform the detection of the scene on the video based on the command statement; and output the detection result obtained from the detection.