Video analysis device, video analysis method, and video analysis program
Patent Information
- Application Number
- PCT/JP2025/012426
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012426_01102026_PF_FP_ABST
Abstract
Description
Video analysis apparatus, video analysis method, and video analysis program
[0001] The present invention relates to a video analysis apparatus, a video analysis method, and a video analysis program.
[0002] Diminished Reality (DR) is a technology that reduces the visibility of objects (e.g., regions, objects) in a video. For example, this technology identifies, among various objects that may attract a user's attention, objects that are irrelevant to the user's task. This technology performs processing to reduce visibility of the identified object (hereinafter referred to as "visibility reduction processing"). This execution can help the user focus on the task.
[0003] The conventional technology performs visibility reduction processing on an object that is consciously selected by the user (for example, see Non-Patent Document 1).
[0004] Hagiyama Naoki et al., "Study on Inattention Suppression System at Work Using Diminished Reality", IEICE Technical Report, vol.123, no.228, MVE2023-22, pp.1-5, 2023.
[0005] However, the conventional technology cannot perform visibility reduction processing on an object that is not consciously selected by the user. As a result, the conventional technology cannot perform visibility reduction processing on an object that the user unconsciously pays attention to, among various objects that may attract the user's attention.
[0006] An object of the present invention is to support execution of visibility reduction processing on an object that the user unconsciously pays attention to.
[0007] A video analysis apparatus according to an embodiment includes an acquisition unit, a region extraction unit, and an object selection unit. The acquisition unit acquires a field-of-view video related to a user's field of view and the user's gaze point. The region extraction unit extracts a gaze region including the gaze point from the field-of-view video. The object selection unit selects the gaze region as an object to be processed for reducing visibility.
[0008] According to the present invention, it is possible to support execution of visibility reduction processing on an object that the user unconsciously pays attention to.
[0009] Figure 1 is a functional configuration diagram of the video analysis device according to the first embodiment. Figure 2 is a flowchart of the video analysis method according to the first embodiment. Figure 3 is a diagram showing the field of view image according to the first embodiment. Figure 4 is a diagram showing the composite image according to the first embodiment. Figure 5 is a functional configuration diagram of the video analysis device according to the second embodiment. Figure 6 is a flowchart of the video analysis method according to the second embodiment. Figure 7 is a functional configuration diagram of the video analysis device according to the third embodiment. Figure 8 is a flowchart of the video analysis method according to the third embodiment. Figure 9 is a diagram showing the field of view image according to the third embodiment. Figure 10 is a hardware configuration diagram of the video analysis device according to each embodiment.
[0010] The embodiments will be described below with reference to the drawings. Multiple parts assigned the same reference numeral will be considered identical, and redundant explanations will be omitted as appropriate.
[0011] (First Embodiment) Figure 1 is a functional configuration diagram of the video analysis device 1 according to the first embodiment. The video analysis device 1 is a device that analyzes video. For example, the video analysis device 1 is a head-mounted display (HMD). The video analysis device 1 comprises various functional units (shooting unit 10A, gaze point measurement unit 10B, acquisition unit 11, region extraction unit 12, video processing unit 13, synthesis unit 14, display unit 15).
[0012] The shooting unit 10A is a means for performing various types of shooting. For example, the shooting unit 10A is a camera capable of shooting video. The shooting unit 10A shoots video related to the user's field of view (hereinafter referred to as "field of view video FV"). The shooting unit 10A transmits the field of view video FV to the acquisition unit 11.
[0013] The gaze point measurement unit 10B is a means for measuring various gaze points. For example, the gaze point measurement unit 10B is an eye tracker. The gaze point measurement unit 10B measures the user's gaze point GP. The gaze point measurement unit 10B transmits the gaze point GP to the acquisition unit 11.
[0014] The HMD may be a standalone type (or a network type). A standalone HMD may have all of the functional components. A network type HMD may have some of the functional components. Another information processing device (e.g., a smartphone, a computer) may have the remaining functional components. The network type HMD and the other information processing device may constitute an information processing system.
[0015] The camera (or eye tracker) may be attached to the user's head along with the HMD. The camera (or eye tracker) may be positioned near the user or in a location that allows observation of the user's field of view.
[0016] The acquisition unit 11 is a means for performing various acquisitions. For example, the acquisition unit 11 acquires the field of view image FV from the imaging unit 10A and the gaze point GP from the gaze point measurement unit 10B. The acquisition unit 11 may also acquire the field of view image FV and gaze point GP from a storage device (not shown). The acquisition unit 11 transmits the field of view image FV and gaze point GP to the region extraction unit 12. The acquisition unit 11 transmits the field of view image FV to the synthesis unit 14.
[0017] The region extraction unit 12 is a means for extracting various regions. For example, the region extraction unit 12 extracts a video region (hereinafter referred to as "gaze region GR") that includes the gaze point GP from the field of view video FV. The region extraction unit 12 transmits the field of view video FV, the gaze point GP, and the gaze region GR to the video processing unit 13.
[0018] The video processing unit 13 is a means for performing various video processing. For example, the video processing unit 13 performs visibility reduction processing. The video processing unit 13 includes a target selection unit 131, a visual effect determination unit 132, and an editing data generation unit 133.
[0019] The target selection unit 131 is a means for selecting various targets. For example, the target selection unit 131 selects the gaze area GR as the target T for visibility reduction processing. The target selection unit 131 transmits the target T to the visual effect determination unit 132.
[0020] The visual effect determination unit 132 is a means for determining various visual effects. For example, the visual effect determination unit 132 determines the visual effect VE for the target T. The visual effect determination unit 132 transmits the target T and the visual effect VE to the editing data generation unit 133.
[0021] The editing data generation unit 133 is a means for generating various types of editing data. For example, the editing data generation unit 133 generates editing data ED for executing the visual effect VE on the target T. The editing data generation unit 133 transmits the editing data ED to the synthesis unit 14.
[0022] The compositing unit 14 is a means for performing various types of compositing. For example, the compositing unit 14 obtains a composite image SV by compositing the field of view image FV and the editing data ED. The compositing unit 14 transmits the composite image SV to the display unit 15.
[0023] The display unit 15 is a means for performing various displays. For example, the display unit 15 displays a composite image SV.
[0024] Figure 2 is a flowchart of the video analysis method according to the first embodiment. The video analysis device 1 executes the video analysis method related to steps S1A to S9A through each functional unit.
[0025] (Step S1A) First, the imaging unit 10A captures a field of view image FV related to the user's field of view. For example, the imaging unit 10A obtains a field of view image FV by capturing the user's field of view in real time. The field of view image FV includes multiple images that are continuous in time (see Figure 3).
[0026] (Step S2A) Next, the gaze point measurement unit 10B measures the user's gaze point GP. For example, the gaze point measurement unit 10B obtains the gaze point GP by measuring the user's line of sight in real time. The gaze point GP may be expressed as coordinates. The gaze point GP may be coordinates in the real world or coordinates in the field of view image FV. The coordinates in the real world and the coordinates in the field of view image FV may be associated with each other by a predetermined relational expression. Step S2A may be performed before step S1A.
[0027] (Step S3A) Next, the acquisition unit 11 acquires the field of view image FV and the point of gaze GP.
[0028] (Step S4A) Next, the region extraction unit 12 extracts a gaze region GR including the gaze point GP from the field of view image FV. For example, the region extraction unit 12 extracts the image region centered on the gaze point GP as the gaze region GR. The gaze region GR may have any shape (e.g., rectangle, circle). The gaze region GR may have any size (or number of pixels). The region extraction unit 12 may use existing object recognition technology to detect an object present at the gaze point GP. The region extraction unit 12 may extract the image region corresponding to the detected object as the gaze region GR.
[0029] (Step S5A) Next, the target selection unit 131 selects the gaze area GR as the target T for the visibility reduction processing. For example, the target selection unit 131 assigns a flag (e.g., the numerical value "1") to the gaze area GR to indicate that the visibility reduction processing will be performed. This assignment selects the gaze area GR as the target T.
[0030] (Step S6A) Next, the visual effect determination unit 132 determines the visual effect VE for the target T. For example, the visual effect determination unit 132 determines one visual effect VE from a plurality of visual effects. The visual effect determination unit 132 may determine the visual effect VE based on the characteristics of the target T (e.g., transparency, contour lines, resolution, hue, saturation, brightness). The visual effect VE may be determined from the following six types of identification numbers: (1) transparency, (2) contouring, (3) blurring, (4) desaturation, (5) desaturation, (6) decontrast (see, for example, Non-Patent Document 1).
[0031] (1) Transparency increases the transparency (or decreases the opacity) of the object T. (2) Contouring represents the object T with contours (or edges). (3) Blurring reduces the resolution of the object T. (4) De-saturation reduces the hue of the object T. (5) De-saturation reduces the saturation of the object T. (6) De-contrast reduces the brightness of the object T.
[0032] (Step S7A) Next, the editing data generation unit 133 generates editing data ED for executing the visual effect VE on the target T. For example, the editing data generation unit 133 obtains editing data ED by combining data that identifies the target T (e.g., shape and size of the region) and data that identifies the visual effect VE (e.g., identification number of the visual effect).
[0033] (Step S8A) Next, the compositing unit 14 combines the field of view image FV and the editing data ED. For example, the compositing unit 14 identifies a target T in the field of view image FV based on the editing data ED. The compositing unit 14 obtains a composite image SV by applying the visual effect VE to the identified target T.
[0034] (Step S9A) Finally, the display unit 15 displays the composite image SV. The display unit 15 may display information that attracts the user's attention (also called "attention-attracting information") before displaying the composite image SV (see Figure 4).
[0035] Figure 3 shows a field of view image FV according to the first embodiment. The field of view image FV includes three objects J (computer J1, desk J2, and beverage can J3). The user engages in a task using a cursor C displayed on the screen of computer J1. While engaging in the task, the user may unconsciously pay attention to the beverage can J3 (i.e., the point of focus GP is attracted to the beverage can J3). As a result, the user cannot concentrate on the task using computer J1.
[0036] Figure 4 shows a composite image SV according to the first embodiment. In the composite image SV, a visibility reduction process is performed on the gaze area GR, which includes the user's gaze point GP. This process reduces the visibility of the beverage can J3. As a result, the visibility of the computer J1 is relatively improved, allowing the user to concentrate on tasks using the computer J1. The gaze point GP (or gaze area GR) may or may not be displayed in the field of view image FV (or composite image SV).
[0037] According to the first embodiment described above, the video analysis device 1 extracts a gaze region GR, which includes the user's point of gaze GP, from the field of view video FV relating to the user's field of view. The video analysis device 1 selects the gaze region GR as the target T for visibility reduction processing.
[0038] Generally, a user's gaze point GP is selectively directed towards external stimuli that the user pays attention to, regardless of whether the user is conscious of it or not. Based on this property, the video analysis device 1 extracts the gaze region GR, which includes the user's gaze point GP, and selects the gaze region GR as the target T for visibility reduction processing. As a result, the video analysis device 1 can select, as the target T for visibility reduction processing, objects that the user unconsciously pays attention to, from among various objects that can attract the user's attention. In other words, the video analysis device 1 can assist in performing visibility reduction processing on objects that the user unconsciously pays attention to.
[0039] (Second Embodiment) Figure 5 is a functional configuration diagram of the video analysis device 1 according to the second embodiment. The video analysis device 1 according to the second embodiment has the same functional configuration as the first embodiment. The video analysis device 1 includes a gaze point input unit 10C instead of a gaze point measurement unit 10B. The video analysis device 1 may also include a gaze point measurement unit 10B and a gaze point input unit 10C.
[0040] The gaze point input unit 10C is a means for inputting various gaze points. For example, the gaze point input unit 10C may be a mouse, keyboard, etc. The gaze point input unit 10C may also receive a gaze point GP specified by the user. The gaze point input unit 10C transmits the gaze point GP to the acquisition unit 11.
[0041] Figure 6 is a flowchart of the video analysis method according to the second embodiment. The video analysis device 1 executes the video analysis method related to steps S1B to S9B through each functional unit. Steps S1B to S9B, excluding step S2B, are the same as steps S1A to S9A, excluding step S2A (see Figure 2).
[0042] (Step S2B) Here, the gaze point input unit 10C inputs the user's gaze point GP. For example, the gaze point input unit 10C inputs the gaze point GP specified by the user. The gaze point GP may be any point on the screen in the field-of-view image FV (e.g., the center point).
[0043] According to the second embodiment described above, the video analysis apparatus 1 does not need to measure the user's gaze point GP using an eye tracker or the like. As a result, the video analysis apparatus 1 can simply acquire the user's gaze point GP.
[0044] (Third Embodiment) FIG. 7 is a functional configuration diagram of the video analysis apparatus 1 according to the third embodiment. The video analysis apparatus 1 according to the third embodiment has the same functional configuration as that of the first embodiment. The video analysis apparatus 1 further includes an exclusion region input unit 10D.
[0045] The exclusion region input unit 10D is a means for inputting various exclusion regions. For example, the exclusion region input unit 10D is a mouse, a keyboard, or the like. The exclusion region input unit 10D may input the exclusion region ER specified by the user. The exclusion region input unit 10D transmits the exclusion region ER to the acquisition unit 11.
[0046] The acquisition unit 11 further transmits the exclusion region ER to the region extraction unit 12. The region extraction unit 12 further transmits the exclusion region ER to the video processing unit 13. The combining unit 14 transmits the field-of-view image FV (or the combined image SV) to the display unit 15. The display unit 15 displays the field-of-view image FV (or the combined image SV).
[0047] FIG. 8 is a flowchart of the video analysis method according to the third embodiment. The video analysis apparatus 1 executes the video analysis method related to steps S1C to S11C, S7D, and S11D through each functional unit.
[0048] (Steps S1C to S2C) Steps S1C to S2C are respectively the same as steps S1A to S2A (see FIG. 2).
[0049] (Step S3C) Next, the exclusion area input unit 10D inputs an exclusion area ER of the field of view image FV. For example, the exclusion area input unit 10D inputs an exclusion area ER specified by the user. The exclusion area ER may be any area on the screen in the field of view image FV (e.g., the central area). The exclusion area ER may be an area related to the user's task (also called a "task-related area"). The exclusion area ER may be an area identified by existing object recognition technology. Steps S1C, S2C, and S3C may be executed in any order.
[0050] (Step S4C) Next, the acquisition unit 11 acquires the field of view image FV, the point of gaze GP, and the excluded area ER.
[0051] (Step S5C) Step S5C is the same as step S4A (see Figure 2).
[0052] (Step S6C) Next, the target selection unit 131 determines whether or not the gaze area GR is located in the exclusion area ER. For example, the target selection unit 131 determines whether or not the gaze area GR is located in the exclusion area ER of the field of view image FV. If the gaze area GR is not located in the exclusion area ER (NO), the process proceeds to step S7C. If the gaze area GR is located in the exclusion area ER (YES), the process proceeds to step S7D.
[0053] (Step S7C) Next, the target selection unit 131 selects the gaze area GR as the target T for visibility reduction processing. The target selection unit 131 may also select a portion of the gaze area GR that is not included in the exclusion area ER as the target T. Step S7C is the same as step S5A (see Figure 2).
[0054] (Step S7D) Alternatively, the target selection unit 131 does not select the gaze area GR as the target T for the visibility reduction processing. For example, the target selection unit 131 assigns a flag (e.g., the numerical value "0") to the gaze area GR to indicate that the visibility reduction processing will not be performed. This assignment prevents the gaze area GR from being selected as the target T. The target selection unit 131 does not need to select a portion of the gaze area GR that is included in the excluded area ER as the target T. After step S7D, the process proceeds to step S11D.
[0055] (Steps S8C to S11C) Steps S8C to S11C are the same as steps S6A to S9A, respectively (see Figure 2).
[0056] (Step S11D) Finally, the display unit 15 displays the field of view image FV (see Figure 9).
[0057] Figure 9 shows a field of view image FV according to the third embodiment. The field of view image FV includes an exclusion area ER in the central region of the screen. The point of focus GP is located at cursor C on the screen of computer J1. The gaze area GR includes cursor C.
[0058] The target selection unit 131 determines that the gaze area GR is located in the exclusion area ER (YES). The target selection unit 131 does not select the gaze area GR as the target T for the visibility reduction processing. In the field of view image FV, the visibility reduction processing is not performed on the gaze area GR, which includes the user's point of focus GP. As a result, the user can concentrate on tasks using the computer J1. The exclusion area ER may or may not be displayed in the field of view image FV (or composite image SV).
[0059] According to the third embodiment described above, the video analysis device 1 extracts the gaze region GR, which includes the user's point of gaze GP, from the field of view video FV relating to the user's field of view. If the gaze region GR is located in the exclusion region ER, the video analysis device 1 does not select the gaze region GR as the target T for visibility reduction processing.
[0060] For example, the excluded area ER is specified for an area related to the user's task. In this case, the video analysis device 1 does not perform visibility reduction processing on that area, thus helping the user concentrate on the task.
[0061] Figure 10 is a hardware configuration diagram of the video analysis device 1 according to each embodiment. The video analysis device 1 includes a CPU 101, RAM 102, ROM 103, storage 104, display device 105, input device 106, and communication device 107 as its components. Each component is connected to each other via an internal bus so as to be able to communicate with each other. The video analysis device 1 may include at least one of each component.
[0062] The CPU 101 is a processor that executes various processes according to a program. The CPU 101 uses a predetermined area of the RAM 102 as a working area. The CPU 101 realizes each processing unit (e.g., acquisition unit 11, area extraction unit 12, video processing unit 13, target selection unit 131, visual effect determination unit 132, editing data generation unit 133, synthesis unit 14) by reading and executing each program stored in the ROM 103 or storage 104. Each processing unit may be realized by a dedicated hardware circuit (e.g., ASIC). The CPU 101 is an example of a processing unit.
[0063] RAM 102 is a memory that stores various types of data in a rewritable format. RAM 102 is an SDRAM (Synchronous Dynamic Random Access Memory), etc. ROM 103 is a memory that stores various types of data in a non-rewritable format. Storage 104 is various types of storage media. Storage 104 may also be a drive device that writes or reads various types of data to or from the storage media. Storage 104 may write or read various types of data to or from the storage media in accordance with the control of the CPU 101. RAM 102, ROM 103, or storage 104 are examples of storage units.
[0064] The display device 105 is a device that displays various images (or videos). The display device 105 is an LCD (Liquid Crystal Display), etc. The display device 105 displays various images based on display signals from the CPU 101. The display device 105 is an example of a display unit (e.g., display unit 15).
[0065] The input device 106 is a device that receives various input operations from the user. The input device 106 is a mouse, keyboard, etc. The input device 106 receives the operations entered by the user as instruction signals and transmits the received instruction signals to the CPU 101. The input device 106 is an example of an input unit (e.g., a gaze point input unit 10C, an exclusion area input unit 10D).
[0066] The communication device 107 is a device that communicates with external devices via a network in accordance with the control of the CPU 101. The communication device 107 is an example of a communication unit.
[0067] Each embodiment of the present invention is presented as an example and does not limit the scope of the invention. Each embodiment can be carried out in various forms without departing from the spirit of the invention. Each embodiment may be combined with one another, in which case the combined effects can be obtained. Each embodiment includes a plurality of components, and various inventions can be obtained by various combinations of these components. Each embodiment or each combination of components is included within the scope of the invention.
[0068] 1...Video analysis device, 10A...Shooting unit, 10B...Point of focus measurement unit, 10C...Point of focus input unit, 10D...Exclusion area input unit, 11...Acquisition unit, 12...Area extraction unit, 13...Video processing unit, 14...Composition unit, 15...Display unit, 101...CPU, 102...RAM, 103...ROM, 104...Storage, 105...Display device, 106...Input device, 107...Communication device, 131...Target selection unit, 132...Visual effect determination unit, 133...Editing data generation unit, C...Cursor, ED...Editing data, ER...Exclusion area, FV...Field of view image, GP...Point of focus, GR...Viewing area, J...Object, J1...Computer, J2...Desk, J3...Beverage can, SV...Composite image, T...Target, VE...Visual effect
Claims
1. An image analysis device comprising: an acquisition unit that acquires a field of view image relating to the user's field of view and the user's point of gaze; a region extraction unit that extracts a gaze region including the point of gaze from the field of view image; and a target selection unit that selects the gaze region as the target of processing to reduce visibility.
2. The video analysis apparatus according to claim 1, wherein the target selection unit does not select the gaze region as the target if the gaze region exists in the exclusion region of the field of view video.
3. A video analysis method comprising: a computer acquiring a field of view image related to the user's field of view and the user's point of gaze; extracting a gaze region including the point of gaze from the field of view image; and selecting the gaze region as the target of a process to reduce visibility.
4. A video analysis program that causes a computer to function as a component of the video analysis apparatus described in claim 1.