Processing device, processing device control method, and program
The processing device improves visibility of superimposed information on human body areas by adjusting transparency based on movement, addressing clarity issues in explanatory movements.
Patent Information
- Application Number
- JP2023073644
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-10-27
- Estimated Expiration
- 2043-04-27
AI Technical Summary
Existing technologies struggle to make superimposed information on human body areas in videos clearly visible, especially when the human body is making explanatory movements, leading to potential misunderstandings.
A processing device that extracts human body regions, determines if the body is making explanatory movements, and adjusts the transparency of overlapping information areas accordingly, ensuring clarity by making the overlapping regions transparent during movements.
Enhances visibility of superimposed information by making it transparent when the human body is explaining, allowing viewers to better understand the content.
Smart Images

Figure 0007760550000001 
Figure 0007760550000002 
Figure 0007760550000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a processing device, a control method for the processing device, and a program. [Background technology]
[0002] In the past, automatic filming of lecture scenes has increasingly involved overlaying information that the person is explaining onto the video of the person. In such cases, measures have been taken to ensure that the background of the overlaid video is not difficult to see.
[0003] In Patent Document 1, when a second image (CG person or sign language interpreter) is superimposed on a first image (background), the display position and transparency of the second image are controlled using image information extracted from the first image, making the background easier to see. The image information extracted from the first image is a saliency map created from program information, and areas that attract people's attention. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 6046961 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in Patent Document 1, when a second video is superimposed on a first video, the transparency of the area overlapping with the human body area cannot be changed depending on whether the human body is making an explanatory movement or not. Furthermore, when information (the second video referred to in Patent Document 1) is superimposed on a video of a human body (the first video referred to in Patent Document 1), it may be difficult to visually recognize the information in the superimposed video. This raises concerns that viewers may not be able to fully understand what the human body is explaining through the superimposed information.
[0006] Therefore, an object of the present invention is to provide a processing device that can make information more easily visible when superimposing information on a video image. [Means for solving the problem]
[0007] In order to achieve the above object, a processing device according to one aspect of the present invention includes a first extraction means for extracting a region of a human body in an image, a superimposition means for superimposing predetermined superimposition information on the image, and a processing means for superimposing predetermined superimposition information on the human body performing a predetermined movement. Or judgement do a first determination means; In the area of the predetermined superimposition information overlapping with the region of the human body It is an area overlap area of a second extraction means for extracting the When the first determination means determines that the human body is performing the predetermined movement, at least In the overlapping region Okeru A transparency change means for changing the transparency to be higher. and when the first determination means determines that the human body is not performing the predetermined motion, the transparency change means does not change the transparency in the overlapping region. It is characterized by: [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a processing device that can make information more easily visible when superimposing information on a video. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram illustrating the configuration of an automatic photography system according to a first embodiment. [Figure 2] 1 is a block diagram illustrating the configuration of an image superimposing device according to a first embodiment. [Figure 3] 1A to 1C are diagrams illustrating the results of human body skeleton estimation according to the first embodiment. [Figure 4] 10A to 10C are diagrams illustrating a method for determining explanatory movements of a human body using a skeleton estimation result of the human body according to the first embodiment. [Figure 5] 10A and 10B are diagrams illustrating how explanatory material is superimposed on a video of a human body when the human body is not making an explanatory movement according to the first embodiment. [Figure 6]10A and 10B are diagrams illustrating how explanatory material in which an area overlapping with a human body area is made transparent is superimposed on a video of the human body when the human body is making an explanation motion according to the first embodiment. [Figure 7] 10 is a diagram illustrating a state in which explanatory material is superimposed on a video of a human body when the human body is making an explanation motion, with the area overlapping with the arm area of the human body being explained made transparent in embodiment 1. FIG. [Figure 8] 10A and 10B are diagrams illustrating a state in which explanatory material is superimposed on a video of a human body when the human body is making an explanation motion, with the area overlapping with the face area of the human body made transparent, according to the first embodiment. [Figure 9] 10A and 10B are diagrams illustrating how explanatory material is superimposed on a video of a human body when the human body is not making an explanatory movement according to the first embodiment. [Figure 10] 4 is a flowchart showing the processing procedure of the automatic photography system according to the first embodiment. [Figure 11] FIG. 10 is a block diagram illustrating the configuration of an automatic photography system according to a second embodiment. [Figure 12] 10A and 10B are diagrams illustrating how explanatory material is superimposed on a video of a human body, with the area overlapping the human body made transparent, when the human body is speaking, according to the second embodiment. [Figure 13] 10 is a diagram showing a speech section of a human body and a transparent section of a material area according to the second embodiment. FIG. [Figure 14] 10 is a flowchart showing the processing procedure of the automatic photography system according to the second embodiment. [Figure 15] FIG. 10 is a block diagram illustrating the configuration of an automatic photography system according to a third embodiment. [Figure 16] FIG. 11 is an explanatory diagram for identifying an explanation area based on the content of an utterance and the content of an explanatory material according to the third embodiment. [Figure 17] 10 is a flowchart showing the processing procedure of the automatic photography system according to the third embodiment. [Figure 18] 10 is a flowchart showing the processing procedure of the automatic photography system according to the third embodiment. [Figure 19] FIG. 10 is a block diagram illustrating the configuration of an automatic photography system according to a fourth embodiment. [Figure 20] 13 is a diagram showing how an explanation region is identified from an explanation action of a human body according to the fourth embodiment. FIG. [Figure 21] This figure explains how explanatory material is superimposed on a human body image in which the area overlapping the human body area is made transparent when the human body does not overlap with the explanation area when the human body performs an explanation motion in embodiment 4. [Figure 22] A figure explaining how explanatory material that has been emphasized is superimposed on an explanatory area on a human body image captured in a case where the human body and the explanatory area overlap when the human body performs an explanatory motion in embodiment 4. [Figure 23] A figure explaining how explanatory material that has been emphasized is superimposed on an explanation area on a human body image captured in a case where the human body and the explanation area overlap when the human body gives a verbal explanation in embodiment 4. [Figure 24] 10 is a flowchart showing the processing procedure of the automatic photography system according to the fourth embodiment. [Figure 25] 10 is a flowchart showing the processing procedure of the automatic photography system according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] The following describes in detail embodiments of the present invention. The embodiments described below are merely examples for realizing the present invention, and should be appropriately modified or adjusted depending on the configuration of the device to which the present invention is applied and various conditions. The present invention is not limited to the following embodiments. In addition, parts having the same functions in all figures are designated by the same numerals, and repeated explanations thereof will be omitted.
[0011] <Embodiment 1> An example of the configuration of the image superimposing device A1003 according to the first embodiment will be described below with reference to Fig. 1. Fig. 1 is a block diagram showing the functional configuration of an automatic photography system A1000 including the image superimposing device A1003 according to the first embodiment. The image superimposing device A1003 functions as a processing device (information processing device) that performs various processes such as human body extraction processing, explanatory action determination processing, and image superimposition processing using various functional units, etc., which will be described later.
[0012] The automatic photography system A1000 detects a human body (person) from a captured video (video information) and determines the explanatory behavior of the detected human body. It then makes transparent (changes the transparency) the area of the acquired explanatory material (superimposed information superimposed on the video) that overlaps with the area of the human body performing the explanatory behavior. It is a processing system that then superimposes the explanatory material with the changed transparency on the video information (video of the human body) and displays the superimposed image on a monitor.
[0013] The automatic photography system A1000 is configured to include an image capture device A1001, a document capture device A1002, an image superimposition device A1003, and a monitor device A1013. The image superimposition device A1003 is communicably connected to the image capture device A1001, the document capture device A1002, and the monitor device A1013. The image superimposition device A1003 and the monitor device A1013 are also connected via a line such as a video interface.
[0014] The video acquisition device A1001 is a device that acquires images by capturing images of the surroundings of the video acquisition device A1001 and generates a captured video from the multiple captured images, and is composed of an imaging device such as a camera. The video acquisition device A1001 has an imaging unit (not shown), which is composed of a lens unit for focusing light, an imaging element that converts the focused light into an analog signal, and a signal processing unit. The imaging unit also functions as a capturing means that captures an image by capturing an image of a subject. The video acquisition device A1001 outputs video information generated from the multiple captured images to the image superimposition device A1003.
[0015] The material acquisition device A1002 is a device that acquires explanatory materials such as presentation materials created with Microsoft PowerPoint or Adobe PDF in the form of electronic data. The material acquisition device A1002 outputs the acquired explanatory materials to the image superimposition device A1003. The explanatory materials may be any images or information that can be superimposed on video. The explanatory materials may be, for example, text information. In other words, the explanatory materials may be predetermined superimposition information that can be superimposed on video. Here, the predetermined superimposition information may be images or text, or may be other symbols, icons, etc. The explanatory materials are used, for example, by a human body included in the video. The human body included in the video can explain the contents of the predetermined superimposition information while viewing the video on which the predetermined superimposition information is superimposed via a monitor or the like.
[0016] The image superimposing device A1003 detects a human body from the video input from the video capture device A1001 and determines whether the human body is making an explanation motion. If the human body is making an explanation motion, the image superimposing device A1003 makes the area of the explanatory material that overlaps with the human body transparent and superimposes it on the video information. The superimposed video is then output to the monitor device A1013.
[0017] The image superimposing device A1003 is configured to have, as functional units, a video acquisition unit A1004, a document acquisition unit A1005, a skeletal information estimation unit A1006, a human body movement determination unit A1007, a region segmentation processing unit A1008, and an overlap region extraction unit A1009. The image superimposing device A1003 is further configured to have, as functional units, a transparency change unit A1010, an image superimposition unit A1011, and a video output device A1012. Each of these functional units is realized by the CPU 11, described later, loading a program stored in the ROM 12 into the RAM 13 and executing it. The CPU 11 then stores the execution results of each process, described later, in the RAM 13 or a predetermined storage medium.
[0018] The video acquisition unit A1004 acquires video information. Specifically, the video acquisition unit A1004 acquires video information input from the video acquisition device A1001. However, this is not limiting, and the video acquisition unit A1004 may acquire video information from a device or a server other than the video acquisition device A1001. The video acquisition unit A1004 outputs the acquired video information to the skeletal information estimation unit A1006 and the region segmentation processing unit A1008.
[0019] The skeletal information estimation unit A1006 estimates skeletal information of a human body. Specifically, it detects a human body from the video information input from the video acquisition unit A1004 and estimates skeletal information, which is information about the human body's skeleton. The skeletal information estimation unit A1006 detects a human body from (based on) an image included in the video information and estimates skeletal information of the detected human body. When estimating the skeletal information of the human body, the skeletal information estimation unit A1006 extracts the coordinates of the human body from the video information and estimates the skeletal information of the human body using a skeletal estimation technique. The skeletal information estimation unit A1006 then outputs the video information and the estimated skeletal information of the human body to the human body movement determination unit A1007 as a skeletal estimation result. In this embodiment, the skeletal information estimation unit A1006 also functions as an estimation means when estimating the skeleton of the detected human body and outputting the skeletal estimation result.
[0020] In recent years, many deep learning-based skeletal estimation technologies have emerged, making it possible to estimate the human skeleton with high accuracy. Among these, some technologies are provided as open source software (OSS), such as OpenPose and DeepPose, making it easier to perform skeletal estimation. In the first embodiment, no particular skeletal estimation technology is used, but one of the above-described deep learning-based skeletal estimation technologies is used.
[0021] The human body movement determination unit A1007 determines whether the human body is performing a predetermined movement. Specifically, it determines whether the human body is performing an explanation movement as a predetermined movement using human body skeletal information, which is the estimation result acquired from the skeletal information estimation unit A1006. If the human body movement determination unit A1007 determines that the human body has performed an explanation movement, it outputs the determination result, video information, and skeletal estimation result to the overlap region extraction unit A1009. On the other hand, if it determines that the human body is not performing an explanation movement, it outputs the determination result and video information to the image superimposition unit A1011. The determination process performed by the human body movement determination unit A1007 will be described below with reference to FIGS. 3 and 4. In this embodiment, the human body movement determination unit A1007 also functions as a first determination means that determines whether the human body is performing a predetermined movement and outputs the determination result.
[0022] Figure 3 shows the skeletal information of the shoulders, arms, and neck required to determine the explanatory motion from the human body skeletal estimation results acquired from the human body motion determination unit A1007. D001 represents video information. P001 represents the human body. P002 represents the left hand, P003 the left elbow, and P004 the left shoulder. P005 represents the neck. P006 represents the right shoulder, P007 the right elbow, and P008 the right hand.
[0023] FIG. 4 is a diagram illustrating how an explanatory motion made by a human body's right arm is determined. D101 represents video information. P101 represents the human body. P106 represents the right shoulder, P107 represents the right elbow, and P108 represents the right hand. The angle between the right shoulder P106 and the right elbow P107 is P109. The angle between the right elbow P107 and the right hand P108 is P110.
[0024] The human body movement determination unit A1007 can determine that an explanation movement is being made when, for example, P108 and P109 are between 0° and 90°. Note that this is merely an example, and any method can be used as long as it is possible to determine an explanation movement using skeletal information. For example, it may be determined that an explanation movement is being made when either P108 or P109 is between 0° and 90°. It may also be determined that an explanation movement is being made when, for example, the body or neck is rotated a predetermined amount, when both hands are spread or brought to the chest, when a finger is raised, or the like.
[0025] The region segmentation processing unit A1008 extracts regions of human bodies and objects in an image. Specifically, it performs region segmentation processing using video information input from the video acquisition unit A1004 to obtain information such as the region and type of the human body or object. The region segmentation processing unit A1008 functions as a first extraction unit that extracts a human body region from an image included in the video information (based on the image). The region segmentation processing unit A1008 outputs the obtained information on the region and type of the human body or object as region information to the overlap region extraction unit A1009. Note that various techniques are known for the region segmentation processing performed by the region segmentation processing unit A1008, such as region split, super-parsing, and full CNN (Convolutional Neural Network) using Deep Learning. In the first embodiment, it is assumed that full CNN is used because it can perform region segmentation with high accuracy, but any technique may be used. Region split, super-parsing, full CNN, etc. are well-known technologies, so detailed description thereof will be omitted.
[0026] The overlapping area extraction unit A1009 extracts overlapping areas from the explanatory material. Specifically, it extracts overlapping areas from the explanatory material using the human body movement determination result and skeleton estimation result input from the human body movement determination unit A1007, the region information input from the region segmentation processing unit A1008, and the explanatory material input from the material acquisition unit A1005. Note that the overlapping area extraction unit A1009 extracts overlapping areas when the determination result input from the human body movement determination unit A1007 indicates that an explanatory movement is being performed. The overlapping area is an area in the explanatory material that overlaps with the region information of the human body performing the explanatory movement in the region information. The overlapping area extraction unit A1009 outputs the extracted overlapping area, the explanatory material, and video information to the transparency change unit A1010. In this embodiment, the overlap area extraction unit A1009 also functions as a second extraction means that extracts an area of explanatory material that overlaps with the human body area as an overlap area based on the judgment result of the human body movement, the human body area, and the explanatory material.
[0027] The overlapping region extraction unit A1009 may also combine region information including a human body region with the skeleton estimation result to extract a partial region of the human body, such as the face or arms, as an overlapping region. The overlapping region extraction unit A1009 may also extract an overlapping region in the human body even when the determination result indicates that the person is not performing an explanation movement.
[0028] The transparency change unit A1010 changes the transparency of at least a portion of the explanatory material. Specifically, it changes the transparency of the explanatory material input from the overlap area extraction unit A1009 so as to increase the transparency of the area that overlaps with the human body area, i.e., the transparency of the overlap area. Note that the transparency change unit A1010 may, for example, change the transparency of the entire explanatory material, or may change the transparency of so-called blank areas in the explanatory material where no figures or text are present. The transparency may also change over time. Any transparency level may be used, such as semi-transparent or completely transparent. The transparency change unit A1010 outputs the explanatory material with the changed transparency and the video information to the image superimposition unit A1011. In this embodiment, the transparency change unit A1010 also functions as a transparency change unit that increases the transparency of at least a portion of the explanatory material in accordance with the overlap area.
[0029] The image superimposition unit A1011 superimposes explanatory materials on the video information. Specifically, when explanatory materials with changed transparency and video information are input from the transparency change unit A1010, the image superimposition unit A1011 superimposes the explanatory materials with changed transparency on the video information. On the other hand, when explanatory materials with changed transparency are not input from the transparency change unit A1010, the image superimposition unit A1011 superimposes explanatory materials with unchanged transparency (explanatory materials input from the material acquisition unit A1005) on the video information. In other words, the image superimposition unit A1011 functions as a superimposition unit that performs a process of superimposing explanatory materials with different transparencies on images in the video information depending on whether the transparency has been changed or not, and generates a superimposed image. The image superimposition unit A1011 outputs the video on which these explanatory materials have been superimposed as a superimposed video to the video output device A1012.
[0030] The video output device A1012 outputs video and image information. Specifically, it outputs the superimposed video input from the image superimposition unit A1011 to the monitor device A1013. In this embodiment, the video output device A1012 also functions as a display control means for displaying the superimposed video composed of the superimposed images on the screen of the monitor device A1013. The monitor device A1013 is a display device that displays the superimposed video input from the video output device A1012 on its screen.
[0031] An example of how the image superimposing unit A1011 superimposes explanatory material on video information will be described below with reference to Figs. 5 to 9. Figs. 5 and 9 are diagrams illustrating how the area overlapping with a human body is not made transparent when the human body is not making an explanatory movement. Figs. 5(A) and 9(A) are diagrams illustrating an example of an image of a human body. Figs. 5(B) and 9(B) are diagrams illustrating how explanatory material is superimposed on video information. Figs. 5 and 9 show how the image superimposing unit A1011 superimposes explanatory material with its transparency unchanged on video information.
[0032] In FIG. 5, D201 represents a video of a human body. D202 represents a video (superimposed video) in which explanatory material is superimposed on D201. P203 represents the explanatory material. P202 represents the human body before it moves, and P202 represents the human body after it moves. In video D201, the human body has simply moved and is not performing any explanatory action, so the image superimposition unit A1011 superimposes explanatory material with its transparency unchanged on the video of the human body. Therefore, in superimposed video D202, the area of explanatory material P203 that overlaps with the human body area is not made transparent.
[0033] In FIG. 9, D601 represents a video of a human body. D602 represents a video in which explanatory material is superimposed on D601 (superimposed video). P603 represents explanatory material. P601, P602, and P604 each represent a human body. Specifically, P601 represents the human body before moving, P602 represents the human body after moving, and P604 represents the human body superimposed on D602. In video D601, the human body has simply moved and is not performing an explanatory action, so the image superimposition unit A1011 superimposes explanatory material with its transparency unchanged on the video of the human body. Therefore, in the superimposed video of superimposed video D602, the area of explanatory material P603 that overlaps with the human body area is not made transparent.
[0034] Fig. 6 is a diagram illustrating a state in which an area overlapping a human body is made transparent when the human body is making an explanatory motion. That is, Fig. 6 shows a state in which the image superimposing unit A1011 superimposes explanatory material with changed transparency on video information. Fig. 6(A) is a diagram showing an example of an image of a human body. Fig. 6(B) is a diagram showing the state in which explanatory material is superimposed on video information.
[0035] In FIG. 6, D301 represents an image of a human body. D302 represents an image (superimposed image) in which explanatory material is superimposed on D301. P301 and P302 represent human bodies. P303 represents explanatory material. Because human body P301 shown in FIG. 6 is performing an explanation action, image superimposition unit A1011 superimposes explanatory material with changed transparency on the image of the human body. Therefore, in superimposed image D302, the area of explanatory material P303 that overlaps with human body P302 (superimposed area) is made transparent, making it possible to visually see which part of the explanatory material human body P301 is explaining.
[0036] 7 and 8 are diagrams illustrating how an area overlapping a part of a human body is made transparent when the human body is making an explanatory motion. FIGS. 7(A) and 8(A) are diagrams illustrating an example of an image of a human body. FIGS. 7(B) and 8(B) are diagrams illustrating how explanatory material is superimposed on video information. FIGS. 7 and 8 illustrate how the image superimposition unit A1011 superimposes explanatory material with changed transparency on video information.
[0037] In FIG. 7, D401 represents an image of a human body. D402 represents an image (superimposed image) in which explanatory material is superimposed on D401. P401 represents a human body. P402 represents an arm used in the explanation. P403 represents the explanatory material. Because the human body P401 shown in FIG. 7 is performing an explanation action, the image superimposition unit A1011 superimposes explanatory material with changed transparency on the image of the human body. Therefore, in the superimposed image D402, the area of the explanatory material P403 that overlaps with the area of the arm P402 of the human body P401 used in the explanation (superimposed area) is made transparent, making it possible to see which part of the explanatory material the human body P401 is explaining.
[0038] In FIG. 8, D501 represents a video of a human body. D502 represents a video (superimposed video) in which explanatory material is superimposed on D501. P501 represents a human body. P502 represents the head of the human body. P503 represents the explanatory material. Because human body P501 shown in FIG. 8 is performing an explanation action, image superimposition unit A1011 superimposes explanatory material with changed transparency on the video of the human body. Therefore, in superimposed video D502, the area of explanatory material P503 that overlaps with the area of head P502 of human body P501 (superimposed area) is made transparent, making it possible to see what human body P501 is explaining.
[0039] In this way, by making the overlapping area, which is the area that overlaps with the human body that is making the explanation motion, transparent only when the human body is making the explanation motion, it is possible to confirm on the screen which part of the explanatory material the human body is explaining when the human body is explaining. Furthermore, when the human body is not explaining, it is possible to confirm the entire explanatory material on the screen.
[0040] Fig. 2 is a diagram showing an example of the hardware configuration of the image superimposing device A1003. As shown in Fig. 2, the image superimposing device A1003 is composed of a CPU 11, a ROM 12, a RAM 13, a secondary storage device 14, a communication device 15, an input device 16, a communication I / F unit 17, and a connection bus 18.
[0041] The CPU (processor) 11 is a central processing unit that controls the automatic photography system A1000, including the image superimposing device A1003, by executing control programs stored in the ROM 12 and RAM 13. In other words, it comprehensively controls each component of the automatic photography system A1000, including the image superimposing device A1003. CPU is an acronym for Central Processing Unit.
[0042] The ROM 12 is a non-volatile memory that stores a control program and various parameter data. The control program is executed by the CPU 11 to realize various processes performed by the image superimposition device A1003, which will be described later. The RAM 13 is a volatile memory that temporarily stores images, videos, the control program, and the execution results thereof.
[0043] The secondary storage device 14 is a rewritable secondary storage device such as a hard disk or flash memory, and stores data received via the communication device 15. It also stores control programs, various setting contents, processing results, etc. This information is output to the RAM 13 and used by the CPU 11 to execute the program.
[0044] The communication device 15 is a wired communication unit that communicates with various devices. Note that the communication device 15 is not limited to a wired communication unit, and may be a wireless communication unit. The input device 16 is a mouse, keyboard, touch panel display, etc. that accepts user input.
[0045] The communication I / F unit 17 is an interface for connecting to a detachable device and includes, for example, a power supply and an attachment mechanism such as a detachable socket for attaching and detaching the detachable device. The image superimposing device A1003 performs data communication with the detachable device via the communication I / F unit 17. The connection bus 18 connects the CPU 11, ROM 12, RAM 13, secondary storage device 14, communication device 15, input device 16, and communication I / F unit 17 that constitute the image superimposing device A1003, and performs data communication among them.
[0046] In this embodiment, the processing in each functional unit is realized by software using the CPU 11 of the image superimposition device A1003, but part or all of the processing of the image superimposition device A1003 may be realized by hardware, which may be a dedicated circuit (ASIC) or a processor (reconfigurable processor, DSP), etc.
[0047] Here, the processing procedure of the automatic photography system A1000 according to the first embodiment will be described with reference to the flowchart in Fig. 10. Fig. 10 is a flowchart showing the processing procedure of the automatic photography system A1000 according to the first embodiment. Each of the following processes is realized by the CPU 11 of the image superimposition device A1003 executing a program stored in the ROM 12 or the like. Each process (step) is denoted by adding an S to the beginning to omit the notation of the process (step). The automatic photography system A1000 starts the automatic photography system when activated by a user operation via the input device 16 or the like.
[0048] First, in S001, the video acquisition unit A1004 acquires video information from the video acquisition device A1001. Then, the process proceeds to S002.
[0049] Next, in S002, the material acquisition unit A1005 acquires explanatory materials from the material acquisition device A1002. After acquiring the explanatory materials, the material acquisition unit A1005 outputs the acquired explanatory materials to the overlap region extraction unit A1009 and the image superimposition unit A1011. Then, the process proceeds to S003.
[0050] Next, in S003, the region division processing unit A1008 performs region division processing using the video information acquired from the video acquisition unit A1004 (first extraction step). Then, the region division processing unit A1008 outputs the divided region information to the overlap region extraction unit A1009. After that, the process proceeds to S004.
[0051] Next, in S004, the skeletal information estimation unit A1006 estimates skeletal information of the human body using the video information acquired from the video acquisition unit A1004. The skeletal information estimation unit A1006 outputs the estimated skeletal information as a skeletal estimation result to the human body movement determination unit A1007. Then, the process proceeds to S005.
[0052] Next, in S005, the human body movement determination unit A1007 estimates a human body movement using the human body skeleton estimation result acquired from the skeleton information estimation unit A1006, and determines whether the movement is an explanatory movement (first determination step). If the determination result indicates an explanatory movement (Yes in S005), the human body movement determination unit A1007 outputs the determination result and the skeleton estimation result to the overlap region extraction unit A1009. Then, the process proceeds to S006. On the other hand, if the movement is not an explanatory movement (No in S005), the determination result is output to the image superimposition unit A1011. Then, the process proceeds to S008.
[0053] Next, in S006, the overlapping area extraction unit A1009 extracts an area in the explanatory material that overlaps with the human body that made the explanatory movement (second extraction step). Specifically, based on the determination result and skeleton estimation result input from the human body movement determination unit A1007, the human body area information input from the area division processing unit A1008, and the explanatory material input from the material acquisition unit A1005, the overlapping area in the explanatory material that overlaps with the human body that made the explanatory movement is extracted. The overlapping area extraction unit A1009 outputs the extracted area (overlapping area) and the explanatory material to the transparency change unit A1010. Then, the process proceeds to S007.
[0054] Next, in S007, the transparency change unit A1010 changes the transparency of the explanatory material according to the overlapping area input from the overlapping area extraction unit (changing step). Specifically, the transparency change unit A1010 changes the transparency of the overlapping area to be higher. Then, the transparency change unit A1010 outputs the explanatory material with the changed transparency to the image superimposition unit A1011. Then, the process proceeds to S008.
[0055] Next, in S008, the image superimposing unit A1011 superimposes explanatory material on the video information (superimposing step). Here, if the determination result obtained from the human body movement determination unit A1007 is that the human body is not making an explanatory movement, the image superimposing unit A1011 superimposes the explanatory material obtained from the material acquisition unit A1005 on the video information obtained from the human body movement determination unit A1007. On the other hand, if the explanatory material with changed transparency is obtained from the transparency change unit A1010, the image superimposing unit A1011 superimposes the explanatory material with changed transparency on the video information input from the transparency change unit A1010. The image superimposing unit A1011 outputs the superimposed video information (superimposed video) to the video output device A1012. Then, the process proceeds to S009.
[0056] Next, in S009, the video output device A1012 outputs the video information (superimposed video) input from the image superimposition unit A1011 to the monitor device A1013. When the video information is input from the video output device A1012, the monitor device A1013 displays the video or image in the video information on the screen. Then, the process proceeds to S010.
[0057] Next, in S010, it is determined whether or not to end the process. Specifically, it is determined whether or not the user has operated the On / Off switch of the automatic photography system (not shown) to stop the automatic photography process. If the result of the determination is that the operation to stop the automatic photography process has not been performed (NO in S010), the process proceeds to S001 and the same process is repeated. On the other hand, if the operation to stop the automatic photography process has been performed (YES in S010), the automatic photography process is ended, and this processing flow is terminated.
[0058] As described above, when the automatic photography system A1000 in the first embodiment superimposes explanatory material on video information, it can change the transparency of the area of the explanatory material that overlaps with the human body area only when the human body is making an explanatory motion. This allows the viewer to see on the screen which part of the explanatory material the human body is explaining when the human body is making an explanatory motion. When the human body is not making an explanatory motion, the entire explanatory material can be seen on the screen. Therefore, the viewer viewing the explanatory material can view the explanatory material more clearly.
[0059] <Embodiment 2> An example of the configuration of an image superimposing device B1003 according to the second embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the functional configuration of an automatic photography system B1000 including an image superimposing device B1003 according to the second embodiment. Note that detailed description of the configuration of devices and functional units similar to those of the automatic photography system A1000 according to the first embodiment will be omitted below. Furthermore, the hardware configuration is also similar to that of the automatic photography system A1000 according to the first embodiment, and therefore description thereof will be omitted.
[0060] The automatic photography system B1000 detects a human body from the captured video and determines whether the human body is speaking from audio information acquired from a microphone. If the human body is speaking, it determines that the human body is providing an audio explanation, and makes the area overlapping with the human body area providing the explanation (overlapping area) transparent (by changing the transparency) and superimposes it on the video information (video of the human body). The processing system then displays the superimposed image on a monitor.
[0061] The automatic photography system B1000 is configured to have an image capture device A1001, a document capture device A1002, an image superimposition device B1003, a monitor device A1013, and an audio capture device B1014. The image superimposition device B1003, the image capture device A1001, the document capture device A1002, the monitor device A1013, and the audio capture device B1014 are communicably connected. The image superimposition device B1003 and the monitor device A1013 are connected via a line such as a video interface.
[0062] The image superimposing device B1003 acquires human body region information from the video input from the video capturing device A1001 and determines whether the human body is speaking from the audio information input from the audio capturing device B1014. If the human body is speaking, it is assumed that the human body is providing an audio explanation, and the image superimposing device B1003 makes the region of the explanatory material that overlaps with the human body region transparent and superimposes it on the video information. After the speech ends, the image superimposing device B1003 continues to make the region of the explanatory material that overlaps with the human body region transparent for a certain period (predetermined period) assuming that the explanation is continuing. The image superimposing device B1003 then outputs the superimposed image, which is the superimposed image, to the monitor device A1013.
[0063] The image superimposing device B1003 is configured to have, as functional units, a video acquisition unit A1004, a document acquisition unit A1005, a region division processing unit A1008, an overlap region extraction unit A1009, a transparency change unit A1010, and an image superimposing unit B1011. The image superimposing device B1003 is further configured to have, as functional units, a video output device A1012, a sound input unit B1015, an utterance determination unit B1016, an elapsed time measurement unit B1017, and a judgment unit B1018. As in the first embodiment, each of these functional units is realized by the CPU 11 loading a program stored in the ROM 12 into the RAM 13 and executing it. The CPU 11 then stores the execution results of each process, which will be described later, in the RAM 13 or a predetermined storage medium.
[0064] The voice capturing device B1014 is a device that generates sound information by collecting sounds around the voice capturing device B1014 with a microphone. The voice capturing device B1014 outputs the generated sound information to the sound input unit B1015.
[0065] The sound input unit B1015 receives as input sound information generated by the voice acquisition unit B1014. The sound input unit B1015 outputs the sound information to the speech determination unit B1016 as voice information (voice data).
[0066] The speech determination unit B1016 determines whether or not there is speech from the audio information input from the sound input unit B1015. When making the determination, the speech determination unit B1016 determines that there is audio explanation if there is speech. To determine whether or not there is speech, voice activity detection (VAD) is used, which determines speech intervals and other intervals from sound data containing speech and other sounds. Note that voice activity detection is a known technology, so a detailed description will be omitted. The speech determination unit B1016 performs voice activity detection on the audio data, and determines that there is speech if a voice interval is present. In this embodiment, the speech determination unit B1016 also functions as a second determination means for determining whether or not a human body is speaking based on audio information.
[0067] If it is determined that there is speech, the speech determination unit B1016 outputs first information, which is information that the human body is explaining by voice, to the overlap region extraction unit A1009. On the other hand, if there is no speech, it outputs second information, which is information that there is no speech by the human body, to the elapsed time measurement unit B1017.
[0068] The elapsed time measurement unit B1017 measures the elapsed time since the human body's utterance ended. Specifically, the elapsed time measurement unit B1017 measures the elapsed time since the human body's utterance ended based on the second information input from the utterance determination unit B1016. Note that the elapsed time measurement unit B1017 ends time measurement if the second information is not input. The elapsed time measurement unit B1017 outputs the measured elapsed time since the utterance ended to the determination unit B1018 as a measured time (time elapsed since the utterance ended). In this embodiment, the elapsed time measurement unit B1017 also functions as a measurement means that measures the time since the human body's utterance ended based on the second information and outputs the measured time.
[0069] The determination unit B1018 determines (judges) whether to continue making the explanatory material transparent. Specifically, based on the measured time input from the elapsed time measurement unit B1017, it determines whether to continue making the explanatory material transparent. That is, it determines whether to continue changing the transparency of the overlapping region, which is the region in the explanatory material that overlaps with the human body. If the determination unit B1018 determines to continue making the explanatory material transparent, it outputs information indicating that the transparency of the explanatory material will continue to the overlapping region extraction unit A1009. On the other hand, if it determines not to continue making the explanatory material transparent, it outputs information indicating that the transparency will not be continued to the image superimposition unit B1011. The processing of the determination unit B1018 in this embodiment will be described below with reference to FIG. 13.
[0070] FIG. 13 is a diagram showing speech intervals of a human body and transparent intervals of a document area according to the second embodiment. In FIG. 13, P801, P802, P803, P804, and P805 represent speech intervals of a human body, respectively. P806, P807, P808, P809, and P810 represent the duration of transparency of the explanatory material for each speech time, respectively. FIG. 13 shows that transparency begins at the start of each speech and continues for a certain period of time after the speech ends. As described above, the transparency change unit A1010 in this embodiment continues to change the transparency of the explanatory material until a predetermined time has elapsed since the speech of the human body ended, based on the time measured by the elapsed time measurement unit B1017. Furthermore, in this embodiment, the predetermined period (predetermined time) for which transparency of the explanatory material continues after the speech ends is set to 10 seconds. However, this is only a guideline, and the duration of transparency of the explanatory material may be any number of seconds.
[0071] The image superimposition unit B1011 superimposes explanatory material on the video information. Specifically, when explanatory material with changed transparency and video information are input from the transparency change unit A1010, the image superimposition unit B1011 superimposes the explanatory material with changed transparency on the video information. Furthermore, when explanatory material with changed transparency is not input from the transparency change unit A1010, the image superimposition unit B1011 superimposes explanatory material with unchanged transparency (explanatory material input from the material acquisition unit A1005) on the video information. Furthermore, when information indicating that the change in transparency will not be continued is input from the determination unit B1018, the image superimposition unit B1011 superimposes explanatory material that has not been made transparent on the video information. The image superimposition unit B1011 outputs the video information on which these explanatory materials have been superimposed to the video output device A1012. Here, an example of the superimposition process of the image superimposition unit B1011 in the second embodiment will be described below with reference to FIG. 12.
[0072] FIG. 12 is a diagram illustrating superimposing explanatory material, with the region overlapping the human body region made transparent, on video information when the human body is speaking, according to the second embodiment. In FIG. 12, D701 represents a video of a human body. D702 represents a video in which explanatory material is superimposed on the video D701. P701 and P703 each represent a human body. P702 represents an example of a human body speaking. P704 represents explanatory material. Here, the human body P701 shown in FIG. 12 is speaking. Therefore, it is determined that an explanation is being given through speech. Therefore, the image superimposing unit B1011 superimposes explanatory material, with the transparency changed, on video information in which the human body P703 is captured, as shown in explanatory material D702. In this way, the region of the explanatory material overlapping the human body region (superimposed region) in the superimposed video on which explanatory material D702 is superimposed is made transparent.
[0073] In this way, when a human body gives a speech explanation, the area of the explanatory material that overlaps with the area of the human body is made transparent, making it possible to visually see which part of the explanatory material the human body is explaining.
[0074] Here, the processing procedure of the automatic photography system B1000 according to the second embodiment will be described with reference to the flowchart in Fig. 14. Fig. 14 is a flowchart showing the processing procedure of the automatic photography system B1000 according to the second embodiment. Each of the following processes is realized by the CPU 11 of the image superimposition device B1003 executing a program stored in the ROM 12 or the like. Each process (step) is denoted by adding an S to the beginning to omit the notation of the process (step). The automatic photography system B1000 starts the automatic photography system when started by a user operation via the input device 16 or the like.
[0075] First, in S101, the video acquisition unit A1004 acquires video information from the video acquisition device A1001, and then the process proceeds to S102.
[0076] Next, in S102, the sound input unit B1015 acquires sound information from the sound acquisition device B1014. After that, the process proceeds to S103.
[0077] Next, in S103, the material acquisition unit A1005 acquires explanatory materials from the material acquisition device A1002. After acquiring the explanatory materials, the material acquisition unit A1005 outputs the acquired explanatory materials to the overlap region extraction unit A1009 and the image superimposition unit B1011. Then, the process proceeds to S104.
[0078] Next, in S104, the region division processing unit A1008 performs region division processing using the video information acquired from the video acquisition unit A1004. Then, the region division processing unit A1008 outputs the divided region information to the overlap region extraction unit A1009. After that, the process proceeds to S104.
[0079] Next, in S105, the speech determination unit B1016 performs speech section detection using the audio information input from the sound input unit B1015, and determines whether or not a human body is providing an explanation by voice. If the determination result shows that a human body is providing an explanation by voice (Yes in S105), the speech determination unit B1016 outputs information (first information) that the human body is providing an explanation by voice to the overlap region extraction unit A1009. Then, the process proceeds to S108. On the other hand, if no explanation by voice is being provided (No in S105), the speech determination unit B1016 outputs information (second information) that the human body is not providing an utterance to the elapsed time measurement unit B1017. Then, the process proceeds to S106.
[0080] Next, in S106, the elapsed time measurement unit B1017 measures the elapsed time since the end of the utterance based on the second information input from the utterance determination unit B1016. Then, the elapsed time measurement unit B1017 outputs the measured time to the determination unit B1018. After that, the process proceeds to S107.
[0081] Next, in S107, the determination unit B1018 determines whether the elapsed time since the end of the utterance input from the elapsed time measurement unit B1017 has exceeded a certain period of time. If the determination result shows that the certain period of time has passed (Yes in S107), information indicating that the transparency should not be changed (transparency should not be continued) is output to the image superimposition unit B1011. Then, the process proceeds to S110. On the other hand, if the certain period of time has not passed (No in S107), information indicating that the transparency should be changed (transparency should be continued) is output to the overlapping area extraction unit A1009. Then, the process proceeds to S108.
[0082] Next, in S108, the overlap region extraction unit A1009 extracts an overlap region from the explanatory material. Specifically, when first information is input from the speech determination unit B1016 or information to change transparency is input from the judgment unit B1018, the overlap region is extracted using the human body region information input from the region segmentation processing unit and the explanatory material input from the material acquisition unit A1005. Then, the overlap region extraction unit A1009 outputs the extracted overlap region to the transparency change unit A1010. Then, the process proceeds to S109.
[0083] Next, in S109, the transparency change unit A1010 changes the transparency of the explanatory material (makes it transparent) using the explanatory material and the overlapping area input from the overlapping area extraction unit A1009. Then, the transparency change unit A1010 outputs the explanatory material with the changed transparency to the image superimposition unit B1011. After that, the process proceeds to S110.
[0084] Next, in S110, the image superimposition unit B1011 superimposes explanatory material on the video information. Here, if explanatory material with changed transparency is input from the transparency change unit A1010, the image superimposition unit B1011 superimposes the explanatory material with the changed transparency on the video information. If explanatory material with changed transparency is not input from the transparency change unit A1010, the image superimposition unit B1011 superimposes explanatory material with unchanged transparency acquired from the material acquisition unit A1005 on the video information. If information indicating that the transparency will not be changed is input from the determination unit B1018, the image superimposition unit B1011 superimposes the explanatory material with unchanged transparency acquired from the material acquisition unit A1005 on the video information. Then, the image superimposition unit B1011 outputs the superimposed image to the video output device A1012. After that, the process proceeds to S111.
[0085] Next, in S111, the video output device A1012 outputs the video information (superimposed video) input from the image superimposition unit B1011 to the monitor device A1013. When the video information is input from the video output device A1012, the monitor device A1013 displays the video or image in the video information on the screen. Then, the process proceeds to S112.
[0086] Next, in S112, it is determined whether or not to end the process. Specifically, it is determined whether or not the user has operated the On / Off switch of the automatic photography system (not shown) to stop the automatic photography process. If the result of the determination is that the operation to stop the automatic photography process has not been performed (NO in S112), the process proceeds to S101 and the same process is repeated. On the other hand, if the operation to stop the automatic photography process has been performed (YES in S112), the automatic photography process is ended, and this processing flow is terminated.
[0087] As described above, when the automatic photography system B1000 in the second embodiment superimposes explanatory material on video information, it can change the transparency of the area of the explanatory material that overlaps with the human body area only when the human body is giving an explanation through voice. This makes it possible to visually confirm on the screen which part of the explanatory material the human body is explaining when the human body is giving an explanation. Furthermore, when the human body is not giving an explanation, the entire explanatory material can be visually confirmed on the screen.
[0088] <Embodiment 3> An example of the configuration of an image superimposing device C1000 according to the third embodiment will be described with reference to Fig. 15. Fig. 15 is a block diagram showing the functional configuration of an automatic photography system C1000 including an image superimposing device C1000 according to the third embodiment. Note that detailed description of the configuration of devices and functional units similar to those of the automatic photography systems A1000 and B1000 according to the first and second embodiments will be omitted below. Furthermore, the hardware configuration is also similar to that of the automatic photography system A1000 according to the first embodiment, and therefore description thereof will be omitted.
[0089] The automatic photography system C1000 detects a human body from the captured video and determines whether the human body is speaking from the audio information acquired from the microphone. If the human body is speaking, it extracts predetermined keywords from the audio information and explanatory materials. If the keywords from both sources match or are similar, it determines that the human body is providing an audio explanation, and makes transparent (changes the transparency of) the area that overlaps with the human body area providing the explanation in the acquired explanatory materials and superimposes it on the video information. The processing system then displays the superimposed image on a monitor.
[0090] The automatic photography system C1000 is configured to have an image capture device A1001, a document capture device A1002, an image superimposition device C1003, a monitor device A1013, and an audio capture device B1014. The image superimposition device C1003 is communicably connected to the image capture device A1001, the document capture device A1002, the monitor device A1013, and the audio capture device B1014. The image superimposition device C1003 and the monitor device A1013 are connected via a line such as a video interface.
[0091] The image superimposing device C1003 acquires human body region information from the video input from the video capturing device A1001 and determines whether a human body is speaking from the audio information input from the audio capturing device B1014. It then extracts predetermined keywords from the audio information and explanatory materials, and if the keywords match or are similar, it assumes that a human body is providing an audio explanation, changes the transparency of the region in the explanatory materials that overlaps with the human body region, and superimposes it on the video information. It then continues to change the transparency of the region in the explanatory materials that overlaps with the human body region (overlapping region) for a certain period after the end of the speech, assuming that the explanation is continuing. The image superimposing device C1003 then outputs the superimposed image to the monitor device A1013.
[0092] The image superimposing device C1003 is configured to have, as functional units, a video acquisition unit A1004, a document acquisition unit A1005, a region division processing unit A1008, an overlap region extraction unit A1009, a transparency change unit A1010, an image superimposing unit B1011, and a video output device A1012. The image superimposing device B1003 is further configured to have, as functional units, a sound input unit B1015, an utterance determination unit B1016, an elapsed time measurement unit C1017, a judgment unit B1018, a keyword extraction unit C1019, and a match determination unit C1020. As in the first embodiment, each of these functional units is realized by the CPU 11 loading a program stored in the ROM 12 into the RAM 13 and executing it. The CPU 11 then stores the execution results of each process, which will be described later, in the RAM 13 or a predetermined storage medium.
[0093] The keyword extraction unit C1019 extracts keywords from both the audio information and the explanatory materials. Specifically, it extracts predetermined keywords from the audio information input from the speech determination unit B1016 and the explanatory materials input from the material acquisition unit A1005. Keywords are extracted from the audio information using voice recognition technology. Keywords are extracted from the explanatory materials using OCR (Optical Character Recognition) or from tag information previously embedded in the explanatory materials. The keyword extraction unit C1019 outputs the extracted keywords to the match determination unit C1020. In this embodiment, the keyword extraction unit C1019 functions as a third extraction means that extracts keywords from both the audio information and the explanatory materials.
[0094] The match determination unit C1020 determines whether the keywords extracted by the keyword extraction unit C1019 match or are similar. Specifically, it determines whether the keywords extracted from the utterance content input from the keyword extraction unit C1019 match or are similar to the keywords extracted from the explanatory materials. If the utterance content and the explanatory content match or are similar, the match determination unit C1020 outputs third information, which is information on the match or similarity, to the overlap area extraction unit A1009. On the other hand, if the utterance content and the explanatory content do not match or are similar, the match determination unit C1020 outputs fourth information, which is information on the match or dissimilarity, to the elapsed time measurement unit C1017 and the image superimposition unit B1011. In this embodiment, the match determination unit C1020 also functions as a third determination means for determining whether the keywords extracted by the keyword extraction unit C1019 match or are similar.
[0095] Fig. 16 is an explanatory diagram for identifying an explanatory area from the content of a human utterance and the content of explanatory material according to embodiment 3. Fig. 16(A) is a diagram showing an example of explanatory material and an area in the explanatory material. Fig. 16(B) is a diagram showing an example of keywords extracted from the utterance. An explanatory area is an area showing explanatory figures, tables, images, videos, animations, characters, etc. displayed in the explanatory material.
[0096] In FIG. 16(A), P901 represents the explanatory material. P902 is an area within the explanatory material, representing area 1 in the table. P903 is an area within the explanatory material, representing area 2 in the table. In FIG. 16(B), the table shows keywords extracted from the utterance, specifically, keywords extracted from area P902, which is area 1, and area P903, which is area 2. Using FIG. 16 as an example, when the keywords of the utterance content are compared with the keywords of area P902, which is area 1, and area P903, which is area 2, the keywords for area and area match the keywords of area P903, which is area 2. Therefore, in the example shown in FIG. 16, the match determination unit C1020 can determine that a human body is providing an explanation by voice to area P903, which is area 2.
[0097] The elapsed time measurement unit C1017 measures the elapsed time since the human body's utterance ended. Specifically, based on the second information input from the utterance determination unit B1016 and the fourth information input from the match determination unit C1020, it measures the elapsed time since the utterance that matches the content of the explanatory material ended. If the second information is not input, the time measurement ends. Then, the elapsed time measurement unit C1017 outputs the measured time to the determination unit B1018.
[0098] Here, the processing procedure of the automatic photography system C1000 according to the third embodiment will be described with reference to the flowcharts of Fig. 17 and Fig. 18. Fig. 17 and Fig. 18 are flowcharts showing the processing procedure of the automatic photography system C1000 according to the third embodiment. Each of the following processes is realized by the CPU 11 of the image superimposition device C1003 executing a program stored in the ROM 12 or the like. Each process (step) is denoted by adding an S to the beginning to omit the notation of the process (step). The automatic photography system C1000 starts the automatic photography system when started by a user operation via the input device 16 or the like.
[0099] First, in S201, the video acquisition unit A1004 acquires video information from the video acquisition device A1001, and then the process proceeds to S202.
[0100] Next, in S202, the sound input unit B1015 acquires sound information from the sound acquisition device B1014. After that, the process proceeds to S203.
[0101] Next, in S203, the material acquisition unit A1005 acquires explanatory materials from the material acquisition device A1002. After acquiring the explanatory materials, the material acquisition unit A1005 outputs the acquired explanatory materials to the overlap region extraction unit A1009 and the image superimposition unit B1011. Then, the process proceeds to S204.
[0102] Next, in S204, the region division processing unit A1008 performs region division processing using the video information acquired from the video acquisition unit A1004. Then, the region division processing unit A1008 outputs the divided region information to the overlap region extraction unit A1009. After that, the process proceeds to S204.
[0103] Next, in S205, the speech determination unit B1016 performs speech section detection using the audio information input from the sound input unit B1015, and determines whether or not a human body is providing an explanation by voice. If the result of the determination is that an explanation by voice is being provided (Yes in S205), the speech determination unit B1016 outputs information (first information) that the human body is providing an explanation by voice to the keyword extraction unit C1019. Then, the process proceeds to S206. On the other hand, if an explanation by voice is not being provided (No in S205), the process outputs information (second information) that the human body is not providing an explanation by voice to the elapsed time measurement unit C1017. Then, the process proceeds to S209.
[0104] Next, in S206, the keyword extraction unit C1019 extracts keywords from the voice information input from the utterance determination unit B 1016. Then, the keyword extraction unit C1019 outputs the extracted keywords to the match determination unit C 1020. After that, the process proceeds to S207.
[0105] Next, in S207, the keyword extraction unit C1019 extracts keywords from the explanatory materials input from the material acquisition unit A1005. The keyword extraction unit C1019 outputs the extracted keywords to the match determination unit C1020. Thereafter, the process proceeds to S208. Note that the order of the processes of S206 and S207 may be reversed.
[0106] Next, in S208, the match determination unit C1020 determines whether the audio information input from the keyword extraction unit C1019 and the keywords extracted from the explanatory materials match or are similar. If the determination result shows that the audio information and the keywords match or are similar (Yes in S208), the match determination unit C1020 outputs information that matches the explanatory content (third information) to the overlap area extraction unit A1009. Then, the process proceeds to S211. On the other hand, if the audio information and the keywords do not match or are not similar (No in S208), the match determination unit C1020 outputs information that does not match the explanatory content (fourth information) to the image superimposition unit B1011 and the elapsed time measurement unit C1017. Then, the process proceeds to S213.
[0107] Next, in S209, the elapsed time measurement unit C1017 measures the elapsed time since the end of the utterance whose content does not match the explanatory material, based on the second information input from the utterance determination unit B1016 and the fourth information input from the match determination unit C1020. The elapsed time measurement unit C1017 outputs the measured time (measured time) to the determination unit B1018. Then, the process proceeds to S210.
[0108] Next, in S210, the determination unit B1018 determines whether a certain period of time has passed since the utterance inconsistent with the explanatory material ended or the utterance inconsistent with the explanatory material ended, based on the measured time input from the elapsed time measurement unit C1017. If the determination result shows that the utterance inconsistent with the explanatory material ended or the certain period of time has passed (Yes in S210), the determination unit B1018 outputs information to the image superimposition unit B1011 not to make the explanatory material transparent. Then, the process proceeds to S213. On the other hand, if the utterance inconsistent with the explanatory material has not ended or the certain period of time has not passed (No in S210), the determination unit B1018 outputs information to make the explanatory material transparent to the overlap region extraction unit A1009. Then, the process proceeds to S211.
[0109] Next, in S211, the overlapping area extraction unit A1009 extracts an overlapping area using the human body area information input from the area division processing unit A1008 and the explanatory materials input from the material acquisition unit A1005. Then, the overlapping area extraction unit A1009 outputs the extracted overlapping area to the transparency change unit A1010. After that, the process proceeds to S212.
[0110] Next, in S212, the transparency change unit A1010 changes the transparency of the explanatory material (makes it transparent) using the explanatory material and the overlapping area input from the overlapping area extraction unit A1009. Then, the transparency change unit A1010 outputs the explanatory material with the changed transparency to the image superimposition unit B1011. After that, the process proceeds to S213.
[0111] Next, in S213, the image superimposition unit B1011 superimposes explanatory material on the video information. Here, if explanatory material with changed transparency is input from the transparency change unit A1010, the image superimposition unit B1011 superimposes the explanatory material with changed transparency on the video information. Also, if information not to change the transparency is input from the determination unit B1018, the image superimposition unit B1011 superimposes the explanatory material with unchanged transparency input from the material acquisition unit A1005 on the video information. Also, if information indicating that the spoken content and the content of the explanatory material do not match (fourth information) is input from the match determination unit C1020, the image superimposition unit B1011 superimposes the explanatory material with unchanged transparency input from the material acquisition unit A1005 on the video information. Then, the image superimposition unit B1011 outputs the superimposed video to the video output device A1012. Thereafter, the process proceeds to S214.
[0112] Next, in S214, the video output device A1012 outputs the video information (superimposed video) input from the image superimposition unit B1011 to the monitor device A1013. When the video information is input from the video output device A1012, the monitor device A1013 displays the video or image in the video information on the screen. Then, the process proceeds to S215.
[0113] Next, in S215, it is determined whether or not to end the process. Specifically, it is determined whether or not the user has operated the On / Off switch of the automatic photography system (not shown) to stop the automatic photography process. If the result of the determination is that the operation to stop the automatic photography process has not been performed (NO in S215), the process proceeds to S201 and the same process is repeated. On the other hand, if the operation to stop the automatic photography process has been performed (YES in S215), the automatic photography process is ended, and this processing flow is terminated.
[0114] As described above, when the automatic photography system A1000 in the third embodiment superimposes explanatory material on video information, it changes the transparency of the area of the explanatory material that overlaps with the human body area only if the audio explanation by the human body matches or is similar to the content of the explanatory material. This makes it possible to visually identify on the screen which part of the explanatory material the human body is explaining when the human body is explaining, and to visually identify the entire explanatory material when the human body is not explaining.
[0115] <Embodiment 4> An example of the configuration of an image superimposing device D1003 according to the fourth embodiment will be described with reference to Fig. 19. Fig. 19 is a block diagram showing the functional configuration of an automatic photography system B1000 including an image superimposing device D1003 according to the fourth embodiment. Note that detailed descriptions of the configurations of devices and functional units similar to those of the automatic photography systems A1000, B1000, and C1000 according to the first, second, and third embodiments will be omitted below. Furthermore, the hardware configuration is also similar to that of the automatic photography system A1000 according to the first embodiment, and therefore description thereof will be omitted.
[0116] The automatic photography system D1000 detects a human body from the captured video. Then, when determining that a human body is making explanatory movements based on skeletal information, or when determining that an explanation is being given based on the similarity between audio information acquired from a microphone and the content of the explanatory material, if the area of the human body overlaps with the explanatory material, the system performs an emphasis process on the explanatory material. On the other hand, if the area of the human body does not overlap with the explanatory material, the system makes the overlapping area, which is the area of the explanatory material that overlaps with the human body, transparent (changes the transparency), and superimposes it on the captured video of the human body, and displays the result on a monitor.
[0117] The automatic photography system D1000 is configured to have an image capture device A1001, a document capture device A1002, an image superimposition device D1003, a monitor device A1013, and an audio capture device B1014. The image superimposition device D1003 is communicably connected to the image capture device A1001, the document capture device A1002, the monitor device A1013, and the audio capture device B1014. The image superimposition device D1003 and the monitor device A1013 are connected via a line such as a video interface.
[0118] The image superimposing device D1003 detects a human body from the video input from the video capturing device A1001 and determines whether an explanatory action is being performed based on the skeletal information of the detected human body. Then, when audio information is input from the audio capturing device B1014, it determines from the audio whether the human body is speaking and whether the audio matches the content of the explanatory material. If it determines that an explanation is being performed using an action or audio, it identifies the explanatory area being explained from the explanatory material and determines whether the explanatory area overlaps with the human body area. If the human body area and the explanatory area overlap, it performs a highlighting process on the explanatory area without performing a transparency process on the overlapping area and explanation area, which are the areas of the explanatory material overlapping with the human body. If they do not overlap, it makes the overlapping area transparent (changes its transparency) and superimposes it on the video information. The superimposed image is output to the monitor device A1013.
[0119] The image superimposition device D1003 is configured to include, as functional units, a video acquisition unit A1004, a document acquisition unit A1005, a region segmentation processing unit A1008, an overlap region extraction unit A1009, a transparency change unit A1010, a video output device A1012, and a sound input unit B1015. Furthermore, the image superimposition device D1003 is configured to include, as functional units, an utterance determination unit B1016, an elapsed time measurement unit C1017, a transparency continuation determination unit C1018, a keyword extraction unit C1019, and a match determination unit C1020. Furthermore, the image superimposition device D1003 is configured to include, as functional units, a description region identification unit D1021, an overlap determination unit D1022, a highlight frame superimposition unit D1023, and an image superimposition unit D1011. As in the first embodiment, each of these functional units is realized by the CPU 11 loading a program stored in the ROM 12 into the RAM 13 and executing it. Then, the CPU 11 stores the execution results of each process, which will be described later, in the RAM 13 or a predetermined storage medium.
[0120] The explanation area identification unit D1021 identifies an explanation area in the explanatory material, which is an area where the human body is making an explanatory movement. Specifically, the explanation area is identified using human body skeletal information, video information, region information, the explanatory material, and information on an explanation area that matches the audio explanation. The skeletal information and video information are input from the human body movement determination unit A1007. The region information is input from the region segmentation processing unit A1008. The explanatory material is input from the material acquisition unit A1005. Information on an explanation area that matches the audio explanation is input from the match determination unit C1020. The explanation area identification unit D1021 outputs the identified explanation area, region information, video information, and explanatory material to the overlap determination unit D1022.
[0121] FIG. 20 is a diagram illustrating how an explanation area is identified from an explanation action of a human body according to the fourth embodiment. In FIG. 20, P1001 represents a human body. P1002 and P1003 represent areas within the explanation area. P1004 represents a half-ray passing through the arm of the human body providing explanation. When half-ray P1004 intersects with each area, that area is identified as the area to be explained (explanation area). Whether a line intersects with a rectangular area can be determined by intersection determination. In FIG. 20, half-ray P1004 and area P1003 intersect, so the area where the human body is providing explanation can be identified. When information about an explanation area that matches the explanation provided by voice is input, that information is identified as the area to be explained (explanation area).
[0122] The overlap determination unit D1022 determines whether the human body region overlaps with the identified explanatory region. Specifically, it determines whether the human body region overlaps with the identified explanatory region using the identified explanatory region, region information, video information, and explanatory material input from the explanatory region determination unit D1021. If the determination result shows that the human body region overlaps with the identified explanatory region, the overlap determination unit D1022 outputs the identified explanatory region, video information, and explanatory material to the highlight frame superimposition unit D1023. On the other hand, if the human body region does not overlap with the identified explanatory region, it outputs the region information, video information, and explanatory material to the overlap region extraction unit A1009.
[0123] The highlight frame superimposing unit D1023 highlights the explanatory region of the explanatory material. Specifically, it superimposes a highlight frame on the explanatory region of the explanatory material using the identified explanatory region, video information, and explanatory material input from the overlap determination unit D1022. The highlight frame superimposing unit D1023 outputs the superimposed explanatory material and video information to the image superimposing unit D1011. The highlight frame to be displayed may be colored, and the thickness of the border may also be made variable. Furthermore, the color within the explanatory region may be changed to be different from the color outside the explanatory region, and the color, font, and size of figures and text in the explanatory region may be changed. Furthermore, the highlight frame may be made to blink, or a combination of these may be used to highlight the explanatory region.
[0124] The image superimposition unit D1011 superimposes explanatory material on the video information. Specifically, when explanatory material with a highlight frame superimposed thereon input from the highlight frame superimposition unit D1023 and explanatory material with an area overlapping with a human body made transparent input from the video information and the transparency change unit A1010 are input, the image superimposition unit D1011 superimposes the explanatory material with the changed transparency on the video information. Furthermore, when explanatory material with changed transparency is not input from the transparency change unit A1010, the image superimposition unit D1011 superimposes explanatory material with an unchanged transparency (explanatory material input from the material acquisition unit A1005) on the video information. Furthermore, when information indicating that the change in transparency will not be continued is input from the determination unit B1018, the image superimposition unit D1011 superimposes explanatory material that has not been made transparent on the video information. The image superimposition unit D1011 outputs the video information with these explanatory materials superimposed thereon to the video output device A1012. An example of the superimposition process of the image superimposition unit D1011 of the third embodiment will now be described with reference to FIGS.
[0125] Fig. 21 is a diagram showing an example of a case where a human body and an explanation region do not overlap. Fig. 21(A) is a diagram showing an example of video information. Fig. 21(B) is a diagram showing an example of explanatory material. Fig. 21(C) is a diagram showing an example of overlap between a human body region and an explanatory material region. Fig. 21(D) is a diagram showing an example of an explanation target region (explanation region).
[0126] In Figure 21, D1101 represents video information. D1102 represents explanatory material. D1103 is a diagram showing the overlap between the human body region and the explanatory material region. D1104 represents video information in which explanatory material D1102 is superimposed on video information D1101. P1101 represents a human body. P1102 represents a human body region. P1103 represents a region that is not the subject of explanation. P1104 represents an explanatory region. P1105, P1106, and P1107 are similar to P1102, P1103, and P1104, respectively, and therefore their explanation will be omitted. P1108 represents the same human body as P1101. P1109 represents explanatory material. In this case, the human body region P1105 and the explanatory region P1107 do not overlap.
[0127] Therefore, in D1104, the area (overlapping area) that overlaps with the human body area in explanatory material P1109 is made transparent and superimposed on the video information. As shown in Fig. 21, since the human body and the explanation area do not overlap, the explanation area can be seen even if the explanation area that overlaps with the human body area is made transparent.
[0128] Fig. 22 is a diagram showing an example of a case where a human body and an explanation region overlap. Fig. 22(A) is a diagram showing an example of video information. Fig. 22(B) is a diagram showing an example of explanatory material. Fig. 22(C) is a diagram showing an example of overlap between a human body region and an explanatory material region. Fig. 22(D) is a diagram showing an example of an explanation target region (explanation region).
[0129] In Figure 22, D1201 represents video information. D1202 represents explanatory material. D1203 is a diagram showing the overlap between the human body region and the explanatory material region. D1204 represents video information in which explanatory material D1202 is superimposed on video information D1201. P1201 represents a human body. P1202 represents a human body region. P1203 represents a region that is not the subject of explanation. P1204 represents an explanation region. P1205, P1206, and P1207 are similar to P1202, P1203, and P1204, respectively, and therefore their explanation will be omitted. P1208 represents a highlighted frame superimposed on the explanation region. P1209 represents explanatory material. In this case, the human body region P1205 and the explanation region P1207 overlap.
[0130] Therefore, in D1204, a highlight frame P1208 is superimposed on the explanation area in explanation area P1209. In this way, when the explanation area overlaps with the human body that is providing an explanation through action (explanatory action), highlighting processing is performed on the explanation area, so that the explanation area can be confirmed even if the human body cannot be seen.
[0131] Fig. 23 is a diagram showing an example of a case where a human body and an explanation region overlap. Fig. 23(A) is a diagram showing an example of video information. Fig. 23(B) is a diagram showing an example of explanatory material. Fig. 23(C) is a diagram showing an example of overlap between a human body region and an explanatory material region. Fig. 23(D) is a diagram showing an example of an explanation target region (explanation region).
[0132] In Figure 23, D1301 represents video information. D1302 represents explanatory material. D1303 is a diagram showing the overlap between the human body region and the explanatory material region. D1304 represents video information in which explanatory material D1302 is superimposed on video information D1301. P1301 represents a human body. P1302 represents the human body region, and P1303 represent regions that are not the subject of explanation. P1304 represents the explanation region. P1305, P1306, and P1307 are similar to P1302, P1303, and P1304, respectively, and therefore their explanation will be omitted. P1308 represents a highlighted frame superimposed on the explanation region. P1309 represents explanatory material. In this case, the human body region P1305 and the explanation region P1307 overlap.
[0133] Therefore, in D1304, a highlight frame is superimposed on the explanation area in the explanatory material P1309. When the human body providing the audio explanation overlaps with the explanation area in this way, highlighting the explanation area makes it possible to confirm the explanation area even if the human body cannot be seen.
[0134] Here, the procedure for processing the automatic photography system D1000 will be described with reference to the flowcharts of Fig. 24 and Fig. 25. Fig. 24 and Fig. 25 are flowcharts showing the processing procedure of the automatic photography system D1000 according to the fourth embodiment. Each of the following processes is realized by the CPU 11 of the image superimposition device D1003 executing a program stored in the ROM 12 or the like. Each process (step) is denoted by adding an S to the beginning to omit the notation of the process (step). The automatic photography system D1000 starts the automatic photography system when activated by a user operation via the input device 16 or the like.
[0135] First, in S301, the video acquisition unit A1004 acquires video information from the video acquisition device A1001, and then the process proceeds to S302.
[0136] Next, in S302, the sound input unit B1015 acquires sound information from the sound acquisition device B1014. After that, the process proceeds to S303.
[0137] Next, in S303, the material acquisition unit A1005 acquires explanatory materials from the material acquisition device A1002. After acquiring the explanatory materials, the material acquisition unit A1005 outputs the acquired explanatory materials to the overlap region extraction unit A1009 and the image superimposition unit B1011. Then, the process proceeds to S304.
[0138] Next, in S304, the region division processing unit A1008 performs region division processing using the video information acquired from the video acquisition unit A1004. Then, the region division processing unit A1008 outputs the divided region information to the overlap region extraction unit A1009. After that, the process proceeds to S305.
[0139] Next, in S305, the skeletal information estimation unit A1006 estimates skeletal information of the human body using the video information acquired from the video acquisition unit A1004. The skeletal information estimation unit A1006 outputs the estimated skeletal information as a skeletal estimation result to the human body movement determination unit A1007. Then, the process proceeds to S306.
[0140] Next, in S306, the human body movement determination unit A1007 estimates a human body movement using the human body skeleton estimation result acquired from the skeleton information estimation unit A1006, and determines whether the movement is an explanatory movement. If the determination result indicates that the movement is an explanatory movement (Yes in S306), the human body movement determination unit A1007 outputs the determination result and the skeleton estimation result to the explanatory region identification unit D1021. Then, the process proceeds to S313. On the other hand, if the movement is not an explanatory movement (No in S005), the determination result is output to the image superimposition unit D1011. Then, the process proceeds to S307.
[0141] Next, in S307, the utterance determination unit B1016 performs voice activity detection using the audio information input from the sound input unit B1015, and determines whether or not a human body is providing an explanation by voice. If the result of the determination is that an explanation by voice is being provided (Yes in S307), the utterance determination unit B1016 outputs information (first information) that a human body is providing an explanation by voice to the keyword extraction unit C1019. Then, the process proceeds to S308. On the other hand, if an explanation by voice is not being provided (No in S307), the process outputs information (second information) that a human body is not providing an utterance to the elapsed time measurement unit C1017. Then, the process proceeds to S311.
[0142] Next, in S308, the keyword extraction unit C1019 extracts keywords from the voice information input from the utterance determination unit B 1016. Then, the keyword extraction unit C1019 outputs the extracted keywords to the match determination unit C 1020. After that, the process proceeds to S309.
[0143] Next, in S309, the keyword extraction unit C1019 extracts keywords from the explanatory materials input from the material acquisition unit A1005. The keyword extraction unit C1019 outputs the extracted keywords to the match determination unit C1020. Thereafter, the process proceeds to S310. Note that the order of the processes of S308 and S309 may be reversed.
[0144] Next, in S310, the match determination unit C1020 determines whether the audio information input from the keyword extraction unit C1019 and the keywords extracted from the explanatory materials match or are similar. If the determination result shows that the audio information and the keywords match or are similar (Yes in S310), the match determination unit C1020 outputs information (third information) that matches the explanatory content to the explanatory area identification unit D1021. Then, the process proceeds to S313. On the other hand, if the audio information and the keywords do not match or are not similar (No in S310), the match determination unit C1020 outputs information (fourth information) that does not match the explanatory content to the image superimposition unit D1011. Then, the process proceeds to S318.
[0145] Next, in S311, the elapsed time measurement unit C1017 measures the elapsed time since the end of the utterance whose content does not match the explanatory material, based on the second information input from the utterance determination unit B1016 and the fourth information input from the match determination unit C1020. The elapsed time measurement unit C1017 outputs the measured time (measured time) to the determination unit B1018. Then, the process proceeds to S312.
[0146] Next, in S312, the determination unit B1018 determines whether a certain period of time has passed since the utterance inconsistent with the explanatory material ended or the utterance inconsistent with the explanatory material ended, based on the measured time input from the elapsed time measurement unit C1017. If the determination result shows that the utterance inconsistent with the explanatory material ended or the certain period of time has passed (Yes in S312), the determination unit B1018 outputs information to the image superimposition unit D1011 not to make the explanatory material transparent. Then, the process proceeds to S318. On the other hand, if the utterance inconsistent with the explanatory material has not ended or the certain period of time has not passed (No in S312), the determination unit B1018 outputs information to make the explanatory material transparent to the overlap region extraction unit A1009. Then, the process proceeds to S313.
[0147] Next, in S313, the explanation area identification unit D1021 identifies an area where the human body is explaining (explanation area). When identifying an area where the human body is explaining, the information including the human body skeleton estimation result, video information, explanatory material, and information that matches the explanation content (third information) is used. The human body skeleton estimation result and video information are input from the human body movement determination unit A1007. The area information is input from the area division processing unit A1008. The explanatory material is input from the material acquisition unit A1005. The information that matches the explanation content (third information) is input from the match determination unit C1020. Then, the explanation area identification unit D1021 outputs the area information, the identified explanation area, explanatory material, and video information to the overlap determination unit D1022. After that, the process proceeds to S314.
[0148] Next, in S314, the overlap determination unit D1022 determines whether the human body region and the explanatory region overlap based on the region information input from the explanatory region identification unit D1021, the identified explanatory region, explanatory material, and video information. If the human body region and the explanatory region overlap (Yes in S314), the overlap determination unit D1022 outputs the identified explanatory region, video information, and explanatory material to the highlight frame superimposition unit D1023. Then, the process proceeds to S315. On the other hand, if the human body region and the explanatory region do not overlap (No in S314), the overlap determination unit D1022 outputs the video information, explanatory material, and region information to the overlap region extraction unit A1009. Then, the process proceeds to S316.
[0149] Next, in S315, the highlight frame superimposition unit D1023 superimposes a highlight frame on the specified explanatory area of the explanatory material using the specified explanatory area, video information, and explanatory material input from the overlap determination unit D1022. The highlight frame superimposition unit D1023 outputs the superimposed explanatory material and video information to the image superimposition unit D1011. Then, the process proceeds to S318.
[0150] Next, in S316, the overlap region extraction unit A1009 extracts an overlap region using the human body region information input from the overlap determination unit D1022 and the explanatory materials input from the material acquisition unit A1005. The overlap region extraction unit A1009 then outputs the extracted overlap region to the transparency change unit A1010. After that, the process proceeds to S317.
[0151] Next, in S317, the transparency change unit A1010 changes the transparency of the explanatory material (makes it transparent) using the explanatory material and the overlapping area input from the overlapping area extraction unit A1009. Then, the transparency change unit A1010 outputs the explanatory material with the changed transparency to the image superimposition unit B1011. After that, the process proceeds to S318.
[0152] Next, in S318, the image superimposition unit D1011 superimposes explanatory material on the video information. Here, if explanatory material with changed transparency is input from the transparency change unit A1010, the image superimposition unit D1011 superimposes the explanatory material with changed transparency on the video information. Furthermore, if information not changing the transparency is input from the determination unit B1018, the image superimposition unit D1011 superimposes the explanatory material with unchanged transparency input from the material acquisition unit A1005 on the video information. Furthermore, if information indicating that the spoken content and the content of the explanatory material do not match (fourth information) is input from the match determination unit C1020, the image superimposition unit D1011 also superimposes the explanatory material with unchanged transparency input from the material acquisition unit A1005 on the video information. Furthermore, if an explanatory area with a superimposed highlight frame is input from the highlight frame superimposition unit D1023, the image superimposition unit D1011 superimposes the explanatory material with the superimposed highlight frame on the video information. The image superimposition unit D1011 then outputs the superimposed video to the video output device A1012. Then proceed to S319.
[0153] Next, in S319, the video output device A1012 outputs the video information (superimposed video) input from the image superimposition unit B1011 to the monitor device A1013. When the video information is input from the video output device A1012, the monitor device A1013 displays the video or image in the video information on the screen. Then, the process proceeds to S320.
[0154] Next, in S320, it is determined whether or not to end the process. Specifically, it is determined whether or not the user has operated the On / Off switch of the automatic photography system (not shown) to stop the automatic photography process. If the result of the determination is that the operation to stop the automatic photography process has not been performed (NO in S320), the process proceeds to S301 and the same process is repeated. On the other hand, if the operation to stop the automatic photography process has been performed (YES in S320), the automatic photography process is terminated and this processing flow is ended.
[0155] As described above, when the automatic photography system A1000 of the fourth embodiment superimposes explanatory material on a video of a human body, if the human body and the area of the explanation subject overlap when the human body is providing an explanation through voice or movement, a highlight frame is superimposed on the area of the explanation subject. On the other hand, if the human body and the area of the explanation subject do not overlap, the transparency of the area of the explanatory material that overlaps with the human body area is changed, and the explanatory material is superimposed on the video information. This makes it possible to confirm the area of the explanation subject on the screen even when the human body overlaps the area of the explanation subject.
[0156] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and various modifications and variations are possible within the spirit and scope of the present invention. Furthermore, while multiple exemplary embodiments have been described above, the present invention can also be embodied as, for example, a system, an apparatus, a method, a program, or a recording medium (storage medium). Specifically, the present invention may be applied to a system consisting of multiple devices (e.g., a host computer, an interface device, an imaging device, a web application, etc.), or to an apparatus consisting of a single device. Furthermore, for example, some or all of the functions of each functional unit shown in FIG. 1 may be included in a device other than the image superimposition device A1003. Specifically, these functional units may be included in a device or storage device other than the image superimposition device A1003, and the functions of each embodiment may be realized by communicating with the image superimposition device A1003 via a wired or wireless connection. Examples of such a device include the video capture device A1001, the document capture device A1002, an information processing device (not shown), or a server (not shown).
[0157] 1 may be realized by one or more computers different from the image superimposition device A1003. Furthermore, the image superimposition device A1003 may have functions equivalent to those of the video acquisition device A1001, the material acquisition device A1002, and the monitor device A1013. In this case, for example, the image superimposition device A1003 may be configured to acquire images and generate videos from the acquired images. Furthermore, for example, the image superimposition device A1003 may be configured to acquire explanatory materials. Furthermore, for example, the image superimposition device A1003 may be configured to display videos or images such as superimposed videos or superimposed images.
[0158] Also, some or all of the functions of the functional units in Fig. 1 may be provided by one or more devices other than the image superimposing device A1003, or the image superimposing device A1003 may have all of the functions in Fig. 1. The same applies to Figs. 11, 15, and 19. That is, the same can be applied to the image superimposing devices B1003, C1003, and D1003.
[0159] The disclosure of this embodiment includes the following configuration, method, and program.
[0160] (Configuration 1) a first extraction means for extracting a region of a human body from an image; a superimposing means for superimposing predetermined superimposing information on the image; a first determination means for determining whether the human body is performing a predetermined motion and outputting a determination result; a second extraction means for extracting, as an overlapping region, a region of the predetermined superimposition information that overlaps with the region of the human body based on the determination result, the region of the human body, and the predetermined superimposition information; a transparency changing unit that changes the transparency of at least a part of the predetermined superimposition information to be higher in accordance with the overlapping area.
[0161] (Configuration 2) The processing device according to configuration 1, characterized in that the superimposing means, when the determination result is that the human body is not performing the predetermined movement, superimposes the predetermined superimposing information with the transparency of the image unchanged.
[0162] (Configuration 3) The processing device according to configuration 1 or 2, wherein the second extraction means extracts the overlapping area from the image when the determination result indicates that the human body is performing the predetermined movement.
[0163] (Configuration 4) 4. The processing device according to any one of configurations 1 to 3, wherein the transparency changing means changes the transparency of the overlapping region.
[0164] (Configuration 5) an estimation means for detecting a human body from the image, estimating a skeleton of the detected human body, and outputting an estimation result; The processing device according to any one of configurations 1 to 4, wherein the first determination means determines whether or not the human body is performing the predetermined movement based on the estimation result estimated by the estimation means.
[0165] (Configuration 6) 6. The processing device according to configuration 5, wherein the second extraction means extracts a part of the human body as the overlapping region when extracting the overlapping region.
[0166] (Configuration 7) 7. The processing device according to any one of configurations 1 to 6, wherein the predetermined motion is at least an arm movement.
[0167] (Configuration 8) 8. The processing device according to any one of configurations 1 to 7, further comprising display control means for displaying the image or video superimposed by the superimposing means on a screen of a display device.
[0168] (Configuration 9) a second determination means for determining whether the human body is speaking based on voice information; The processing device according to any one of configurations 1 to 8, wherein the second determination means, when determining that the human body is speaking, determines that the human body is giving an explanation by voice.
[0169] (Configuration 10) The processing device according to configuration 9, wherein the second determination means outputs first information when it determines that the human body is speaking, and outputs second information when it determines that the human body is not speaking.
[0170] (Configuration 11) 11. The processing device according to configuration 10, further comprising a measuring means for measuring the time since the human body finished speaking based on the second information and outputting the measured time.
[0171] (Configuration 12) The processing device according to configuration 11, wherein the transparency changing means continues to change the transparency of the predetermined superimposition information based on the measured time until a predetermined time has elapsed since the human body stopped speaking.
[0172] (Configuration 13) a third extraction means for extracting a predetermined keyword from each of the voice information and the predetermined superimposition information; The processing device according to configuration 11 or 12, characterized in that the third extraction means has a third determination means for determining whether the keywords extracted from the audio information and the predetermined superimposition information respectively match or are similar.
[0173] (Configuration 14) The processing device described in configuration 13, wherein the third determination means outputs third information when it determines that the keywords extracted from the audio information and the predetermined superimposed information match or are similar, and outputs fourth information when it determines that the keywords extracted from the audio information and the predetermined superimposed information do not match or are similar.
[0174] (Configuration 15) 15. The processing device according to configuration 14, wherein the measuring means measures the time since the human body finished speaking based on the second information and the fourth information, and outputs the measured time.
[0175] (Configuration 16) 16. The processing device according to any one of configurations 1 to 15, further comprising: an identification means for identifying an explanation area, which is an area where the human body is explaining using the predetermined movement or voice, from the predetermined superimposition information.
[0176] (Configuration 17) 17. The processing device according to configuration 16, further comprising highlighting means for highlighting the description area.
[0177] (Configuration 18) A method for controlling a processing device, comprising: a first extraction step of extracting a region of a human body in an image; a superimposing step of superimposing predetermined superimposing information on the image; a determination step of determining whether the human body is performing a predetermined motion and outputting a determination result; a second extraction step of extracting, as an overlapping region, a region of the predetermined superimposition information that overlaps with the region of the human body based on the determination result, the region of the human body, and the predetermined superimposition information; and changing the transparency of at least a part of the predetermined superimposition information to increase it in accordance with the overlapping area.
[0178] (Configuration 19) A program for causing a computer to execute a control method for a processing device, the program including: a first extraction step of extracting a region of a human body in an image; a superimposing step of superimposing predetermined superimposing information on the image; a determination step of determining whether the human body is performing a predetermined motion and outputting a determination result; a second extraction step of extracting, as an overlapping region, a region of the predetermined superimposition information that overlaps with the region of the human body based on the determination result, the region of the human body, and the predetermined superimposition information; and a modifying step of modifying the transparency of at least a part of the predetermined superimposition information to increase the transparency in accordance with the overlapping area.
[0179] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0180] A1000 Automatic Photography System A1001 Image acquisition device A1002 Data acquisition device A1003 Image superimposition device A1004 Video acquisition unit A1005 Material acquisition department A1006 Skeleton information estimation unit A1007 Human body motion determination section A1008 Area division processing unit A1009 Overlapping area extraction part A1010 Transparency change section A1011 Image superimposition unit A1012 Video output device A1013 Monitor device
Claims
1. a first extraction means for extracting a region of a human body from an image; a superimposing means for superimposing predetermined superimposing information on the image; a first determination means for determining whether the human body is performing a predetermined movement; a second extraction means for extracting an overlapping area that is an area that overlaps with the area of the human body in the area of the predetermined superimposition information; a transparency change means for changing the transparency of at least the overlapping region to be higher when the first determination means determines that the human body is performing the predetermined movement, the transparency changing means does not change the transparency in the overlapping region when the first determining means determines that the human body is not performing the predetermined movement. A processing device characterized by:
2. 2. The processing device according to claim 1, wherein the second extraction means extracts the overlapping area from the image when the first determination means determines that the human body is performing the predetermined movement.
3. an estimation means for detecting a human body from the image, estimating a skeleton of the detected human body, and outputting an estimation result; 2. The processing device according to claim 1, wherein the first determining means determines whether or not the human body is performing the predetermined movement based on the estimation result obtained by the estimating means.
4. 2. The processing device according to claim 1, wherein the second extraction means extracts a region corresponding to a part of the human body as the overlapping region when extracting the overlapping region.
5. The processing device according to claim 1 , wherein the predetermined motion is at least an arm motion.
6. 2. The processing device according to claim 1, further comprising display control means for displaying the image or video superimposed by said superimposing means on a screen of a display device.
7. a second determination means for determining whether the human body is speaking based on voice information; 2. The processing device according to claim 1, wherein the second determination means, when determining that the human body is speaking, determines that the human body is giving an explanation by voice.
8. 8. The processing device according to claim 7, wherein the second determination means outputs first information when it determines that the human body is speaking, and outputs second information when it determines that the human body is not speaking.
9. 9. The processing device according to claim 8, further comprising: a measuring means for measuring the time elapsed since the human body finished speaking based on the second information, and outputting the measured time.
10. 10. The processing device according to claim 9, wherein the transparency changing means continues to change the transparency of the predetermined superimposition information based on the measured time until a predetermined time has elapsed since the human body stopped speaking.
11. a third extraction means for extracting a predetermined keyword from each of the voice information and the predetermined superimposed information; 10. The processing device according to claim 9, further comprising a third determination means for determining whether the keywords extracted by the third extraction means from the audio information and the predetermined superimposed information respectively match or are similar.
12. The processing device according to claim 11, characterized in that the third determination means outputs third information when it determines that the keywords extracted from the voice information and the specified superimposition information match or are similar, and outputs fourth information when it determines that the keywords extracted from the voice information and the specified superimposition information do not match or are similar.
13. 13. The processing device according to claim 12, wherein the measuring means measures the time from when the human body finishes speaking based on the second information and the fourth information, and outputs the measured time.
14. 2. The processing device according to claim 1, further comprising: a specifying unit for specifying, from the predetermined superimposition information, an explanation area in which the human body is explaining by the predetermined motion or voice.
15. 15. The processing device according to claim 14, further comprising highlighting means for highlighting the description area.
16. A method for controlling a processing device, comprising: a first extraction step of extracting a region of a human body in an image; a superimposing step of superimposing predetermined superimposing information on the image; a determining step of determining whether the human body is performing a predetermined movement; a second extraction step of extracting an overlapping region that is a region that overlaps with the region of the human body in the region of the predetermined superimposition information; a change step of increasing transparency in at least the overlapping region when it is determined in the determination step that the human body is performing the predetermined movement, In the changing step, if it is determined in the determining step that the human body is not performing the predetermined movement, the transparency in the overlapping region is not changed. A method for controlling a processing apparatus comprising:
17. A program for causing a computer to execute a control method for a processing device, the program including: a first extraction step of extracting a region of a human body in an image; a superimposing step of superimposing predetermined superimposing information on the image; a determining step of determining whether the human body is performing a predetermined movement; a second extraction step of extracting an overlapping region that is a region that overlaps with the region of the human body in the region of the predetermined superimposition information; a change step of increasing transparency in at least the overlapping region when it is determined in the determination step that the human body is performing the predetermined movement; In the changing step, if it is determined in the determining step that the human body is not performing the predetermined movement, the transparency in the overlapping region is not changed. A program characterized by:
Citation Information
Patent Citations
Manufacture of coner tile
JP1985046961A
Image display method, program, image display apparatus, and image display system
JP2005287004A
Image display and image display method therefor
JP2010039125A
Explanation support device, explanation support method, and explanation support program
JP2016194877A
Moving image output device, moving image output method, moving image output program, and moving image output system
JP2020155961A