Image analysis device, meeting support system, image analysis method, and program
The image analysis device captures and records gaze direction to identify and extract discussion objects, addressing the challenge of sharing focused discussion points in meetings, enhancing information accessibility for all participants.
Patent Information
- Application Number
- JP2021197171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-12-03
AI Technical Summary
Existing technologies fail to effectively share information about focused discussion points in meetings, especially for non-participants, as they struggle to accurately identify and record which sticky notes or on-screen content received attention during discussions.
An image analysis device that captures and analyzes gaze direction to identify objects of discussion, extracts relevant image areas, and records this data for later review, allowing participants and non-participants to understand the discussion focus.
Enables anyone to easily know what was highlighted during the discussion by capturing and recording the objects participants were gazing at, facilitating better information sharing across meeting participants.
Smart Images

Figure 0007771688000001 
Figure 0007771688000002 
Figure 0007771688000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image analysis device, a conference support system, an image analysis method, and a program, and more particularly to an image analysis device, a conference support system, an image analysis method, and a program that analyze the status of conference participants by performing image analysis on a second captured image acquired by an imaging device. [Background technology]
[0002] Various related technologies have been proposed to help stimulate discussions in meetings. For example, an information processing device described in Patent Document 1 uses a camera, a microphone, and various sensors to identify the status of participants in a meeting, such as their positions, voice volume, hand movements, and activity levels, which are indicators of the level of activity in the discussion. If the information processing device described in Patent Document 1 determines that the level of activity in the discussion is low and that the discussion is stagnating, it presents information to promote the discussion (for example, hints for ideas).
[0003] In brainstorming and the KJ method, meeting participants freely express their ideas and opinions, and then post their ideas and opinions on sticky notes on a whiteboard or poster. The presenter may also explain the contents of the meeting materials in detail by drawing diagrams and supplementary explanations on a screen on which the data of the meeting materials is projected.
[0004] In online meetings, participants in remote locations cannot directly see sticky notes or drawings, making it difficult to accurately understand the discussion in the meeting room. Related technologies have been developed to solve this problem.
[0005] For example, Patent Document 2 describes a method in which additional information is generated by capturing images of characters, figures, etc. drawn on a screen by a presenter, and the generated additional information is stored in a recording medium in association with data on the conference materials projected on the screen. After the conference ends, conference participants and non-participants (such as absentees) can review the discussions that took place in the conference by checking the additional information and data on the conference materials stored in the recording medium.
[0006] Patent document 3 also describes a method of identifying sticky notes attached to the screen from a second captured image of the screen, generating a partial image including the image area of the identified sticky note, and saving the data of the partial image in a storage device. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] International Publication No. 2020 / 070733 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-184333 [Patent Document 3] Japanese Patent Application Laid-Open No. 2014-186823 Summary of the Invention [Problem to be solved by the invention]
[0008] In the related technology described in Patent Document 2, even meeting participants, let alone non-participants, may be unable to remember what explanation was attached to the characters or figures drawn on the screen. Furthermore, in the related technology described in Patent Document 3, it may be unclear which sticky notes contained ideas or opinions that received particular attention during the discussion. As a result, it is difficult to share information such as ideas and opinions that received particular attention from participants in a meeting room with people who were not present at the actual discussion.
[0009] The present invention has been made in view of the above-mentioned problems, and its purpose is to enable anyone to easily know what was focused on during the course of a discussion. [Means for solving the problem]
[0010] An image analysis device according to one embodiment of the present invention comprises an acquisition means for acquiring a first captured image and a second captured image taken by an imaging device, an estimation means for estimating the gaze direction of the meeting participants appearing in the first captured image, an identification means for identifying the object of discussion that the meeting participants are gazing at in the second captured image, an extraction means for extracting an image area of the identified object of discussion in the second captured image, and a recording means for recording data of the image area of the object of discussion.
[0011] A conference support system according to one embodiment of the present invention comprises an image analysis device including an acquisition means for acquiring a first captured image and a second captured image taken by a photographing device, an estimation means for estimating the gaze direction of the conference participants appearing in the first captured image, an identification means for identifying the object of discussion that the conference participants are gazing at in the second captured image, an extraction means for extracting an image area of the identified object of discussion in the second captured image, and a recording means for recording data of the image area of the object of discussion; a photographing device for transmitting the second captured image to the image analysis device; and a storage device in which data of the image area of the object of discussion is recorded.
[0012] An image analysis method according to one embodiment of the present invention includes an estimation means for acquiring a first captured image and a second captured image by a photographing device, estimating the gaze direction of the meeting participants appearing in the first captured image, identifying the object of discussion that the meeting participants are gazing at in the second captured image, extracting an image area of the identified object of discussion in the second captured image, and recording data of the image area of the object of discussion.
[0013] A program according to one embodiment of the present invention causes a computer to acquire a first captured image and a second captured image taken by a photographing device, estimate the direction in which the meeting participants in the first captured image are looking, identify in the second captured image the object of discussion that the meeting participants are looking at, extract an image area of the identified object of discussion in the second captured image, and record data of the image area of the object of discussion. [Effects of the Invention]
[0014] One aspect of the present invention is to allow anyone to easily see what was highlighted during the course of a discussion. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram schematically illustrating an example of the configuration of an image analyzing device including the image analyzing device according to any one of the first to third embodiments. [Figure 2] 1 is a block diagram showing the configuration of an image analysis device according to a first embodiment. [Figure 3] 3 is a flowchart showing the operation of the image analyzing device according to the first embodiment. [Figure 4] 3 is a diagram showing an example of the data structure of image region data recorded in a storage device by a recording unit of the image analysis device according to the first embodiment. FIG. [Figure 5] FIG. 10 is a block diagram showing the configuration of an image analysis device according to a second embodiment. [Figure 6] 10 is a flowchart showing the operation of the image analyzing device according to the second embodiment. [Figure 7] 10 is a flowchart showing the operation of a sound recording unit included in the image analyzing device according to the second embodiment. [Figure 8] FIG. 10 is a diagram showing an example of information indicating an audio file created by a sound recording unit of an image analyzing device according to a second embodiment. [Figure 9] FIG. 10 is a block diagram showing the configuration of an image analysis device according to a third embodiment. [Figure 10]10 is a flowchart showing the operation of the image analyzing device according to the third embodiment. [Figure 11] 11 is a flowchart showing the operation of a generating unit included in the image analyzing device according to the third embodiment. [Figure 12] FIG. 10 is a diagram schematically illustrating an example of a virtual space generated by a generating unit included in the image analyzing device according to the third embodiment. [Figure 13] FIG. 1 is a diagram illustrating an example of a hardware configuration of an image analyzing device according to any one of the first to third embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0016] (Meeting Support System 1) A conference support system 1 that supports online conferences will be described with reference to FIG. 1. FIG. 1 is a diagram that schematically illustrates an example of the configuration of the conference support system 1. In an online conference, some or all of the conference participants enter a virtual conference room from remote locations via a network. Here, a remote location refers to any location different from the conference room. For example, a remote location could be a building separate from the building where the conference room is located, a satellite office, a coworking space, or a participant's home.
[0017] In an online conference, application software that can be installed on the user terminal 100, such as software for a Web conference system, is used.
[0018] 1, the conference support system 1 includes an image analyzing device 10 (20, 30) according to embodiments 1 to 3 described below. Here, the "image analyzing device 10 (20, 30)" means any one of the image analyzing devices 10, 20, and 30.
[0019] The conference support system 1 also includes a user terminal 100 used by participants in the conference room and a user terminal 100 used by participants in remote locations. The conference support system 1 further includes an image capturing device 200 and a storage device 300 installed in the conference room.
[0020] The user terminal 100 and the photographing device 200 that constitute the conference support system 1 are communicably connected to the image analyzing device 10 (20, 30) via a network. The storage device 300 is accessible from at least the image analyzing device 10 (20, 30). The network may be a local area network or the Internet.
[0021] The user terminal 100 is a communication device used by the participants of the conference. The user terminal 100 is, for example, a mobile phone, a smart device, or a personal computer. The user terminal 100 is equipped with a microphone, a camera, a speaker, and a display. The image capturing device 200 is installed in the conference room. The image capturing device 200 is, for example, a surveillance camera for bird's-eye view photography or a 360-degree camera.
[0022] The storage device 300 stores data of captured images generated by the imaging device 200. The storage device 300 also stores data and information generated by an image analysis device 10 (20, 30) or an AR (Augmented Reality) processing device (embodiment 3) described below. The storage device 300 is, for example, a network server. The image analysis device 10 (20, 30) is configured to be able to access the storage device 300.
[0023] [Embodiment 1] The first embodiment will be described with reference to FIGS.
[0024] (Image analysis device 10) Fig. 2 is a block diagram showing the configuration of the image analysis device 10 according to the present embodiment 1. As shown in Fig. 2, the image analysis device 10 includes an acquisition unit 11, an estimation unit 12, an identification unit 13, an extraction unit 14, and a recording unit 15.
[0025] The acquisition unit 11 acquires an image (hereinafter referred to as a captured image) captured by the image capturing device 200. The acquisition unit 11 is an example of an acquisition means.
[0026] In one example, the acquisition unit 11 acquires data of the captured image via a network (e.g., the Internet or a local area network) from the image capturing device 200 (FIG. 1) of the conference support system 1. Alternatively, the acquisition unit 11 may acquire data of the captured image recorded in the storage device 300.
[0027] In yet another example, the acquisition unit 11 remotely controls the operation of the photographing device 200 via a network, thereby causing the photographing device 200 to perform photographing. Then, the acquisition unit 11 acquires the photographed image obtained by photographing from the photographing device 200.
[0028] The acquisition unit 11 outputs data of the acquired first captured image to the estimation unit 12. The first captured image shows the faces of the conference participants (an example of "people" in the following description). The acquisition unit 11 also outputs data of the acquired second captured image to the identification unit 13. The second captured image shows the entire interior of the conference room. Here, the entire interior includes the inner walls of the conference room, objects in the conference room, and people in the conference room. The first captured image and the second captured image may be captured by the same image capture device 200, or may be captured by different image capture devices 200.
[0029] The estimation unit 12 estimates the gaze direction of a person appearing in the first captured image. The estimation unit 12 is an example of an estimation means.
[0030] In one example, the estimation unit 12 receives data of a first captured image from the acquisition unit 11. The estimation unit 12 detects a person's face from the received data of the first captured image using a face detection related technology. The estimation unit 12 estimates the orientation of the detected person's face.
[0031] Alternatively, the estimation unit 12 detects the eye region of the detected person. Then, the estimation unit 12 estimates the gaze of the person based on the deviation of the pupil in the eye region of the person. The estimation unit 12 identifies the gaze direction of the person based on the estimated face direction of the person or the estimated gaze direction of the person.
[0032] The estimation unit 12 outputs information indicating the direction in which the person is gazing to the identification unit 13.
[0033] The identification unit 13 identifies the object of discussion that the person is gazing at in the second captured image. The identification unit 13 is an example of an identification means.
[0034] In one example, the identification unit 13 receives data of the second captured image from the acquisition unit 11. The identification unit 13 also receives information from the estimation unit 12 indicating the direction in which the person is gazing.
[0035] The identification unit 13 detects an object in the direction of the person's face or the direction of the person's line of sight, that is, an object that the person is gazing at, in the second captured image. The object is the subject of discussion. The object is, for example, a wall of a conference room, a display, a whiteboard, a poster, or a sticky note attached to a wall or whiteboard.
[0036] The identification unit 13 can detect an object that the person is gazing at from the second captured image using a technique related to object detection in the technical field of image analysis (for example, edge detection or contrast analysis).
[0037] Note that if the first captured image and the second captured image are captured by different image capturing devices 200, the identification unit 13 needs to detect, in the second captured image, the person detected in the first captured image by the estimation unit 12. For this purpose, the identification unit 13 receives, from the estimation unit 12, information indicating the gaze direction of the person as well as the person's identification information. For example, the person's identification information includes the person's facial features and / or the person's position coordinates in the absolute coordinate system.
[0038] Furthermore, except for the case where object identification information (e.g., object ID) has already been issued for the object that the person is gazing at, the identification unit 13 issues identification information for identifying the object that the person is gazing at. Thereafter, the identification unit 13 associates the information for identifying the object that the person is gazing at with the object identification information and records it in the storage device 300 or the like.
[0039] The identification unit 13 outputs information for identifying the object that the person is gazing at, together with the data of the second captured image, to the extraction unit 14. The information for identifying the object that the person is gazing at includes identification information of the object (e.g., object ID), information indicating the time when the object was gazed at, and information indicating the position coordinates (range) of the object in the second captured image (FIG. 4).
[0040] The identification unit 13 may detect only a specific type of object. In this case, the identification unit 13 may detect a specific object in advance that appears in the second captured image before executing the process of identifying the object that the person is gazing at. More specifically, the specific type of object is an object that is the subject of discussion, and specific examples of such an object include a whiteboard, a display, and a sticky note.
[0041] For example, the identification unit 13 can use a classifier that has learned the characteristics of a specific type of object. The identification unit 13 uses the trained classifier to extract an image area of the specific type of object from the second captured image. Thereafter, the identification unit 13 determines whether or not an image area of the specific type of object exists in the direction of gaze of the person in the second captured image.
[0042] As described above, the identification unit 13 can detect a specific type of object in the direction of a person's gaze. An example in which the specific type of object is a "sticky note" will be described in the third embodiment.
[0043] The extraction unit 14 extracts an image area of the identified object from the second captured image. The extraction unit 14 is an example of an extraction means.
[0044] In one example, the extraction unit 14 receives data of the second captured image and information for identifying the object that the person is gazing at from the identification unit 13. The extraction unit 14 extracts an image area of the object that the person is gazing at from the second captured image, using the information for identifying the object that the person is gazing at.
[0045] For example, the extraction unit 14 extracts information indicating the time when the object was gazed at and information indicating the position coordinates (range) of the object in the second captured image from the information for identifying the object gazed at by the person. Here, the "position coordinates (range)" means a specific point on the object or an area occupied by the outline of the object.
[0046] The extraction unit 14 extracts an image area specified by information indicating the position coordinates (range) of the object from the second captured image captured during the time when the object was gazed at.
[0047] In this way, the extraction unit 14 extracts, from the second captured image, the image region of the object identified by the identification unit 13, that is, the image region of the object that the person is gazing at.
[0048] Furthermore, the extraction unit 14 associates the data of the image area of the object that the person is gazing at with the object identification information (for example, object ID) received from the identification unit 13, and records it in the storage device 300 or the like (FIG. 4).
[0049] The extraction unit 14 outputs data of the image area of the object that the person is gazing at to the recording unit 15.
[0050] The recording unit 15 records data of the image area of the object. The recording unit 15 is an example of a recording means.
[0051] In one example, the recording unit 15 receives data on an image area of an object that a person is gazing at from the extraction unit 14. The recording unit 15 records the data on the image area of the object in the storage device 300 (FIG. 1).
[0052] (Operation of image analysis device 10) The operation of the image analyzing device 10 according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of processing executed by each unit of the image analyzing device 10.
[0053] As shown in FIG. 3, first, the acquisition unit 11 acquires a first captured image and a second captured image captured by the image capturing device 200 (FIG. 1) (S101).
[0054] The acquisition unit 11 outputs the acquired data of the first photographed image to the estimation unit 12. The acquisition unit 11 also outputs the acquired data of the second photographed image to the identification unit 13.
[0055] The estimation unit 12 receives data of the first captured image acquired by the acquisition unit 11. The estimation unit 12 estimates the gaze direction of the person appearing in the first captured image (S102).
[0056] The estimation unit 12 outputs information indicating the direction in which the person is gazing to the identification unit 13.
[0057] The identification unit 13 receives the data of the second captured image acquired by the acquisition unit 11. The identification unit 13 identifies an object that the person is gazing at in the second captured image (S103).
[0058] The identification unit 13 outputs information for identifying the object that the person is gazing at to the extraction unit 14 together with the data of the second captured image.
[0059] The extraction unit 14 receives data of the second captured image and information for identifying the object that the person is gazing at from the identification unit 13. The extraction unit 14 extracts an image area of the identified object in the second captured image (S104).
[0060] The extraction unit 14 outputs data of the image area of the object that the person is gazing at to the recording unit 15.
[0061] The recording unit 15 receives data on the image area of the object that the person is gazing at from the extraction unit 14. The recording unit 15 records the data on the image area of the object (S105).
[0062] This completes the operation of the image analyzing device 10 according to the first embodiment.
[0063] (An example of information for identifying an object a person is gazing at) 4 shows an example of information for identifying an object that a person is gazing at. The information for identifying an object that a person is gazing at is recorded in the storage device 300 or the like by the identification unit 13 and extraction unit 14 of the image analysis device 10 described above.
[0064] Specifically, the identification unit 13 associates information indicating the time the object was gazed at and information indicating the position coordinates (range) of the object in the second captured image with the object's identification information (e.g., object ID) and records it in a storage device 300, etc.
[0065] 4, "20:00:00" and "20:10:00" represent information indicating the time when the object was gazed at, while "AA:AA:AA" and "BB:BB:BB" represent the position coordinates (range) of the object in the second captured image.
[0066] Furthermore, the extraction unit 14 associates data of the image area of the object that the person is gazing at with identification information (for example, object ID) of the object, and records the data in the storage device 300, etc. In Fig. 4, "aaaaa" and "bbbbb" represent data of the image area of the object that the person is gazing at.
[0067] (Effects of this embodiment) According to the configuration of this embodiment, the acquisition unit 11 acquires a first captured image and a second captured image captured by a photographing device. The estimation unit 12 estimates the gaze direction of a person appearing in the first captured image. The identification unit 13 identifies the object of discussion that the person is gazing at in the second captured image. The extraction unit 14 extracts the image area of the identified object in the second captured image. The recording unit 15 records the data of the image area of the object. Not only meeting participants but also non-participants can learn the object that the person was gazing at during the meeting by referring to the data of the image area of the object. This makes it possible for anyone to easily know what was being focused on during the discussion.
[0068] [Embodiment 2] A second embodiment will be described with reference to Figures 5 to 8. In the second embodiment, a configuration will be described in which discussions, conversations, or presentations between meeting participants are recorded to create audio files, and the created audio files are associated with data on image areas of objects that the participants are gazing at during the meeting.
[0069] In the second embodiment, the same components as those in the first embodiment are denoted by the same reference numerals, and the overlapping description with the first embodiment will be omitted.
[0070] (Image analysis device 20) Fig. 5 is a block diagram showing the configuration of an image analysis device 20 according to embodiment 2. As shown in Fig. 5, the image analysis device 20 includes an acquisition unit 11, an estimation unit 12, an identification unit 13, an extraction unit 14, and a recording unit 15. The image analysis device 20 further includes a sound recording unit 25.
[0071] The recording unit 25 records the discussion between the person and other participants and creates an audio file. The recording unit 25 is an example of a recording means.
[0072] In one example, the recording unit 25 receives audio signals input to a microphone by the voices of conference participants from a user terminal 100 (Figure 1) in a conference room or a user terminal 100 (Figure 1) in a remote location.
[0073] The recording unit 25 creates an audio file, which is digital data, by A / D converting the received audio signal. The recording unit 25 outputs the created audio file to the recording unit 15.
[0074] Furthermore, the recording unit 25 receives information for identifying an object that the person is gazing at from the identification unit 13. The information for identifying an object that the person is gazing at includes identification information of the object (for example, an object ID).
[0075] The recording unit 25 issues identification information (for example, an audio ID) for identifying the audio file. After that, the recording unit 25 associates the audio file with the identification information of the audio file and the identification information of the object, and records the audio file in the storage device 300 or the like (FIG. 8).
[0076] As described in the first embodiment, the identification unit 13 issues identification information (e.g., an object ID) for identifying the object that the person is gazing at, except when identification information has already been issued for the object that the person is gazing at. On the other hand, if the object that the person is gazing at cannot be identified, the identification unit 13 cannot issue identification information for the object. In this case, the recording unit 25 records the audio file in the storage device 300 or the like in association with only the identification information of the audio file (FIG. 8).
[0077] The recording unit 15 receives data on the image area of the object that the person is gazing at from the extraction unit 14. The recording unit 15 also receives the audio file that the audio recording unit 25 has created.
[0078] Recording unit 15 records the data of the image area of the object in storage device 300 (FIG. 1). Recording unit 15 may add information for specifying the time when the object was gazed at to the data of the image area of the object.
[0079] Furthermore, the recording unit 15 may record the audio file and the data of the image area of the object that the person is gazing at in association with each other in the storage device 300 or the like.
[0080] (Operation of image analysis device 20) The operation of the image analyzing device 20 according to the second embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the flow of processing executed by each unit of the image analyzing device 20.
[0081] As shown in FIG. 6, first, the acquisition unit 11 acquires a second captured image captured by the image capturing device 200 (FIG. 1) (S201).
[0082] The acquisition unit 11 outputs the acquired data of the first photographed image to the estimation unit 12. The acquisition unit 11 also outputs the acquired data of the second photographed image to the identification unit 13.
[0083] The estimation unit 12 receives data of the first captured image acquired by the acquisition unit 11. The estimation unit 12 estimates the gaze direction of the person appearing in the first captured image (S202).
[0084] The estimation unit 12 outputs information indicating the direction in which the person is gazing to the identification unit 13.
[0085] The identification unit 13 receives the data of the second captured image acquired by the acquisition unit 11. The identification unit 13 identifies an object that the person is gazing at in the second captured image (S203).
[0086] The identification unit 13 outputs information for identifying the object that the person is gazing at, together with the data of the second captured image, to the extraction unit 14. The identification unit 13 also outputs information for identifying the object that the person is gazing at, at least the identification information of the object (for example, object ID), to the sound recording unit 25.
[0087] The recording unit 25 records the discussion between the person and the other participants and creates an audio file (S203). Details of step S203 will be described later with reference to another flowchart (FIG. 7).
[0088] The recording unit 25 receives information for identifying an object that the person is gazing at from the identification unit 13. The recording unit 25 extracts identification information of the object (for example, an object ID) from the information for identifying the object that the person is gazing at. Then, the recording unit 25 records the audio file in the storage device 300 or the like in association with the identification information of the object.
[0089] The extraction unit 14 receives data of the second captured image and information for identifying the object that the person is gazing at from the identification unit 13. The extraction unit 14 extracts an image area of the identified object in the second captured image (S205).
[0090] The extraction unit 14 outputs data of the image area of the object that the person is gazing at to the recording unit 15.
[0091] The recording unit 15 receives data on the image area of the object that the person is gazing at from the extraction unit 14. The recording unit 15 records the data on the image area of the object (S206). At this time, the recording unit 15 may record the audio file and the data on the image area of the object that the person is gazing at in association with each other.
[0092] This completes the operation of the image analyzing device 20 according to the second embodiment.
[0093] (Details of step S203 executed by the recording unit 25) With reference to FIG. 7, the process executed by the recording unit 25 in step S203 of the above-mentioned flowchart (FIG. 6) will be described in detail.
[0094] As shown in FIG. 7, in step S203, recording unit 25 first creates an audio file by A / D converting an audio signal input to the microphone of user terminal 100 (FIG. 1) (S2031).
[0095] The recording unit 25 records the created audio file in the storage device 300 or the like in association with the identification information of the audio file and the identification information of the object (FIG. 8).
[0096] Next, the recording unit 25 determines whether or not a certain number of people or more are gazing at the same object (S2032).
[0097] For example, the recording unit 25 counts the number of pieces of object identification information associated with the recorded audio files. If the number of pieces of object identification information associated with the recorded audio files is equal to or greater than a certain number, the recording unit 25 determines that a certain number or more people are gazing at the same object.
[0098] If a certain number of people or more are not gazing at the same object (No in S2032), that is, if the number of people gazing at the same object is less than a certain number, this flow ends.
[0099] On the other hand, if a certain number of people or more are gazing at the same object (Yes in S2032), the recording unit 25 associates the identification information of the object for which recording has been stopped with the audio file (S2033).
[0100] Thereafter, the flow proceeds to step S305 in the flowchart (FIG. 6) described above.
[0101] (An example of an audio file associated with an object's identification information) Fig. 8 shows an example of an audio file associated with identification information (e.g., object ID) of an object that a person is gazing at. The object identification information is an index of information (Fig. 4) for specifying the object that a person is gazing at. As shown in Fig. 8, the audio file is associated with the identification information (e.g., audio ID) of the audio file and the identification information (e.g., object ID) of the object, and is recorded in the storage device 300 or the like.
[0102] (Effects of this embodiment) According to the configuration of this embodiment, the acquisition unit 11 acquires a first captured image and a second captured image captured by a photographing device. The estimation unit 12 estimates the gaze direction of a person appearing in the first captured image. The identification unit 13 identifies the object of discussion that the person is gazing at in the second captured image. The extraction unit 14 extracts the image area of the identified object in the second captured image. The recording unit 15 records the data of the image area of the object. Not only meeting participants but also non-participants can learn the object that the person was gazing at during the meeting by referring to the data of the image area of the object. This makes it possible for anyone to easily know what was being focused on during the discussion.
[0103] Furthermore, according to the configuration of this embodiment, the sound recording unit 25 records discussions, conversations, or presentations between a person and other participants to create an audio file. The recording unit 15 records the audio file in association with data on the image area of the object. This allows not only meeting participants but also non-participants to know what object the meeting participants were gazing at and what kind of discussion, conversation, or presentation they were having.
[0104] [Embodiment 3] A third embodiment will be described with reference to Figures 9 to 12. In this third embodiment, a configuration will be described that enables a remote participant to visually recognize an object that a participant is gazing at during a conference in a virtual space. The virtual space here encompasses the concepts of an augmented reality world, an augmented virtual world, and a mixed reality world.
[0105] In the third embodiment, the same components as those in the first or second embodiment are denoted by the same reference numerals, and the description thereof will be omitted.
[0106] (Image analysis device 30) Fig. 9 is a block diagram showing the configuration of an image analysis device 30 according to the third embodiment. As shown in Fig. 9, the image analysis device 30 includes an acquisition unit 11, an estimation unit 12, an identification unit 13, an extraction unit 14, and a recording unit 15. The image analysis device 30 further includes a generation unit 35.
[0107] After the image area of the object is extracted, the generation unit 35 pastes the data of the image area of the object onto a model of the virtual space, thereby generating data of the virtual space corresponding to the real space including the object. The generation unit 35 is an example of a generation means.
[0108] In one example, the generation unit 35 receives information (FIG. 4) for identifying an object that the person is gazing at from the identification unit 13. The generation unit 35 receives at least information indicating the position coordinates (range) of the object in the second captured image. The generation unit 35 also receives data on the image area of the object that the person is gazing at from the extraction unit 14.
[0109] The generation unit 35 generates data of the virtual space by attaching the data of the received image area to a model of the virtual space using information indicating the position coordinates (range) of the object in the second captured image.
[0110] In one example, the generation unit 35 arranges a plurality of image regions in the virtual space in the depth direction so that the plurality of image regions pasted on the model of the virtual space do not overlap each other when viewed from a specific viewpoint in the virtual space. The generation unit 35 may arrange the plurality of image regions in the virtual space in the depth direction when viewed from the specific viewpoint in the order of the time when the object was gazed at.
[0111] The generation unit 35 records the generated virtual space data in the storage device 300 or the like. In addition, the generation unit 35 outputs the virtual space data to which the data of the image area of the object that the person is gazing at has been attached to the recording unit 15.
[0112] The recording unit 15 receives the virtual space data to which the data of the image area of the object that the person is gazing at has been attached from the generation unit 35. The recording unit 15 may record the virtual space data generated by the generation unit 35 instead of the data of the image area of the object that the person is gazing at.
[0113] (Operation of image analysis device 30) The operation of the image analyzing device 30 according to the third embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the flow of processing executed by each unit of the image analyzing device 30.
[0114] As shown in FIG. 10, first, the acquisition unit 11 acquires a second captured image captured by the image capturing device 200 (FIG. 1) (S301).
[0115] The acquisition unit 11 outputs the acquired data of the first photographed image to the estimation unit 12. The acquisition unit 11 also outputs the acquired data of the second photographed image to the identification unit 13.
[0116] The estimation unit 12 receives data of the first captured image acquired by the acquisition unit 11. The estimation unit 12 estimates the gaze direction of the person appearing in the first captured image (S302).
[0117] The estimation unit 12 outputs information indicating the direction in which the person is gazing to the identification unit 13.
[0118] The identification unit 13 receives the data of the second captured image acquired by the acquisition unit 11. The identification unit 13 identifies an object that the person is gazing at in the second captured image (S303).
[0119] The identification unit 13 outputs information for identifying the object that the person is gazing at to the extraction unit 14 together with the data of the second captured image.
[0120] The extraction unit 14 receives data of the second captured image and information for identifying the object that the person is gazing at from the identification unit 13. The extraction unit 14 extracts an image area of the identified object in the second captured image (S304).
[0121] The extraction unit 14 outputs data of the image area of the object that the person is gazing at to the recording unit 15 and the generation unit 35.
[0122] After the image area of the object is extracted, the generation unit 35 pastes the data of the image area of the object onto a model of the virtual space, thereby generating data of the virtual space corresponding to the real space including the object (S305). Details of step S305 will be described later with reference to another flowchart (FIG. 11).
[0123] The generating unit 35 outputs the generated virtual space data to the recording unit 15.
[0124] The recording unit 15 receives data of the image area of the object that the person is gazing at from the extraction unit 14. The recording unit 15 also receives data of the virtual space from the generation unit 35. The recording unit 15 records the data of the image area of the object that the person is gazing at. Alternatively, the recording unit 15 records the data of the virtual space to which the data of the image area of the object that the person is gazing at has been attached (S306).
[0125] This completes the operation of the image analyzing device 30 according to the third embodiment.
[0126] (Details of S305: Processing Executed by the Generation Unit 35) The process executed by the generation unit 35 in step S305 of the above-mentioned flowchart (FIG. 10) will be described in detail with reference to Fig. 11. Here, an example will be taken in which the object that a person is gazing at is a "sticky note."
[0127] As shown in FIG. 11, first, the generating unit 35 receives data of an image area including a sticky note and information for identifying the sticky note that the person is gazing at (FIG. 4) from the identifying unit 13 (S3051).
[0128] Next, the generation unit 35 acquires data of a model of the virtual space from a recording medium such as the storage device 300 (S3052). Alternatively, the generation unit 35 may acquire data of a model of a general-purpose virtual space (for example, a cubic space). The position coordinates of each point in the virtual space correspond one-to-one to the position coordinates of each point in the second captured image. Information indicating the correspondence between the position coordinates is stored in advance in the storage device 300 or the like, or is held by the image analysis device 30.
[0129] Furthermore, the generation unit 35 extracts information indicating the position coordinates (range) of the sticky note from the information for identifying the sticky note that the person is gazing at (S3053).
[0130] Then, the generation unit 35 uses information indicating the position coordinates of the sticky note to attach data of the image area of the sticky note to the model of the virtual space (S3054). At this time, the generation unit 35 converts the position coordinates of the sticky note into position coordinates in the virtual space using information (for example, a function or a correspondence table) that associates the position coordinates of each point in the virtual space with the position coordinates of each point in the second captured image on a one-to-one basis.
[0131] Thereafter, the flow proceeds to step S306 in the flowchart (FIG. 10) described above.
[0132] (An example of a virtual conference room) Fig. 12 is a schematic diagram showing an example of a virtual space generated by the image analysis device generation unit 35 according to the third embodiment. The example shown in Fig. 12 shows a sticky note (an example of an object) in the virtual space. Arrows also indicate the line of sight of a meeting participant who is viewing the virtual space using AR (augmented reality) glasses or the like.
[0133] In the virtual conference room, the meeting participants can see 360 degrees. When the meeting participants move or change direction, an AR processing device (not shown) rotates the sticky notes in the virtual space so that the meeting participants can see the sticky notes.
[0134] The AR processing device may also enlarge or reduce the sticky notes in the virtual space. Furthermore, the AR processing device may move the sticky notes in the virtual space so that they do not overlap with each other as seen by the meeting participants. This can improve the visibility of the sticky notes (e.g., the readability of characters written on the sticky notes) as seen by the meeting participants viewing the virtual space.
[0135] The AR processing device may be part of the image analysis device 30. That is, in one modified example, the image analysis device 30 may have the function of the AR processing device.
[0136] (Effects of this embodiment) According to the configuration of this embodiment, the acquisition unit 11 acquires a first captured image and a second captured image captured by a photographing device. The estimation unit 12 estimates the gaze direction of a person appearing in the first captured image. The identification unit 13 identifies the object of discussion that the person is gazing at in the second captured image. The extraction unit 14 extracts the image area of the identified object in the second captured image. The recording unit 15 records the data of the image area of the object. Not only meeting participants but also non-participants can learn the object that the person was gazing at during the meeting by referring to the data of the image area of the object. This makes it possible for anyone to easily know what was being focused on during the discussion.
[0137] Furthermore, according to the configuration of this embodiment, after the image area of the object is extracted, the generation unit 35 pastes the data of the image area of the object onto a model of the virtual space, thereby generating virtual space data corresponding to the real space including the object. The recording unit 15 records the data of the image area of the object or the virtual space data to which it has been pasted. This allows even conference participants in remote locations to use a device for viewing the virtual world to view in the virtual space the object that the conference participants in the real space are gazing at.
[0138] (About hardware configuration) Each of the components of the image analyzing devices 10, 20, and 30 described in the first to third embodiments represents a functional block. Some or all of these components are realized by an information processing device 900 as shown in Fig. 13. Fig. 13 is a block diagram showing an example of the hardware configuration of the information processing device 900.
[0139] As shown in FIG. 13, the information processing device 900 includes the following configuration, for example.
[0140] CPU(Central Processing Unit)901 ROM (Read Only Memory) 902 RAM (Random Access Memory) 903 Program 904 loaded into RAM 903 A storage device 905 for storing a program 904 A drive device 907 that reads and writes data from and to the recording medium 906 A communication interface 908 that connects to a communication network 909 Input / output interface 910 for inputting and outputting data A bus 911 connecting each component Each of the components of the image analyzing devices 10, 20, and 30 described in the first to third embodiments is realized by the CPU 901 reading and executing a program 904 that realizes the functions of the components. The program 904 that realizes the functions of the components is stored in advance in, for example, the storage device 905 or the ROM 902, and is loaded into the RAM 903 and executed by the CPU 901 as needed. The program 904 may be supplied to the CPU 901 via the communication network 909, or may be stored in advance in the recording medium 906, and the drive device 907 may read out the program and supply it to the CPU 901.
[0141] According to the above configuration, the image analyzing devices 10, 20, and 30 described in the first to third embodiments are realized as hardware, and therefore the same effects as those described in any of the first to third embodiments can be achieved.
[0142] (Addendum) One embodiment of the present invention is also described as in the following supplementary notes, but is not limited to the following.
[0143] (Appendix 1) an acquisition means for acquiring a first captured image and a second captured image captured by the image capturing device; an estimation means for estimating the gaze direction of the conference participants appearing in the first captured image; an identification means for identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; An extraction means for extracting an image area of the identified subject of discussion in the second captured image; and recording means for recording data of the image area of the object of discussion. Image analysis device.
[0144] (Appendix 2) the object of discussion is a sticky note, The identification means identifies a sticky note that is being gazed at by the participant of the meeting in the second captured image. 2. The image analysis device according to claim 1,
[0145] (Appendix 3) Further, a recording means is provided for recording a discussion between the participant of the conference and other participants to create an audio file; The recording means records the audio file and the data of the image area of the object of discussion in association with each other. 3. The image analysis device according to claim 1 or 2.
[0146] (Appendix 4) The recording means adds information for identifying the object of discussion that the participants of the conference gazed at and information for identifying the time when the object of discussion was gazed at to the data of the image area of the object of discussion. 4. The image analysis device according to claim 1, wherein:
[0147] (Appendix 5) The method further includes a generating means for generating virtual space data corresponding to the real space including the object of discussion by pasting data of the image area of the object of discussion onto a model of the virtual space after the image area of the object of discussion is extracted. 3. The image analysis device according to claim 1 or 2.
[0148] (Appendix 6) The generating means arranges the plurality of image areas in the virtual space in a depth direction so that the plurality of image areas pasted on the model of the virtual space do not overlap with each other when viewed from a specific viewpoint in the augmented space. 6. The image analysis device according to claim 5,
[0149] (Appendix 7) The generating means arranges the plurality of image regions in the virtual space in the order of gaze times of the object of discussion in the depth direction when viewed from the specific viewpoint. 7. The image analysis device according to claim 5 or 6,
[0150] (Appendix 8) an acquisition means for acquiring a first captured image and a second captured image captured by the image capturing device; an estimation means for estimating the gaze direction of the conference participants appearing in the first captured image; an identification means for identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; An extraction means for extracting an image area of the identified subject of discussion in the second captured image; and recording means for recording data of the image area of the object of discussion. an image analyzer; an imaging device that transmits the second captured image to the image analysis device; a storage device in which data of the image region of the subject of discussion is recorded; A conference support system equipped with
[0151] (Appendix 9) acquiring a first photographed image and a second photographed image photographed by an imaging device; Estimating the gaze direction of the conference participants shown in the first captured image; Identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; Extracting an image area of the identified subject of discussion from the second captured image; recording data of said image region of said subject of discussion; Image analysis methods.
[0152] (Appendix 10) acquiring a first captured image and a second captured image captured by an imaging device; Estimating the gaze direction of the conference participants shown in the first captured image; Identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; Extracting an image area of the identified subject of discussion in the second captured image; recording data of the image region of the object of discussion; A program that causes a computer to execute the following.
[0153] (Appendix 11) the subject of the discussion is a whiteboard; The identifying means identifies a whiteboard that the participants of the meeting are gazing at in the second captured image. 8. The image analysis device according to claim 1, wherein: [Industrial Applicability]
[0154] The present invention can be used, for example, in a conference support system that supports active discussion in an online conference. [Explanation of symbols]
[0155] 1. Meeting support system 10 Image analysis equipment 11 Acquisition Department 12 Estimation part 13 Specific section 14 Extraction part 15 Recording section 20 Image analysis equipment 25 Recording Section 30 Image analysis equipment 35 Generation part 100 user terminals 200 Imaging Device 300 storage device 900 Information Processing Equipment 901 CPU 902 ROM 903 RAM 904 Program 905 Storage device 906 Recording Media 907 Drive unit 908 Communication Interface 909 Communication Network
Claims
1. an acquisition means for acquiring a first captured image and a second captured image captured by the imaging device; an estimation means for estimating the gaze direction of a conference participant appearing in the first captured image; an identification means for identifying an object of discussion that the participants of the conference are gazing at in the second captured image; An extraction means for extracting an image area of the identified subject of discussion in the second captured image; recording means for recording data of the image region of the subject of discussion; The system further includes a generating means for generating virtual space data corresponding to the real space including the object of discussion by pasting data of the image area of the object of discussion onto a model of the virtual space after the image area of the object of discussion is extracted. Image analysis device.
2. the object of discussion is a sticky note, The identification unit identifies a sticky note that is being gazed at by the participant of the meeting in the second captured image.
2. The image analysis device according to claim 1.
3. a recording means for recording the discussion between the participant of the conference and other participants to create an audio file; The recording means records the audio file and the data of the image area of the object of discussion in association with each other.
3. The image analysis device according to claim 1, wherein the image analysis device is a computer.
4. The recording means adds information for specifying a time period during which the object of discussion is gazed at to the data of the image area of the object of discussion.
4. The image analysis device according to claim 1, wherein the image analysis device is a computer.
5. The generating means arranges the plurality of image areas in the virtual space in a depth direction so that the plurality of image areas pasted on the model of the virtual space do not overlap with each other when viewed from a specific viewpoint in the virtual space.
3. The image analysis device according to claim 1, wherein the image analysis device is a computer.
6. The generating means arranges the plurality of image regions in the virtual space in the order of gaze times of the object of discussion in the depth direction when viewed from the specific viewpoint.
6. The image analysis device according to claim 5.
7. an acquisition means for acquiring a first captured image and a second captured image captured by the imaging device; an estimation means for estimating the gaze direction of a conference participant appearing in the first captured image; an identification means for identifying an object of discussion that the participants of the conference are gazing at in the second captured image; An extraction means for extracting an image area of the identified subject of discussion in the second captured image; and recording means for recording data of the image area of the object of discussion. an image analyzer; an imaging device that transmits the second captured image to the image analysis device; a storage device in which data of the image region of the subject of discussion is recorded; Equipped with The image analysis device The system further includes a generating means for generating virtual space data corresponding to the real space including the object of discussion by pasting data of the image area of the object of discussion onto a model of the virtual space after the image area of the object of discussion is extracted. Meeting support system.
8. acquiring a first photographed image and a second photographed image photographed by an imaging device; Estimating the gaze direction of the conference participants shown in the first captured image; Identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; Extracting an image area of the identified subject of discussion from the second captured image; 1. A method of image analysis, recording data of the image region of the subject of discussion, comprising: After the image area of the object of discussion is extracted, data of the image area of the object of discussion is pasted onto a model of a virtual space, thereby generating data of a virtual space corresponding to a real space including the object of discussion. Image analysis methods.
9. acquiring a first captured image and a second captured image captured by an imaging device; estimating the gaze direction of the conference participants shown in the first captured image; Identifying an object of discussion that the participants of the meeting are gazing at in the second captured image; Extracting an image area of the identified subject of discussion in the second captured image; and recording data of the image region of the object of discussion, After the image area of the object of discussion is extracted, data of the image area of the object of discussion is pasted onto a model of a virtual space, thereby generating data of a virtual space corresponding to a real space including the object of discussion. A program for causing the computer to execute the above.
Citation Information
Patent Citations
Conference supporting system, information display, program and control method
JP2005124160A
Projection system and additional information recording method used for the same
JP2006184333A
Information processor, information processing method, information processing program, and network conference system
JP2010206307A
Information terminal, information control method of the information terminal, and information control program
JP2011044044A
Fuel cell and fuel cell stack
JP2014186823A