Information processing device, information processing method, and program

The information processing device enhances video conference systems by using sound direction information to accurately identify and pair participants across different orientations, improving composite image generation accuracy.

JP2026085490APending Publication Date: 2026-05-25RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
RICOH CO LTD
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing video conference systems struggle with accurately identifying the same person across multiple participants due to variations in face orientation and positioning, leading to errors in generating composite images.

Method used

An information processing device that utilizes sound direction information to enhance person identification by correlating sound direction with facial features from different camera angles, ensuring accurate pairing of participants in composite images.

Benefits of technology

Improves the accuracy of identifying the same person across different orientations by using sound direction information, reducing errors in composite image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085490000001_ABST
    Figure 2026085490000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of identifying the same person. [Solution] The information processing device includes an audio processing unit that acquires audio direction information indicating the direction of the sound of speech uttered by a plurality of persons engaging in a dialogue from sound collection devices positioned for the persons said by
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0006] , , , , , , , , , , , , ,

[0005] ,

[0003]

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] When multiple people participate in a video conference, there is a video conference system that can distribute an image including the faces of all participants and an image including the face of each participant close - up to other bases.

[0003] For example, Patent Document 1 describes acquiring a plurality of videos including videos of faces of different orientations for each of a plurality of participants, detecting at least the movement of the face of at least one of the participants from the acquired videos of the faces, selecting a video of the orientation of the face of another participant according to the detected movement of the face of the participant, and generating a video for displaying at least the video of the face of the other participant from the selected video on the terminal device.

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technology described in Patent Document 1, in order to select one video for each participant from videos of faces of different orientations for each of a plurality of participants, it is necessary to accurately identify videos of the same person, but the method is not specified in Cited Document 1.

[0005] The present invention has been made in view of the above points, and an object thereof is to improve the accuracy of identifying the same person. [[ID=3 1]]

Means for Solving the Problems

[0006] To solve the above problem, the information processing device includes: a sound processing unit that acquires sound direction information indicating the direction of the sound of speech uttered by multiple people from sound collection devices positioned for multiple people engaging in a dialogue; and a person identification unit that identifies pairs of the same people for each of the people included in a first image taken by a first imaging device of the multiple people and a second image taken by a second imaging device from a different direction than the first imaging device. The person identification unit identifies the people corresponding to a predetermined range in the direction indicated by the sound direction information as the pair in both the first image and the second image. [Effects of the Invention]

[0007] This can improve the accuracy of identifying the same person. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example configuration of a video conferencing system in the first embodiment. [Figure 2] This figure shows an example of the hardware configuration of the display device 20 in the first embodiment. [Figure 3] This is a diagram illustrating the layout of the base in the first embodiment. [Figure 4] This figure shows an example of a front view image. [Figure 5] This figure shows an example of a center image. [Figure 6] This figure shows an example of a composite image. [Figure 7] This figure shows an example of the functional configuration of the display device 20 in the first embodiment. [Figure 8] This diagram illustrates an example of an error in identifying the same person. [Figure 9] This figure shows an example of a composite image based on incorrect identification results of the same person. [Figure 10] This is a diagram illustrating the direction of speech and the speaker's range. [Figure 11]This diagram illustrates the transformation of coordinate values ​​in the front image. [Figure 12] This is a flowchart illustrating an example of a processing procedure performed by the display device 20 in the first embodiment. [Figure 13] This figure shows an example of the functional configuration of the display device 20 in the second embodiment. [Figure 14] This figure shows an example of the configuration of the person information storage unit 252. [Figure 15] This figure shows an example configuration of the conference information storage unit 253. [Figure 16] This is a flowchart illustrating an example of a processing procedure performed by the display device 20 in the second embodiment. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described below with reference to the drawings. Figure 1 is a diagram showing an example of the configuration of a video conferencing system in the first embodiment. Video conferencing refers to a dialogue (communication) such as a meeting in which audio and video are exchanged between two or more locations, so that each location can understand what is happening at the other locations through audio and video. Video conferencing may be video conferencing or web conferencing.

[0010] Figure 1 shows the system configuration when a video conference is conducted between two locations, location X and location Y. A location refers to a space where people participating in the video conference (hereinafter simply referred to as "participants") gather, such as a conference room. Location X includes a display device 20x and a conference terminal 10x. Location Y includes a display device 20y and a conference terminal 10y. When display devices 20x and 20y are not distinguished, they are simply referred to as "display device 20". When conference terminals 10x and 10y are not distinguished, they are simply referred to as "conference terminal 10".

[0011] The conference terminal 10 is a terminal (device) equipped with the hardware and software necessary for video conferencing, such as a camera, microphone, and speaker. The conference terminal 10 transmits images captured by the camera (imaging device) and audio collected by the microphone (sound collecting device) to the display device 20. In this embodiment, the conference terminal 10 is equipped with two ultra-wide-angle cameras and generates a 360° panoramic image (omnidirectional panoramic image) by combining two 180° panoramic images generated based on the images captured by each ultra-wide-angle camera. Therefore, the conference terminal 10 transmits the 360° panoramic image to the display device 20. The conference terminal 10 is also equipped with multiple microphones and is configured to detect the direction of the sound source based on the volume of sound collected by the multiple microphones. The conference terminal 10 also outputs audio (from other locations) transmitted from the display device 20 through its speaker. The conference terminal 10 is basically positioned so that all participants within the location can be captured. However, it is acceptable if some participants are not captured.

[0012] The display device 20 is a device on which a video conferencing application is installed and which has communication functions, UI functions, and information processing functions, and is an example of an information processing device. The display device 20x transmits images taken by the conference terminal 10x at site X and audio collected by the conference terminal 10x to the display device 20y at site Y, and controls the display of images and output of audio from the display device 20y at site Y. For example, a PC (Personal Computer) or an IWB (Interactive White Board: an electronic whiteboard with mutual communication capabilities) may be used as the display device 20. In this embodiment, it is assumed that an IWB is used as the display device 20.

[0013] In FIG. 1, an example is shown in which the display devices 20 at two locations are connected to the server 30 via the network 40, and the communication of the display devices 20 at two locations is mediated via the server 30. However, the display devices 20 at three or more locations may be connected to the server 30 via the network 40. In this case, by mediating the communication between the display devices 20 at any plurality of locations among the three or more locations by the server 30, it is possible to hold a video conference between any plurality of locations.

[0014] FIG. 2 is a diagram showing an example of the hardware configuration of the display device 20 in the first embodiment. As shown in FIG. 2, the display device 20 includes a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, a SSD (Solid State Drive) 204, a network I / F (Interface) 205, and an external device connection I / F 206.

[0015] Among these, the CPU 201 controls the operation of the entire display device 20. The ROM 202 stores programs used for driving the CPU 201 such as the CPU 201 and the IPL (Initial Program Loader). The RAM 203 is used as a work area for the CPU 201. The SSD 204 stores various data such as programs for an electronic blackboard. The network I / F 205 controls communication with the communication network 100. The external device connection I / F 206 is an interface for connecting various external devices. The external devices in the present embodiment are the conference terminal 10 in addition to the speaker 250 and the camera 260. The speaker 250 outputs the voice received from the display devices 20 at other locations and the like. The camera 260 captures an image of the front of the display device 20 (IWB).

[0016] In addition, the display device 20 includes a capture device 211, a GPU (Graphics Processing Unit) 212, a display controller 213, a touch sensor 214, a sensor controller 215, an electronic pen controller 216, a short-range communication circuit 219, an antenna 219a of the short-range communication circuit 219, a power switch 222, and selection switches 223.

[0017] Among these, the capture device 211 causes an external PC (Personal Computer) 270 to display image information as a still image or a moving image on its display. The GPU 212 is a semiconductor chip specialized in handling graphics. The display controller 213 controls and manages screen display in order to output the output image from the GPU 212 to the display 280 or the like. The touch sensor 214 detects that an electronic pen 290, the user's hand H, etc. have touched the display 280. The sensor controller 215 controls the processing of the touch sensor 214. The touch sensor 214 performs coordinate input and coordinate detection by an infrared blocking method. The method of performing this coordinate input and coordinate detection is such that two light emitting and receiving devices installed at both upper ends of the display 280 emit a plurality of infrared rays parallel to the display 280, and the light is reflected by a reflecting member provided around the display 280 and returns on the same optical path as the optical path of the light emitted by the light receiving element, and the light receiving element receives the light. The touch sensor 214 outputs the IDs of the infrared rays emitted by the two light emitting and receiving devices blocked by an object to the sensor controller 215, and the sensor controller 215 specifies the coordinate position that is the contact position of the object. The electronic pen controller 216 determines the presence or absence of pen tip touch and pen butt touch on the display 280 by communicating with the electronic pen 290. The short-range communication circuit 219 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The power switch 222 is a switch for switching the ON / OFF of the power supply of the display device 20. The selection switches 223 are, for example, a group of switches for adjusting the brightness, color tone, etc. of the display of the display 280.

[0018] Furthermore, the display device 20 is equipped with a bus line 210. The bus line 210 is an address bus, data bus, etc., for electrically connecting each component, such as the CPU 201 shown in Figure 2.

[0019] Furthermore, the contact sensor 214 is not limited to the infrared blocking method. Various detection means may be used, such as a capacitive touch panel that identifies the contact position by detecting changes in capacitance, a resistive touch panel that identifies the contact position by voltage changes between two opposing resistive films, or an electromagnetic induction touch panel that identifies the contact position by detecting electromagnetic induction caused by contact between an object and the display. In addition, the electronic pen controller 216 may be configured to determine whether or not there is touch not only at the tip and end of the electronic pen 290, but also at the part of the electronic pen 290 held by the user or other parts of the electronic pen.

[0020] Figure 3 is a diagram illustrating the layout of the base in the first embodiment. Figure 3 is a top view of base X (for example, a conference room). In Figure 3, a table t1 is placed in front of the display device 20x, which acts as an IWB, and six participants A to F (represented by circles in the figure) are seated around the table t1. A conference terminal 10x is placed on the table t1. Each participant seated around the table t1 is photographed by the camera of the conference terminal 10 (hereinafter referred to as the "center camera") and also by the camera 260 of the display device 20 (hereinafter referred to as the "front camera"). That is, each participant is photographed from different directions by the center camera and the front camera.

[0021] The display device 20 generates a composite image for display at another location (location Y in this embodiment) based on an image captured by the front camera (hereinafter referred to as the "front image") and a 360° panoramic image captured by the center camera (hereinafter referred to as the "center image").

[0022] Figure 4 shows an example of a front image, and Figure 5 shows an example of a center image. Both images are examples of images taken in the layout shown in Figure 3. Since the center image is a 360° panoramic image, each participant is arranged horizontally.

[0023] Figure 6 shows an example of a composite image. As shown in Figure 6, the composite image is generated by combining the facial image of each participant (hereinafter referred to as "facial image") with a front body image. When such a composite image is displayed on the display device 20y at base Y, participants at base Y can see the facial expressions of each participant, as well as the overall situation at base X.

[0024] In this embodiment, the composite image is generated based on the idea that, from the viewpoint of the visibility of the participants' facial expressions, it is desirable for each participant's face image to be as frontal as possible. However, the direction of each participant's face is not necessarily constant. Some participants are looking at the display 280 of the display device 20 to view composite images from other locations (i.e., facing the display device 20), while others are looking at other participants at the same location (i.e., facing the conference terminal 10). The display device 20 has a functional configuration that enables it to generate a composite image that includes a face image as close to a frontal view as possible for each participant, even in such situations.

[0025] Figure 7 shows an example of the functional configuration of the display device 20 in the first embodiment. In Figure 7, the display device 20 includes an image storage unit 21, a person detection unit 22, a face feature quantity calculation unit 23, a voice processing unit 24, a same person identification unit 25, a selection unit 26, an image generation unit 27, and a virtual camera 28. Each of these units is realized by processing that one or more programs installed on the display device 20 cause the CPU 201 to execute. The display device 20 also utilizes an image storage unit 251. The image storage unit 251 can be realized using, for example, a RAM 203, an SSD 204, or a storage device that can be connected to the display device 20 via a network.

[0026] The image storage unit 21 temporarily stores the images transmitted from the front camera and the center camera (front image and center image) as image files in the image storage unit 251.

[0027] The person detection unit 22 performs person detection on each image file (front image and center image) stored in the image storage unit 251. Person detection refers to identifying the range (region) of a person's face in an image, and can be performed using publicly known techniques. For each image (front image and center image), the person detection unit 22 outputs coordinate values ​​(hereinafter referred to as "face coordinates") indicating the identified range of the face. The face coordinates identified in the front image are 2D coordinate values ​​in the coordinate system of the front image (see Figure 4). The face coordinates identified in the center image are 2D coordinate values ​​in the coordinate system of the center image (see Figure 5). Since the coordinate systems of the front image and the center image are different, the face coordinates identified in the front image and the face coordinates identified in the center image do not necessarily indicate the same location because they belong to different coordinate systems.

[0028] The face feature calculation unit 23 calculates the feature quantities (hereinafter referred to as "face features") of the image (face image) for each face range in each image, based on the face coordinates in each image output from the person detection unit 22 for both the front image and the center image. The face feature calculation unit 23 associates the calculated face features for the face image within the range indicated by each face coordinate with each face coordinate. Hereinafter, information including the pair of face coordinate and face feature (face coordinate, face feature) will be referred to as face information. Note that the calculation of face image features may be performed using known methods. Face features may also be represented by numerical vectors.

[0029] The audio processing unit 24 acquires the audio of a participant's speech and information on the direction of the speech (the direction from which the speech was made) (hereinafter referred to as "audio direction information") from the conference terminal 10, and inputs the audio and audio direction information into the person identification unit 25. The audio direction information can be any information that indicates the direction in the coordinate system of the center image, for example.

[0030] The Identical Person Identification Unit 25 identifies pairs of identical individuals among participants (people) included in both a front image (an example of a first image) captured by the camera 260 of the display device 20 (an example of a first imaging device) and a center image (an example of a second image) captured by the camera of the conference terminal 10 (an example of a second imaging device) from a different direction than the camera 260. Specifically, the Identical Person Identification Unit 25 identifies identical individuals between participants included in the front image and participants included in the center image based on a comparison of the facial features of each facial information obtained from the front image (hereinafter referred to as "front facial information") and the facial features of each facial information obtained from the center image (hereinafter referred to as "center facial information"). For example, the same person identification unit 25 calculates the similarity (e.g., cosine similarity) between the facial feature quantities of each front face information and the facial feature quantities of each center face information for each front face information, and identifies pairs of face information (pairs of front face information and center face information) whose similarity of facial feature quantities is below a threshold (similarity is higher than a predetermined value) as pairs of face information relating to the same person. In other words, in this embodiment, the identification of the same person is achieved by identifying pairs of face information relating to the same person. If the similarity of the facial feature quantities of one face information to the facial feature quantities of multiple other face information is below a threshold, then the face information relating to the facial feature quantity with the smallest similarity (i.e., the highest similarity) and the one face information in question should be identified as pairs of face information relating to the same person.

[0031] The selection unit 26 selects the face information to be combined (face information corresponding to the face image to be included in the combined image) from among the front face information and center face information. For face information that constitutes a set of face information relating to the same person, the selection unit 26 selects one of the face information to be combined based on factors such as the face image that is more facing forward or the image quality when the face image is enlarged. In other words, if the degree to which each face image in a set of face information relating to the same person is facing forward is about the same, the selection unit 26 selects the face information that includes the face image with a larger face area. For example, a method that is generally known is to detect facial feature points (landmarks) and detect the orientation of the face from the geometric positional relationship of the feature points. On the other hand, the selection unit 26 selects face information that does not constitute a set of face information relating to the same person as is for combination. The selection unit 26 inputs each of the face information selected as a combination target to the image generation unit 27.

[0032] The image generation unit 27 extracts a face image from the front image based on the face coordinates contained in the face information obtained from the front image, among the face information input from the selection unit 26. The image generation unit 27 also extracts a face image from the center image based on the face coordinates contained in the face information obtained from the center image, among the face information input from the selection unit 26. The image generation unit 27 generates a single composite image in which the extracted face images and the front image are arranged in a predetermined format, and outputs this composite image to the virtual camera 28. When generating the composite image, if the area of ​​the face image included in the face information and the image area per person used for the composite image differ, the image is adjusted by scaling. This allows the video conferencing application to use the composite image as the camera image. The virtual camera 28 is a camera that virtually captures a composite image, which is not an image of the real world. The virtual camera 28 outputs the composite image (generated by the image generation unit 27) and transmits the composite image to a display device 20 at another location via the network 40. Alternatively, the display device 20 may not have a virtual camera 28, and the composite image may be transmitted directly via the network 40.

[0033] By the way, if the same person identification unit 25 identifies the same person based solely on facial features, the following errors may occur in the identification result.

[0034] Figure 8 illustrates an example of an error in identifying the same person. In Figure 8, for participants A, B, C, and D, examples are shown where pairs of facial information were identified where the similarity of facial features between the front image and the center image was below a threshold. On the other hand, for participant E, an example is shown where, despite actually being present in both the front and center images, the similarity of facial features between the two images exceeded the threshold, leading to the determination that the two facial images did not belong to the same person. For example, if participant E's front facial image is a profile image that does not include much of the front view, and participant E's center facial image is a direct front view, the facial features of both facial images may not be below the threshold.

[0035] Furthermore, regarding participant F, because they were sitting close to the front camera, their face was outside the camera's field of view, and therefore their face image was not included in the front image, either partially or completely. As a result, facial information could not be obtained from the front image.

[0036] In this case, the selection unit 26 will select only one of either the front face information or the center face information as the target for synthesis for participants A, B, C, and D, but for participant E, it will ultimately select both the front face information and the center face information as the target for synthesis. This is because neither the front face information nor the center face information of participant E constitutes a set of face information relating to the same person. Similarly, since the center face information of participant F does not constitute a set of face information relating to the same person, the selection unit 26 also selects the center face information of participant F as the target for synthesis. As a result, the image generation unit 27 will generate a composite image as shown in Figure 9.

[0037] Figure 9 shows an example of a composite image based on incorrect identification of the same person. In Figure 9, "(Front)" means a face image extracted from the front image, and "(Center)" means a face image extracted from the center image. In Figure 9, the composite image includes face images extracted from both the front and center images for person E. As a result, although there are 6 participating members, a composite image containing the face images of 7 people has been generated.

[0038] Therefore, in order to reduce the possibility of such a situation occurring, the person identification unit 25 identifies pairs of identical individuals based on whether or not they are participants corresponding to a predetermined range in the direction indicated by the sound direction information in both the front image and the center image. In other words, the person identification unit 25 improves the accuracy of person identification by also using sound direction information from the sound processing unit 24.

[0039] For example, suppose that when the front image and center image are captured at the time of composite image generation, person E is speaking. In this case, the voice processing unit 24 inputs the voice and voice direction information for the utterance to the person identification unit 25. The person identification unit 25 identifies a pair of front face coordinates and center face coordinates that does not constitute a pair of face information relating to the same person, and whose face coordinates fall within a predetermined range (hereinafter referred to as the "speaker range") in which the speaker is presumed to be present in the direction indicated by the voice direction information, as a pair of face information relating to the same person.

[0040] Figure 10 is a diagram illustrating the direction of speech and the speaker range. Figure 10 shows the direction of speech and the speaker range a1 when participant E is speaking. In this case, the direction of speech refers to the direction of participant E as seen from the conference terminal 10. The speaker range a1 refers to, for example, the range along an extension line of approximately ±5 degrees from the conference terminal 10 in that direction, and is specified as the range of the coordinate system of the center image.

[0041] In this case, if there is one front face information and one center face information whose face coordinate range is included in speaker range a1, and the face coordinates of other face information are located more than a certain distance away from speaker range a1, then it is highly likely that the pair of front face information and center face information whose face coordinate range is included in speaker range a1 belongs to the same person.

[0042] Therefore, in this case, even if the front face information and center face information do not constitute a set relating to the same person based on the similarity of the face features, the front face information and center face information set whose respective face coordinates are within the speaker's range will be identified as a set of face information relating to the same person. As a result, the possibility of generating a composite image like the one shown in Figure 9 (a composite image containing two face images of the same participant) becomes lower, and the possibility of generating a composite image like the one shown in Figure 6 becomes higher.

[0043] Furthermore, while the speaker range is identified in the coordinate system of the center image, the face coordinates of the front face information are coordinate values ​​in the coordinate system of the front image, so the two cannot be directly compared. Therefore, the person identification unit 25 converts the face coordinates of the front face information into coordinate values ​​in the coordinate system of the center image, and then determines whether or not those face coordinates are included in the speaker range.

[0044] Figure 11 is a diagram illustrating the transformation of coordinate values ​​in the front image. Figure 11 shows the positional relationship (positional relationship as viewed from the side) between the front camera (camera 260 of the display device 20) and the conference terminal 10 (center camera). The positional relationship between the front camera and the center camera can be determined from the size of the conference terminal 10 (i.e., the size of the conference terminal 10 in the front image) and the angle α when the conference terminal 10 (center camera) is photographed from the front camera. In the front image (Figure 4), the position P1 (Figure 4) where the conference terminal 10 (center camera) is directly in front of the display device 20 is the origin (0,0) in the coordinate system of the center image. Based on this origin and the positional relationship between the front camera and the center camera, the coordinate values ​​of the front camera's coordinate system can be transformed into the coordinate values ​​of the center camera's coordinate system.

[0045] Furthermore, correspondence information may be stored in the SSD204 or the like in advance, where the facial features (face features) of a face image are associated with the voice features (hereinafter referred to as "voice features") for each person. The same person identification unit 25 may use this correspondence information to improve the accuracy of identifying the same person. In this case, the speech processing unit 24 calculates the voice features of the utterance. The same person identification unit 25 identifies pairs of the same person based on a comparison between the voice features associated in the correspondence information and the utterance of the voice features of the face image that is most similar to the facial features of the face image included in the front image or center image, for participants corresponding to a predetermined range (speaker range) in the direction indicated by the speech direction information, among the facial features associated with the voice features for each person.

[0046] More specifically, the person identification unit 25 calculates the similarity between each of the two face features of the pair of face information (in the example in Figure 8, the front face information and center face information of participant E) that constitute the pair of face information related to the same person identified using the voice direction information, and each face feature included in the corresponding information. The person identification unit 25 selects one of the two face information that has the smallest similarity to the face feature included in the corresponding information. The person identification unit 25 calculates the similarity (e.g., cosine similarity) between the voice feature associated in the corresponding information with the face feature that had the smallest similarity to the face feature of the selected face information, and the voice feature from the voice processing unit 24. The identical person identification unit 25 determines that if the similarity is below a threshold, the set of facial information identified using the speech direction information is correct (it determines that it is indeed the set of facial information of the participant who made the utterance), and if at least one of the similarities exceeds the threshold, it determines that the set of facial information related to the identical person is not the set of facial information related to the identical person. Therefore, in this case, the set of facial information is not identified as the set of facial information related to the identical person.

[0047] The following describes the processing steps performed by the display device 20. Figure 12 is a flowchart illustrating an example of the processing steps performed by the display device 20 in the first embodiment. The point in time when the processing steps in Figure 12 are being performed is called "time t". The processing steps in Figure 12 are repeated 30 times per second if the frame rate of the video conference is, for example, 30 fps. In this case, time t will occur 30 times per second.

[0048] In step S110, the image storage unit 21 stores the front image captured by the front camera at time t and the center image captured by the center camera at time t in the image storage unit 251.

[0049] Next, the person detection unit 22 performs person detection on both the front image and the center image, and outputs face coordinates for each face detected in each image (S120).

[0050] Next, the face feature calculation unit 23 calculates face features for each face image identified based on the face coordinates in each image, which are output from the person detection unit 22 for both the front image and the center image (S130). The face feature calculation unit 23 generates face information for each face coordinate, including the face coordinate and the face feature corresponding to that face coordinate, and inputs each piece of face information to the same person identification unit 25.

[0051] Meanwhile, the audio processing unit 24 acquires the audio collected by the conference terminal 10 at time t and the audio direction information of said audio from the conference terminal 10, and inputs said audio and said audio direction information to the person identification unit 25 (S140). Step S140 may be executed in parallel with steps S110 to S130.

[0052] Next, the person identification unit 25 identifies pairs of face information relating to the same person based on face features (without using audio direction information) using the method described above, with respect to the front face information obtained from the front image and the center face information obtained from the center image (S150).

[0053] Next, the person identification unit 25 determines whether or not one or more front face information and one or more center face information remain as face information that does not constitute a set of face information relating to the same person (S160).

[0054] If at least one of the front face information and the center face information does not remain as face information that does not constitute a set of face information relating to the same person (No in S160), proceed to step S180. If one or more front face information and one or more center face information remain as face information that does not constitute a set of face information relating to the same person (Yes in S160), the same person identification unit 25 uses the speaker range based on the speech direction information to identify a set of face information relating to the same person from the remaining face information using the method described above (S170), and proceed to step S180.

[0055] In step S180, the selection unit 26 selects the face information to be synthesized from the front face information and the center face information using the method described above. That is, the selection unit 26 selects either the front face information or the center face information that constitutes a set of face information relating to the same person, and selects the front face information or the center face information that does not constitute a set of face information relating to the same person as the target for synthesis.

[0056] Next, the image generation unit 27 extracts a face image corresponding to the face coordinates of the face information from the image (front image or center image) that corresponds to the face information selected by the selection unit 26 (S190).

[0057] Next, the image generation unit 27 generates a composite image including the extracted face image and the front image (S200).

[0058] Next, the image generation unit 27 outputs the generated composite image to the virtual camera 28 (S210).

[0059] Steps S110 to S210 are repeated, for example, until the meeting ends (S220).

[0060] In addition, the same person identification unit 25 may perform same person identification in step S150 using a machine learning model that has already learned the correspondence between face images and people (hereinafter referred to as the "person estimation model"). Specifically, for each piece of face information, the same person identification unit 25 extracts a face image from the corresponding image (front image or center image) based on the face coordinates of the face information. For each piece of face information, the same person identification unit 25 inputs the extracted face image for that face information into the person estimation model and identifies the person corresponding to that face information based on the output from the person estimation model (for example, the probability for each person). The same person identification unit 25 identifies a pair of front face information and center face information that have been identified as the same person as a pair of face information relating to the same person. Even when using a person estimation model, depending on the state of the face image and the learning status of the person estimation model, it may not be possible to identify a person (for example, if the probability is the same for multiple people, or if the probability for any person does not exceed a threshold), and face information may be generated that does not form a set of face information relating to the same person. Therefore, steps S160 and beyond may be executed.

[0061] As described above, according to the first embodiment, the person (participant) included in the front image and the person (participant) included in the center image are identified as the same person not only by the facial image in each image but also by the direction of the speech. Therefore, the accuracy of identifying the same person can be improved.

[0062] Next, a second embodiment will be described. The differences between the second embodiment and the first embodiment will be described. Therefore, points not specifically mentioned may be the same as in the first embodiment.

[0063] Figure 13 shows an example of the functional configuration of the display device 20 in the second embodiment. In Figure 13, the same reference numerals are used for parts that are the same as or corresponding to parts in Figure 7, and their descriptions are omitted.

[0064] As shown in Figure 13, the display device 20 in the second embodiment further utilizes a person information storage unit 252 and a meeting information storage unit 253. Each of these storage units can be implemented using an SSD 204 or a storage device connected to the display device 20 via a network.

[0065] Figure 14 shows an example of the configuration of the person information storage unit 252. As shown in Figure 14, the person information storage unit 252 pre-stores user ID, name, face image, and voice information for each person who is a candidate participant in a video conference. A person's face image is taken in advance and recorded in the person information storage unit 252 in association with that person's user ID. The voice information for a person is the voice information of that person. For example, when that person conducts a video conference, the voice when the speaker matches that person's face image may be recorded as voice information.

[0066] Figure 15 shows an example of the configuration of the meeting information storage unit 253. As shown in Figure 15, the meeting information storage unit 253 stores meeting information for each scheduled meeting, including the meeting name, meeting time, and meeting participant IDs. The meeting name is the name of the meeting. The meeting time is information indicating the time period of the meeting (the time interval from the start time to the end time). The meeting participant IDs are the user IDs of all participants taking part in the meeting. These user IDs are the user IDs stored in the person information storage unit 252.

[0067] In the second embodiment, an example is described in which the accuracy of identifying the same person is improved by using facial images stored in the person information storage unit 252 in advance for each participant in the video conference, as well as the number of participants that can be identified from the meeting information of the video conference.

[0068] For example, in the second embodiment, the same person identification unit 25 identifies a group of identical people based on a comparison between a face image pre-stored in the person information storage unit 252 for each person and the face image of that person included in the front image and the center image, respectively, for individuals who cannot be identified as a group of identical people by the method of the first embodiment. Furthermore, for individuals who cannot be identified as a group of identical people based on a comparison between a face image pre-stored in the person information storage unit 252 for each person and the face image of that person included in the front image and the center image, respectively, the same person identification unit 25 identifies a group of identical people based on a comparison between the voice (speech information) feature quantity associated with a face image pre-stored in the person information storage unit 252 for each person and the speech feature quantity of the utterance at that time.

[0069] For example, suppose there are three meeting participants, A, B, and C. Two of their faces can be identified as belonging to the same person, but one face from the front image and one face from the center image cannot. In this case, the facial feature quantities of the two faces identified as belonging to the same person (from the front image or center image) are compared with the facial feature quantities of each face stored in the person information storage unit 252 to identify who the two people are (i.e., their user IDs). If the two people are A and B, then the two faces that cannot be identified as belonging to the same person can be identified as the pair of C's faces.

[0070] Furthermore, if there are four meeting participants, A, B, C, and D, and the facial information of two of them can be identified as belonging to the same person, as described above, and it can be determined that these two are A and B based on the facial images stored in the person information storage unit 252, then the multiple facial information that cannot be identified as belonging to the same person will be the facial information of C or D. By determining which facial feature quantity of the facial image stored for C or D in the person information storage unit 252 is more similar to the facial feature quantity of each of these multiple facial information, the multiple facial information that cannot be identified as belonging to the same person can be classified into sets of facial information of C and sets of facial information of D.

[0071] Furthermore, by using voice and voice features, it is possible to improve the accuracy of identifying the same person by determining whether it is C or D.

[0072] Figure 16 is a flowchart illustrating an example of a processing procedure performed by the display device 20 in the second embodiment. In Figure 16, steps identical to those in Figure 12 are given the same step numbers, and their explanations are omitted.

[0073] In Figure 16, steps S171 to S176 are added between steps S170 and S180.

[0074] In step S171, the person identification unit 25 determines whether one or more front face information and one or more center face information remain as face information that does not constitute a set of face information relating to the same person. In other words, even using the person identification method in the first embodiment, it is determined whether there are any sets that may not be identified as the face information of the same person.

[0075] If at least one of the front face information and the center face information is not found to be the relevant face information (No in S171), the process proceeds to step S180. If the relevant face information is found (Yes in S171), the same person identification unit 25 identifies the number of participants in the meeting based on the meeting information stored in the meeting information storage unit 253 (Figure 15) for the meeting in question, and determines whether the number of sets of face information already identified as relating to the same person is equal to (number of participants in the meeting - 1) (S172). For example, if the number of participants in the meeting in question is three, it is determined whether two sets of face information relating to the same person have been identified. In other words, it is determined whether only one participant has not had a set of face information relating to the same person identified. Note that the meeting information for the meeting in question may be determined by obtaining the meeting name of the meeting in question from the video conferencing application.

[0076] If the number of sets of facial information identified as relating to the same person is the same as (number of participants in the meeting - 1) (i.e., if there is only one participant for whom a set of facial information relating to the same person has not been identified) (Yes in S172), the same person identification unit 25 identifies the front facial information and center facial information that have not been identified as a set constituting facial information relating to the same person as a set of facial information relating to the same person (S173), and proceeds to step S180.

[0077] On the other hand, if the number of sets of facial information identified as relating to the same person is less than (number of participants in the meeting - 1) (i.e., if sets of facial information relating to the same person have not been identified for two or more participants) (No in S172), the same person identification unit 25 uses the person information stored in the person information storage unit 252 (Figure 14) to identify the same person for the remaining facial information (S174). Specifically, the same person identification unit 25 identifies the participant (user ID) corresponding to each facial information that already constitutes a set of facial information relating to the same person based on a comparison (for example, calculation of similarity) between the facial feature quantities of each facial information that already constitutes a set of facial information relating to the same person and the facial feature quantities of each facial image among the facial images stored in the person information storage unit 252 (Figure 14) that correspond to the participants of the meeting in question. Hereinafter, the set of participants (user IDs) from the participants of the meeting in question, excluding those identified as corresponding to each facial information that already constitutes a set of facial information relating to the same person, will be referred to as "Set A". The person identification unit 25 compares the facial feature quantities of the front face information and center face information, which do not constitute a set of facial information relating to the same person, with the facial feature quantities of each facial image other than the facial image stored in the person information storage unit 252 (Figure 14) for each participant belonging to set A among the participants of the target meeting, to identify the participant (user ID) corresponding to each facial information that does not constitute a set of facial information relating to the same person. The person identification unit 25 identifies the set of front face information and center face information that have identified the same participant (user ID) as a set of facial information relating to the same person.

[0078] Following step S174, the person identification unit 25 determines whether there are still any face information remaining that does not constitute a set of face information relating to the same person (S175). If at least one of the front face information and the center face information does not constitute a set of face information relating to the same person (No in S175), the process proceeds to step S180.

[0079] On the other hand, in step S174, if, for each participant belonging to set A, there are two or more facial features with the highest similarity (facial features with the same minimum similarity value) among the facial features of the facial images stored in the person information storage unit 252 (Figure 14), then the participant cannot be identified for that facial information. In such cases, where a participant cannot be identified by facial features (hereinafter, the set of facial information in question is referred to as "set B"), even after the execution of step S174, there is still a possibility that one or more front facial information and one or more center facial information remain as facial information that does not constitute a set of facial information relating to the same person. In this case (Yes in S175), the same person identification unit 25 identifies the same person by identifying the participant (user ID) for the facial information belonging to set B based on the speech (voice) features (hereinafter, referred to as "speech speech features") spoken at time t and input from the speech processing unit 24 (S176). Specifically, for each piece of face information belonging to set B (i.e., face information in which multiple participants (user IDs) have been identified), the same person identification unit 25 identifies the participant (user ID) corresponding to the face information by selecting the "voice information" stored in the person information storage unit 252 for each participant (user ID) identified for that face information, and which has the smallest similarity between its feature quantity (voice feature quantity) and the utterance voice feature quantity (highest similarity to the generated voice feature quantity). Therefore, if a participant whose face information belongs to set B was speaking at this point t, the participant (user ID) can be identified for the face information belonging to set B. If there is other piece of face information identified in step S174 or S175 that corresponds to the same participant (user ID) as the face information in question, the same person identification unit 25 identifies these two pieces of face information as a pair of face information relating to the same person and proceeds to step S180.

[0080] Steps S180 and beyond are as described in the first embodiment.

[0081] As described above, according to the second embodiment, the accuracy of identifying the same person can be further improved by using meeting information and person information.

[0082] Each function of each embodiment can be implemented by one or more processing circuits. Hereinafter, "processing circuit" as used herein includes processors programmed to execute each function by software, such as processors implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (digital signal processors), FPGAs (field programmable gate arrays), and conventional circuit modules designed to execute the functions described above.

[0083] Although embodiments of the present invention have been described in detail above, the present invention is not limited to these specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention as described in the claims.

[0084] Examples of the present invention are as follows:

[0085] <1> A sound processing unit that acquires sound direction information indicating the direction of the speech of multiple people engaging in a dialogue from sound collection devices positioned in front of the people engaged in the dialogue, A person identification unit that identifies the same group of people in a first image taken by a first imaging device of the aforementioned multiple people, and a second image taken by a second imaging device of the aforementioned multiple people from a different direction than the first imaging device, It has, The aforementioned person identification unit identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the aforementioned pair. An information processing device characterized by the following: <2> A selection unit selects, from among the sets of identical persons identified by the identical person identification unit, the face image of the first image corresponding to the identical person or the face image of the second image corresponding to the identical person that is facing forward the most; An image generation unit that generates a composite image by combining multiple face images selected in the selection unit, An output unit that outputs the composite image, The information processing apparatus according to claim 1, characterized by having the following features.

[0086] <3> The aforementioned speech processing unit further acquires the speech from the sound collection device and calculates the characteristic quantities of the speech. The aforementioned person identification unit identifies the pair based on a comparison between the voice features of the utterance and the voice features of the utterance, for each person whose voice features are associated with a face image feature that is most similar to the face image feature included in the first or second image, for a person whose voice features correspond to a predetermined range in the direction indicated by the speech direction information. Characterized by <1> The information processing device described.

[0087] <4> The aforementioned identical person identification unit identifies identical person pairs for individuals for whom a pair of identical people cannot be identified, based on a comparison between a face image pre-stored for each individual and the face images of the individuals included in the first image and the second image, respectively. Characterized by <1> ~ <3> Any of the information processing devices described above.

[0088] <5> The aforementioned speech processing unit further inputs the speech from the sound collection device and calculates the characteristic quantities of the speech. The aforementioned person identification unit identifies the pair of individuals for whom it is not possible to identify a pair of individuals based on a comparison of the facial images pre-stored for each individual with the facial images of the individuals included in the first image and the second image, based on a comparison of the voice feature quantities associated with the facial images pre-stored for each individual with the voice feature quantities of the utterance. Characterized by <4> The information processing device described.

[0089] <6> A speech processing procedure that acquires speech direction information indicating the direction of speech uttered by multiple people engaging in a dialogue from sound collection devices positioned for each person, A procedure for identifying identical individuals, which identifies sets of individuals included in a first image taken by a first imaging device of the aforementioned multiple individuals, and a second image taken by a second imaging device of the aforementioned multiple individuals from a different direction than the first imaging device, The computer executes this, The aforementioned procedure for identifying identical persons identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the pair. An information processing method characterized by the following:

[0090] <7> A speech processing procedure that acquires speech direction information indicating the direction of speech uttered by multiple people engaging in a dialogue from sound collection devices positioned for each person, A procedure for identifying identical individuals, which identifies sets of individuals included in a first image taken by a first imaging device of the aforementioned multiple individuals, and a second image taken by a second imaging device of the aforementioned multiple individuals from a different direction than the first imaging device, Have the computer run it, The aforementioned procedure for identifying identical persons identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the pair. A program characterized by the following features. [Explanation of symbols]

[0091] 10x conference terminals 10y conference terminal 20x display device 20y display device 30 servers 40 Networks 21 Image storage section 22 Person Detection Unit 23. Facial Feature Calculation Unit 24. Audio Processing Unit 25 Same person identification department 26 Selection Section 27 Image Generation Unit 28 Virtual Camera 251 Image storage unit 252 Person information storage section 253 Conference Information Storage Unit [Prior art documents] [Patent Documents]

[0092] [Patent Document 1] International Publication No. 2022 / 154128

Claims

1. A sound processing unit that acquires sound direction information indicating the direction of the speech of multiple people engaging in a dialogue from sound collection devices positioned in front of the people engaged in the dialogue, A person identification unit identifies the same person for each of the persons included in a first image taken by a first imaging device of the aforementioned multiple persons and a second image taken by a second imaging device of the aforementioned multiple persons from a different direction than the first imaging device, It has, The aforementioned person identification unit identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the aforementioned pair. An information processing device characterized by the following:

2. A selection unit selects, from among the sets of identical persons identified by the identical person identification unit, the face image of the first image corresponding to the identical person or the face image of the second image corresponding to the identical person that has a higher degree of facing forward. An image generation unit that generates a composite image by combining multiple face images selected in the selection unit, An output unit that outputs the composite image, The information processing apparatus according to claim 1, characterized by having the following features.

3. The aforementioned speech processing unit further acquires the speech from the sound collection device and calculates the characteristic quantities of the speech. The aforementioned person identification unit identifies the pair based on a comparison between the voice features of the utterance and the voice features of the utterance, for each person whose voice features are associated with a face image feature that is most similar to the face image feature included in the first or second image, for a person whose voice features correspond to a predetermined range in the direction indicated by the speech direction information. The information processing apparatus according to claim 1, characterized in that it is a product of the present invention.

4. The aforementioned identical person identification unit identifies identical person groups for individuals for whom a group of identical individuals cannot be identified, based on a comparison between a face image pre-stored for each individual and the face images of the individuals included in the first image and the second image, respectively. The information processing apparatus according to claim 1, characterized in that it is a product of the present invention.

5. The aforementioned speech processing unit further inputs the speech from the sound collection device and calculates the characteristic quantities of the speech. The identical person identification unit identifies individuals who cannot be identified as identical based on a comparison of the facial images pre-stored for each individual with the facial images of the individuals included in the first image and the second image, by comparing the voice feature quantities associated with the facial images pre-stored for each individual with the voice feature quantities of the utterance. The information processing apparatus according to feature 4.

6. A speech processing procedure that acquires speech direction information indicating the direction of speech uttered by multiple people engaging in a dialogue from sound collection devices positioned for each person, A procedure for identifying identical individuals, which identifies sets of individuals included in a first image taken by a first imaging device of the aforementioned multiple individuals, and a second image taken by a second imaging device of the aforementioned multiple individuals from a different direction than the first imaging device, The computer executes this, The aforementioned procedure for identifying identical persons identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the pair. An information processing method characterized by the following:

7. A speech processing procedure that acquires speech direction information indicating the direction of speech uttered by multiple people engaging in a dialogue from sound collection devices positioned for each person, A procedure for identifying identical individuals, which identifies sets of individuals included in a first image taken by a first imaging device of the aforementioned multiple individuals, and a second image taken by a second imaging device of the aforementioned multiple individuals from a different direction than the first imaging device, Have the computer run it, The aforementioned procedure for identifying identical persons identifies the person corresponding to a predetermined range in the direction indicated by the audio direction information in both the first image and the second image as the pair. A program characterized by the following features.