Information processing device, online meeting system, information processing method, and computer program
The information processing device estimates and reconstructs facial expressions under occlusions by combining non-occluded facial regions, addressing the challenge of capturing natural-looking faces without masks in crowded settings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2026-04-01
AI Technical Summary
Existing technologies struggle to accurately reconstruct facial expressions of individuals wearing occlusions such as masks, leading to uninteresting photographs in crowded settings where natural-looking faces without masks are desired.
An information processing device and method that estimates the occluded facial expression of a person by analyzing non-occluded facial regions, generating an estimated expression image, and combining it with the original image to create a composite image that reveals the person's natural expression.
Enables the generation of composite images that depict individuals without masks, even when they are wearing them, enhancing the attractiveness of photographs taken in crowded environments.
Smart Images

Figure 0007838641000001 
Figure 0007838641000002 
Figure 0007838641000003
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of information processing apparatuses, information processing methods, and recording media.
Background Art
[0002] In a patent document 1, there is described a technique for determining a shielded area shielded in an input image which is an image representing a face, and performing identification of the input image using an area other than an area associated with a shield pattern based on the shielded area, and further improving recognition accuracy of a face image including the shielded area. In a patent document 2, there is described a technique for inputting a face image, detecting areas including parts such as eyes, nose, mouth, cheeks, etc. included in the face image, filling the inside of the detected part areas, and synthesizing an image of the parts stored in advance with the face image in which the part areas are filled. In a patent document 3, there is described a technique for photographing a front image (moving image) of a user through a head-mounted display from the position of a camera fixed to the head-mounted display, using, as it is, a face area not hidden by the head-mounted display in this moving image, and replacing an area hidden by the head-mounted display with an area cut out by a mask pattern of the head-mounted display from a still image photographed in advance from the same viewpoint and stored in an accumulating means, pasting, by a texture mapping technique, the face image synthesized from the moving image and the still image onto the surface of an appropriate solid such as a cube, and outputting or displaying it as a human head.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] This disclosure aims to provide an information processing device, an information processing method, and a recording medium that improve upon the technologies described in prior art documents. [Means for solving the problem]
[0005] One aspect of the information processing apparatus of this disclosure includes: acquisition means for acquiring information about a person, including at least an image of the person; detection means for detecting a face region including the face of the person from the image; estimation means for estimating the occluded region if at least a portion of the face region is occluded; expression estimation means for estimating the expression of the person based on the information about the person; estimated expression image generation means for generating an estimated expression image of the region corresponding to the occluded region, corresponding to the expression estimated by the expression estimation means; and composite image generation means for generating a composite image based on the image and the estimated expression image.
[0006] One aspect of the information processing method of this disclosure includes acquiring information about a person, including at least an image of the person; detecting a face region including the person's face from the image; estimating the occluded region if at least a portion of the face region is occluded; estimating the person's facial expression based on the information about the person; generating an estimated facial expression image corresponding to the occluded region, corresponding to the estimated facial expression; and generating a composite image based on the image and the estimated facial expression image.
[0007] One aspect of the recording medium of this disclosure contains a computer program that causes a computer to execute an information processing method that involves acquiring information about a person, including at least an image of the person; detecting a face region including the person's face from the image; estimating the occluded region if at least a portion of the face region is occluded; estimating the person's facial expression based on the information about the person; generating an estimated facial expression image corresponding to the occluded region in accordance with the estimated facial expression; and generating a composite image based on the image and the estimated facial expression image. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is a block diagram showing the configuration of the information processing device in the first embodiment. [Figure 2] Figure 2 is a block diagram showing the configuration of the information processing device in the second embodiment. [Figure 3] Figure 3 is a flowchart showing the flow of information processing operations performed by the information processing device in the second embodiment. [Figure 4] Figure 4 is a block diagram showing the configuration of the information processing device in the fourth embodiment. [Figure 5] Figure 5 is a flowchart showing the flow of learning operations performed by the information processing device in the fourth embodiment. [Figure 6] Figure 6 is a flowchart showing the flow of the estimated facial expression image generation operation performed by the information processing device in the fifth embodiment. [Figure 7] Figure 7 is a block diagram showing the configuration of the information processing device in the sixth embodiment. [Figure 8] Figure 8 is a conceptual diagram showing an example of display controlled by the information processing device in the sixth embodiment. [Figure 9] Figure 9 is a conceptual diagram of the online meeting system in the seventh embodiment. [Figure 10] Figure 10 is a block diagram showing the configuration of the online conference control device in the seventh embodiment. [Figure 11] Figure 11 is a flowchart showing the flow of online meeting control operations performed by the online meeting control device in the seventh embodiment. [Modes for carrying out the invention]
[0009] The following describes embodiments of the information processing device, information processing method, and recording medium with reference to the drawings. [1: First Embodiment]
[0010] The first embodiment of an information processing apparatus, an information processing method, and a recording medium will be described. Below, the first embodiment of an information processing apparatus, an information processing method, and a recording medium will be described using an information processing apparatus 1 to which the first embodiment of an information processing apparatus, an information processing method, and a recording medium is applied. [1-1: Configuration of Information Processing Apparatus 1]
[0011] Referring to FIG. 1, the configuration of the information processing apparatus 1 in the first embodiment will be described. FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1 in the first embodiment.
[0012] As shown in FIG. 1, the information processing apparatus 1 includes an acquisition unit 11, a detection unit 12, a region estimation unit 13, an expression estimation unit 14, an estimated expression image generation unit 15, and a composite image generation unit 16. The acquisition unit 11 acquires information about the person including at least an image of the person. The detection unit 12 detects a face region including the face of the person from the image. The region estimation unit 13 estimates a shielded region that is shielded when at least a part of the face region is shielded. The expression estimation unit 14 estimates the expression of the person based on the information about the person. The estimated expression image generation unit 15 generates an estimated expression image of a region corresponding to the shielded region according to the expression estimated by the expression estimation unit 14. The composite image generation unit 16 generates a composite image based on the image and the estimated expression image. [1-2: Technical Effects of Information Processing Apparatus 1]
[0013] Since the information processing apparatus 1 in the first embodiment generates a composite image based on the image and the image corresponding to the estimated expression of the person, even when at least a part of the face region of the person is shielded, an image corresponding to the expression of the person (that is, the composite image) in which the face region of the person is not shielded can be acquired. [2: Second Embodiment]
[0014] The second embodiment of the information processing apparatus, the information processing method, and the recording medium will be described. Hereinafter, the second embodiment of the information processing apparatus, the information processing method, and the recording medium will be described using the information processing apparatus 2 to which the second embodiment of the information processing apparatus, the information processing method, and the recording medium is applied. [2-1: Configuration of Information Processing Apparatus 2]
[0015] The configuration of the information processing apparatus 2 in the second embodiment will be described while referring to FIG. 2. FIG. 2 is a block diagram showing the configuration of the information processing apparatus 2 in the second embodiment.
[0016] As shown in FIG. 2, the information processing apparatus 2 includes an arithmetic unit 21 and a storage unit 22. Further, the information processing apparatus 2 may include a communication device 23, an input device 24, and an output device 25. However, the information processing apparatus 2 may not include at least one of the communication device 23, the input device 24, and the output device 25. The arithmetic unit 21, the storage unit 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0017] The arithmetic unit 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic unit 21 reads a computer program. For example, the arithmetic unit 21 may read a computer program stored in the storage device 22. For example, the arithmetic unit 21 may read a computer program stored in a computer-readable and non-temporary recording medium using a recording medium reading device (not shown) provided by the information processing device 2 (for example, an input device 24 described later). The arithmetic unit 21 may obtain a computer program from a device (not shown) located outside the information processing device 2 via a communication device 23 (or other communication device) (i.e., it may download or read the program). The arithmetic unit 21 executes the read computer program. As a result, logical functional blocks for performing the operations that the information processing device 2 should perform are realized within the arithmetic unit 21. In other words, the arithmetic unit 21 can function as a controller for realizing logical functional blocks necessary for the information processing device 2 to perform its operations (in other words, processing).
[0018] Figure 2 shows an example of a logical functional block implemented within the arithmetic unit 21 to perform information processing operations. As shown in Figure 2, the arithmetic unit 21 implements an acquisition unit 211, which is a specific example of the "acquisition means" described in the appendix below; a detection unit 212, which is a specific example of the "detection means" described in the appendix below; a region estimation unit 213, which is a specific example of the "estimation means" described in the appendix below; a facial expression estimation unit 214, which is a specific example of the "facial expression estimation means" described in the appendix below; a facial expression image generation unit 215, which is a specific example of the "estimated facial expression image generation means" described in the appendix below; and a composite image generation unit 216, which is a specific example of the "composite image generation means" described in the appendix below. The operations of each of the acquisition unit 211, detection unit 212, region estimation unit 213, facial expression estimation unit 214, estimated facial expression image generation unit 215, and composite image generation unit 216 will be described later with reference to Figure 3.
[0019] The storage device 22 is capable of storing desired data. For example, the storage device 22 may temporarily store computer programs executed by the arithmetic unit 21. The storage device 22 may temporarily store data that the arithmetic unit 21 uses temporarily when it is executing a computer program. The storage device 22 may store data that the information processing device 2 stores long-term. The storage device 22 may include at least one of the following: RAM (Random Access Memory), ROM (Read Only Memory), hard disk drive, magneto-optical disk drive, SSD (Solid State Drive), and disk array drive. In other words, the storage device 22 may include non-temporary recording media.
[0020] The communication device 23 can communicate with devices outside the information processing device 2 via a communication network (not shown).
[0021] The input device 24 is a device that receives information input to the information processing device 2 from outside the information processing device 2. For example, the input device 24 may include an operating device (e.g., at least one of a keyboard, mouse, and touch panel) that can be operated by an operator of the information processing device 2. For example, the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be attached externally to the information processing device 2.
[0022] The output device 25 is a device that outputs information to the outside of the information processing device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (so-called display) capable of displaying an image representing the information to be output. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (so-called speaker) capable of outputting sound. For example, the output device 25 may output information onto paper. That is, the output device 25 may include a printing device (so-called printer) capable of printing desired information onto paper. [2-2: Information processing operations performed by the information processing device 2]
[0023] Referring to Figure 3, the flow of information processing operations performed by the information processing device 2 in the second embodiment will be explained. Figure 3 is a flowchart showing the flow of information processing operations performed by the information processing device 2 in the second embodiment.
[0024] As shown in Figure 3, the acquisition unit 211 acquires information about the person, including at least an image of the person (step S20). In addition to the image of the person, the acquisition unit 211 may also acquire other information about the person, such as audio information acquired when the image of the person was generated.
[0025] The detection unit 212 detects a face region containing a person's face from the image (step S21). The detection unit 212 may detect the face region by applying a known face detection process to the image. The detection unit 212 may also detect a region having facial features as a face region. A region having facial features may include characteristic parts that constitute a face, such as eyes, nose, and mouth. There are no particular restrictions on the method by which the detection unit 212 detects the face region. For example, the detection unit 212 may detect the face region based on the extraction of edges or patterns characteristic of the face region.
[0026] The detection unit 212 may detect face regions using a neural network that has been trained in machine learning to detect face regions. The detection unit 212 may be composed of a convolutional neural network (CNN).
[0027] The region estimation unit 213 estimates the occluded region if at least a portion of the face region is occluded (step S22). In the second embodiment, the occluded region in which at least a portion of the face region is occluded may be a mask region occluded by a mask worn by the person. The region estimation unit 213 may estimate the occluded mask region if at least a portion of the face region is occluded by a mask worn by the person. For example, the region estimation unit 213 may determine that the face region includes a mask region if no characteristic points such as the nasal wings and corners of the mouth are detected from the face region. The mask region hidden by the mask may be a predetermined region including the nasal wings, corners of the mouth, etc.
[0028] The facial expression estimation unit 214 estimates the person's facial expression based on information about the person (step S23). If at least a portion of the face area is obscured by a mask worn by the person, the facial expression estimation unit 214 may use information obtained from areas other than the mask area as information about the person. In this case, the facial expression estimation unit 214 may estimate the person's facial expression based on information obtained from areas of the face other than the mask area, for example. Alternatively, the facial expression estimation unit 214 may estimate the person's facial expression based on, for example, the angle of the face, the pose the person is taking, and the gesture the person is making, in addition to or instead of information obtained from areas of the face other than the mask area, for example. Alternatively, the facial expression estimation unit 214 may estimate the person's facial expression based on, for example, audio information obtained when the person's image was generated, in addition to or instead of information obtained from the person's image. The audio information may include at least one of the following: information indicating the state of vocalization and information indicating the content of speech. The state of vocalization may include at least one of the following: tone of vocalization and tempo. Furthermore, the facial expression estimation unit 214 may estimate a person's facial expression based, for example, on information indicating the surrounding circumstances when the person's image was generated, in addition to, or instead of, information about the person themselves. The facial expression estimation unit 214 may also adopt information that improves the accuracy of estimating the person's facial expression as information about the person.
[0029] The facial expression estimation unit 214 may estimate a person's facial expression based on predetermined rules, for example. For example, it may estimate a person's facial expression based on the state of movement of the facial muscles. The state of movement of the facial muscles may include at least one of the following: the eyebrows are raised, the eyebrows are lowered, and the cheeks are raised. The facial expression estimation unit 214 may estimate a person's facial expression by combining multiple states of movement of facial muscles. The facial expression estimation unit 214 may estimate a person's facial expression to be at least one of the following: an expression of joy, an expression of surprise, an expression of fear, an expression of disgust, an expression of anger, an expression of sadness, and a neutral expression. For example, the facial expression estimation unit 214 may estimate a person's facial expression to be joyful if their cheeks are raised above a predetermined level.
[0030] Furthermore, in the second embodiment, the example given was that the shielded area, in which at least a portion of the face area is shielded, is the mask area shielded by a mask worn on the face. However, the shielded area may also be, for example, the area shielded by sunglasses. In this case, the facial expression estimation unit 214 may estimate the person's facial expression from the state of the mouth. The state of the mouth may include, for example, at least one of the following: the upper lip is raised, the corners of the mouth are raised, dimples are formed, and the chin is raised.
[0031] The estimated facial expression image generation unit 215 generates an estimated facial expression image for the region corresponding to the occluded area, corresponding to the facial expression estimated by the facial expression estimation unit 214 (step S24).
[0032] The composite image generation unit 216 generates a composite image based on the image and the estimated facial expression image. The composite image generation unit 216 may generate the composite image such that at least the occluded area is hidden by the estimated facial expression image. That is, the composite image generation unit 216 may fill in the occluded area of the person's face with an image corresponding to the estimated facial expression of the person. [2-3: Technical Effects of Information Processing Device 2]
[0033] In the second embodiment, the information processing device 2 generates a composite image based on an image and an image of a mask region corresponding to the estimated facial expression of a person. Therefore, even when a person is wearing a mask, it is possible to obtain an image corresponding to the person's facial expression in which the person's mouth is not obscured.
[0034] In recent years, due to changes in hygiene awareness, wearing masks is recommended, especially in crowded places. However, when taking commemorative photos in crowded places, such as tourist spots, the photos often end up being dominated by faces wearing masks, resulting in uninteresting pictures. In other words, even in crowded places where removing a mask is discouraged, there is a need to record natural-looking images of faces without masks.
[0035] In contrast, the information processing device 2 in the second embodiment generates a composite image of a person without a mask based on an image of the area corresponding to the mask area, corresponding to the estimated facial expression of the person, when the person is wearing a mask. Therefore, it can provide a natural face image without a mask. Consequently, in photographs taken in crowded places, a natural face image without a mask will be included, making it possible to record an attractive photograph. [3: Third Embodiment]
[0036] A third embodiment of the information processing device, information processing method, and recording medium will be described below. In the following description, the third embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 3 to which the third embodiment of the information processing device, information processing method, and recording medium is applied.
[0037] In the third embodiment, if at least a portion of the facial region is obscured by a mask worn by the face, the facial expression estimation unit 214 may estimate the person's facial expression based on the region around the person's eyes in the facial region, as a region other than the masked region included in the facial region. The facial expression estimation unit 214 may also estimate the person's facial expression based on information obtainable from the region around the eyes included in the facial region.
[0038] The facial expression estimation unit 214 may, for example, extract the area around the eyes from the face region based on the distance between the two eyes included in the face. Alternatively, the facial expression estimation unit 214 may extract the area around the eyes from the face region based on both sides of the lower part of the bridge of the nose included in the face.
[0039] Furthermore, the facial expression estimation unit 214 may estimate a person's facial expression based on, for example, the angle of the face or the pose / gesture taken by the person, in addition to the area information around the eyes included in the face region. Alternatively, the facial expression estimation unit 214 may estimate a person's facial expression based on, for example, audio information acquired when the person's image was generated, in addition to the area information around the eyes included in the face region. Furthermore, the facial expression estimation unit 214 may estimate a person's facial expression based on, in addition to the area information around the eyes included in the face region, information indicating the surrounding circumstances when the person's image was generated. Similar to the second embodiment, the facial expression estimation unit 214 may employ information that improves the accuracy of estimating the person's facial expression as information about the person. [Technical effects of the information processing device 3]
[0040] In the third embodiment, the information processing device 3 can estimate the facial expression under the mask from image information around the eyes and synthesize an appropriate facial image of a face without a mask. [4: Fourth Embodiment]
[0041] A fourth embodiment of the information processing device, information processing method, and recording medium will be described below. In the following description, the fourth embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 4 to which the fourth embodiment of the information processing device, information processing method, and recording medium is applied. [4-1: Configuration of Information Processing Device 4]
[0042] The configuration of the information processing device 4 in the fourth embodiment will be described with reference to Figure 4. Figure 4 is a block diagram showing the configuration of the information processing device 4 in the fourth embodiment.
[0043] As shown in Figure 4, the information processing device 4 in the fourth embodiment includes an arithmetic unit 21 and a storage device 22, similar to the information processing device 2 in the second embodiment and the information processing device 3 in the third embodiment. Furthermore, the information processing device 4 may also include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment and the information processing device 3 in the third embodiment. However, the information processing device 4 does not have to include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 4 in the fourth embodiment differs from the information processing device 2 in the second embodiment and the information processing device 3 in the third embodiment in that the arithmetic unit 21 includes a learning unit 417 and performs learning operations. Other features of the information processing device 4 may be the same as other features of at least one of the information processing device 2 in the second embodiment and the information processing device 3 in the third embodiment. [4-2: Learning operations performed by the information processing device 4]
[0044] Referring to Figure 5, the flow of information processing operations performed by the information processing device 4 in the fourth embodiment will be explained. Figure 5 is a flowchart showing the flow of information processing operations performed by the information processing device 4 in the fourth embodiment.
[0045] As shown in Figure 5, the acquisition unit 211 acquires learning information including sample information about a sample person with a predetermined facial expression and facial expression labels indicating the predetermined facial expression (step S40). The predetermined facial expression may include at least one of the following: a joyful expression, a surprised expression, a frightened expression, a disgusted expression, an angry expression, a sad expression, and a neutral expression. The facial expression labels may be labels indicating each of these expressions. Furthermore, labels may be provided for each of the multiple intensity levels of each facial expression.
[0046] The acquisition unit 211 may acquire learning information stored in the storage device 22 from the storage device 22. The acquisition unit 211 may also acquire learning information from an external device via the communication device 23.
[0047] The detection unit 212 detects a face region containing a person's face from the image (step S21). The facial expression estimation unit 214 estimates the facial expression of the sample person based on the sample information (step S41).
[0048] The learning unit 417 causes the expression estimation unit 214 to learn a method for estimating a person's facial expression based on the facial expression labels and the results of the facial expression estimation unit 214 for a sample person (step S42). The learning unit 417 may construct a facial expression estimation model that can estimate the facial expression of a person whose facial region is obscured by at least a portion of it. The facial expression estimation unit 214 may use the facial expression estimation model to estimate the facial expression of a person whose facial region is obscured by at least a portion of it, based on information about the person. By using the trained facial expression estimation model, the facial expression estimation unit 214 can accurately estimate the facial expression of a person whose facial region is obscured by at least a portion of it.
[0049] The parameters that define the operation of the facial expression estimation model may be stored in the memory device 22. The parameters that define the operation of the facial expression estimation model may also be parameters that are updated by the learning process, such as the weights and biases of a neural network.
[0050] Images used for learning facial expressions obscured by a mask only need to show the state of the person outside the masked area. In other words, learning can be performed using areas outside the masked region. That is, the images used for learning may be images of people wearing masks or images of people not wearing masks. [4-3: Technical effects of the information processing device 4]
[0051] In the fourth embodiment, the information processing device 4 can achieve highly accurate estimation of a person's facial expression through machine learning. [5: Fifth Embodiment]
[0052] A fifth embodiment of the information processing device, information processing method, and recording medium will be described below. The fifth embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 5 to which the fifth embodiment of the information processing device, information processing method, and recording medium is applied.
[0053] The information processing device 5 according to the fifth embodiment will be described with reference to Figure 6. The fifth embodiment describes a specific example of the operation during the generation of estimated facial expression images in the second to fourth embodiments described above (i.e., the operation corresponding to step S24 in Figure 3). In the fifth embodiment, the storage device 22 may have images of people with various facial expressions, at least images of people whose occluded area is not occluded, pre-registered. Other parts of the operation during the generation of estimated facial expression images may be the same as at least one of the second to fourth embodiments. For this reason, the parts that differ from each embodiment already described will be described in detail below, and other overlapping parts will be omitted as appropriate. [5-1: Estimated facial expression image generation operation performed by the information processing device 5]
[0054] Referring to Figure 6, the flow of estimated facial expression image generation by the information processing device 5 according to the fifth embodiment (i.e., the operation when generating an estimated facial expression image) will be explained. Figure 6 is a flowchart showing the flow of the estimated facial expression image generation operation by the information processing device 5 according to the fifth embodiment.
[0055] As shown in Figure 6, the estimated facial expression image generation unit 215 estimates who the person to be processed is (step S50). The estimated facial expression image generation unit 215 may also perform facial recognition using the face region detected by the detection unit 212 to estimate who the person to be processed is.
[0056] The estimated facial expression image generation unit 215 searches for and acquires an image that is estimated to be the image of the person to be processed (hereinafter sometimes referred to as "the person") from among the pre-registered images of people in which at least the occluding area is not occluded (step S51). In step S51, the estimated facial expression image generation unit 215 determines whether or not an image of the person has been acquired (step S52).
[0057] If an image of the person is obtained in step S51 (step S52: Yes), the estimated facial expression image generation unit 215 determines whether or not there is an image of the person with a facial expression corresponding to the facial expression estimated in step S23 (step S53). The facial expressions corresponding to the estimated facial expression may include facial expressions that match or are similar to the estimated facial expression.
[0058] If there is an image of the person with the expression corresponding to the expression estimated in step S23 (step S53: Yes), the estimated expression image generation unit 215 generates an estimated expression image based on a pre-registered image of the person with the expression corresponding to the expression estimated by the expression estimation unit 214 (step S54). The estimated expression image generation unit 215 may also select a pre-registered image of the person with the expression corresponding to the expression estimated by the expression estimation unit 214, and generate an estimated expression image by correcting the brightness of the image, the posture of the person, etc.
[0059] If there is no image of the person with the expression corresponding to the expression estimated in step S23 (step S53: No), the estimated expression image generation unit 215 generates an estimated expression image based on a pre-registered image of the person in which at least the occluded area is not occluded (step S55). If there is no pre-registered image of the person with the expression corresponding to the expression estimated by the expression estimation unit 214, the estimated expression image generation unit 215 may select any image of the person, convert the expression in the image to the expression corresponding to the expression estimated by the expression estimation unit 214, and generate an estimated expression image. The estimated expression image generation unit 215 may also apply deep learning techniques, such as a Generative Adversarial Network (GAN), to generate an image of the expression in the image that corresponds to the expression estimated by the expression estimation unit 214 as an estimated expression image.
[0060] If, in step S51, an image of the person could not be obtained (step S52: No), the estimated facial expression image generation unit 215 may, for example, apply deep learning techniques such as GAN to generate an image of the facial expression corresponding to the facial expression estimated by the facial expression estimation unit 214 as the estimated facial expression image (step S56).
[0061] Furthermore, only one image of the person may be registered per person. In other words, the information image generation unit 215 may omit the operation in step S53 and perform the operation in step S55. Also, the estimated facial expression image generation unit 215 may generate an estimated facial expression image by applying deep learning techniques such as GAN, regardless of whether or not there is an image of the person. In other words, the information image generation unit 215 may omit the operations in steps S50 to S52 and perform the operation in step S56.
[0062] Furthermore, the images generated in this embodiment do not necessarily have to be intended for use in person authentication. Therefore, the estimated facial expression image generation unit 215 may generate a facial image with an expression that is more appropriate to the person's situation at the time the image was generated, rather than focusing on individual characteristics. [5-2: Technical Effects of Information Processing Device 5]
[0063] In the fifth embodiment, the information processing device 5 generates the estimated facial expression image based on a pre-registered image of a person in which at least the occluding area is not obscured, thereby obtaining an image that looks like the person in question. Furthermore, if an image of a person in which at least the occluding area of the facial expression corresponding to the estimated facial expression is pre-registered, the information processing device 5 generates the estimated facial expression image based on that pre-registered image, thereby obtaining an image that looks even more like the person in question. [6: Sixth Embodiment]
[0064] A sixth embodiment of the information processing device, information processing method, and recording medium will be described below. In the following description, the sixth embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 6 to which the sixth embodiment of the information processing device, information processing method, and recording medium is applied. [6-1: Configuration of Information Processing Device 6]
[0065] The configuration of the information processing device 6 in the sixth embodiment will be described with reference to Figure 7. Figure 7 is a block diagram showing the configuration of the information processing device 6 in the sixth embodiment.
[0066] As shown in Figure 7, the information processing device 6 in the sixth embodiment includes a arithmetic unit 21 and a storage device 22, similar to the information processing devices 2 in the second embodiment to 5 in the fifth embodiment. Furthermore, the information processing device 6 may also include a communication device 23, an input device 24, and an output device 25, similar to the information processing devices 2 in the second embodiment to 5 in the fifth embodiment. However, the information processing device 6 does not have to include at least one of the communication device 23, input device 24, and output device 25. The information processing device 6 in the sixth embodiment differs from the information processing devices 2 in the second embodiment to 5 in the sixth embodiment in that the arithmetic unit 21 includes a display control unit 618. Other features of the information processing device 6 may be the same as at least one other feature of the information processing device 2 in the second embodiment to 5 in the fifth embodiment. [6-2: Information processing operations performed by the information processing device 6]
[0067] When the composite image generation unit 216 generates a composite image, the display control unit 618 displays the composite image instead of the original image and overlays information indicating that the image was generated by the composite image generation unit 216 onto the composite image. For example, as illustrated in Figure 8(a), when the composite image generation unit 216 generates a composite image, the display control unit 618 may display text such as "Mask Area Completion Image" in the lower right corner of the display mechanism D. Alternatively, as illustrated in Figure 8(b), when the composite image generation unit 216 generates a composite image, the display control unit 618 may overlay a semi-transparent mask onto the area corresponding to the mask area in the uncomposite image. [6-3: Technical Effects of Information Processing Device 6]
[0068] In the sixth embodiment, when the information processing device 6 displays a composite image, it superimposes information indicating that it is a composite image onto the composite image, so that the user can easily distinguish whether or not the image is a composite image. [7: Seventh Embodiment]
[0069] A seventh embodiment of the online meeting system will be described below. In the following description, the seventh embodiment of the online meeting system will be described using an online meeting system 700 to which the seventh embodiment of the online meeting system is applied. [7-1: Configuration of Online Meeting System 700] As illustrated in Figure 9, the online meeting system 700 in the seventh embodiment may include an online meeting control device 7 and a plurality of terminals 70 that conduct meetings (in Figure 9, terminals 70-1, 70-2, 70-3, ..., terminal 70-N are shown as examples). The online meeting control device 7 can communicate with the plurality of terminals 70. The plurality of terminals 70 may conduct online meetings. The plurality of terminals 70 may conduct web conferences. [7-2: Configuration of Online Meeting Control Device 7]
[0070] The configuration of the online meeting control device 7 will be described with reference to Figure 10. Figure 10 is a block diagram showing the configuration of the online meeting control device 7 in the seventh embodiment.
[0071] As shown in Figure 10, the online conference control device 7 comprises a computing device 71 and a storage device 72. Furthermore, the online conference control device 7 may also comprise a communication device 73, an input device 74, and an output device 75. However, the online conference control device 7 does not have to comprise at least one of the communication device 73, the input device 74, and the output device 75. The computing device 71, the storage device 72, the communication device 73, the input device 74, and the output device 75 may be connected via a data bus 76.
[0072] The arithmetic unit 71 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic unit 71 reads computer programs. For example, the arithmetic unit 71 may read computer programs stored in the storage device 72. For example, the arithmetic unit 71 may read computer programs stored on a computer-readable and non-temporary recording medium using a recording medium reading device (not shown) provided by the online meeting control device 7 (for example, an input device 74 described later). The arithmetic unit 71 may obtain computer programs from a device (not shown) located outside the online meeting control device 7 via a communication device 73 (or other communication device) (i.e., it may download or read them). The arithmetic unit 71 executes the read computer programs. As a result, logical functional blocks for performing the operations that the online meeting control device 7 should perform are realized within the arithmetic unit 71. In other words, the computing unit 71 can function as a controller for realizing logical functional blocks necessary to perform the operations (in other words, processes) that the online meeting control device 7 should perform.
[0073] Figure 10 shows an example of a logical functional block implemented within the computing unit 71 to perform online meeting control operations. As shown in Figure 10, the computing unit 71 implements an acquisition unit 711, which is a specific example of the "acquisition means" described in the appendix below; a detection unit 712, which is a specific example of the "detection means" described in the appendix below; a region estimation unit 713, which is a specific example of the "estimation means" described in the appendix below; a facial expression estimation unit 714, which is a specific example of the "facial expression estimation means" described in the appendix below; an estimated facial expression image generation unit 715, which is a specific example of the "estimated facial expression image generation means" described in the appendix below; a composite image generation unit 716, which is a specific example of the "composite image generation means" described in the appendix below; and an output control unit 719, which is a specific example of the "output control means". The operations of the acquisition unit 711, detection unit 712, region estimation unit 713, facial expression estimation unit 714, estimated facial expression image generation unit 715, composite image generation unit 716, and output control unit 719 will be described later with reference to Figure 11.
[0074] The storage device 72 is capable of storing desired data. For example, the storage device 72 may temporarily store a computer program executed by the arithmetic unit 71. The storage device 72 may temporarily store data that the arithmetic unit 71 temporarily uses when it is executing a computer program. The storage device 72 may store data that the online meeting control device 7 stores long-term. The storage device 72 may include at least one of the following: RAM (Random Access Memory), ROM (Read Only Memory), hard disk drive, magneto-optical disk drive, SSD (Solid State Drive), and disk array device. In other words, the storage device 72 may include a non-temporary recording medium.
[0075] The communication device 73 can communicate with devices outside the online meeting control device 7 via a communication network (not shown). The online meeting control device 7 may also be able to communicate with each of the multiple terminals 70 via the communication device 73.
[0076] The input device 74 is a device that receives information input to the online meeting control device 7 from outside the online meeting control device 7. For example, the input device 74 may include an operating device (e.g., at least one of a keyboard, mouse, and touch panel) that can be operated by the operator of the online meeting control device 7. For example, the input device 74 may include a reader that can read information recorded as data on an external recording medium that can be attached to the online meeting control device 7.
[0077] The output device 75 is a device that outputs information to the outside of the online meeting control device 7. For example, the output device 75 may output information as an image. That is, the output device 75 may include a display device (so-called display) capable of displaying an image that shows the information to be output. For example, the output device 75 may output information as sound. That is, the output device 75 may include an audio device (so-called speaker) capable of outputting sound. For example, the output device 75 may output information on paper. That is, the output device 75 may include a printing device (so-called printer) capable of printing the desired information on paper. [7-3: Online meeting control operations performed by the online meeting control device 7]
[0078] Referring to Figure 11, the flow of online meeting control operations performed by the online meeting control device 7 in the seventh embodiment will be explained. Figure 11 is a flowchart showing the flow of online meeting control operations performed by the online meeting control device 7 in the seventh embodiment.
[0079] As shown in Figure 11, the acquisition unit 711 acquires information about a person, including at least an image of the person, from at least one terminal 70 among the multiple terminals 70 conducting the conference (step S70). The acquisition unit 711 may also acquire information about the person, including at least an image of the person operating the terminal 70. The acquisition unit 711 may also acquire information about the person, including a video of the person operating the terminal 70.
[0080] The detection unit 712 detects a face region containing a person's face from the image (step S71). The region estimation unit 713 estimates the occluded region if at least a portion of the face region is occluded (step S72). The facial expression estimation unit 714 estimates the person's facial expression based on information about the person (step S73). The estimated facial expression image generation unit 715 generates an estimated facial expression image for the region corresponding to the occluded region, according to the facial expression estimated by the facial expression estimation unit 714 (step S74). The composite image generation unit 716 generates a composite image based on the image and the estimated facial expression image (step S75).
[0081] Furthermore, the operations performed by the detection unit 712 may be the same as those performed by at least one of the detection units 212 in the second to sixth embodiments. Also, the operations performed by the region estimation unit 713 may be the same as those performed by at least one of the region estimation units 213 in the second to sixth embodiments. Also, the operations performed by the facial expression estimation unit 714 may be the same as those performed by at least one of the facial expression estimation units 214 in the second to sixth embodiments. Also, the operations performed by the estimated facial expression image generation unit 715 may be the same as those performed by at least one of the estimated facial expression image generation units 215 in the second to sixth embodiments. Also, the operations performed by the composite image generation unit 716 may be the same as those performed by at least one of the composite image generation units 216 in the second to sixth embodiments.
[0082] When the composite image generation unit 716 generates a composite image, the output control unit 719 outputs the composite image to the multiple terminals 70 instead of the image (step S76). When the acquisition unit 711 acquires a video of a person operating the terminal 70, the output control unit 719 may output the image or composite image to the multiple terminals 70 in real time. Alternatively, when the output control unit 719 outputs a composite image to the multiple terminals 70, it may output it at a slower rate compared to when it outputs an image to the multiple terminals 70. When the output control unit 719 outputs a composite image to the multiple terminals 70, it may output it with a delay of, for example, several seconds, compared to when it outputs an image to the multiple terminals 70.
[0083] Furthermore, in at least one of the information processing devices 2 in the second embodiment to 6 in the sixth embodiment, the composite image generation operation may be performed in real time. Alternatively, in at least one of the information processing devices 2 in the second embodiment to 6 in the sixth embodiment, a time lag of, for example, several seconds may occur.
[0084] Furthermore, if the acquisition unit 711 acquires a still image of a person operating the terminal 70, the learning unit 717 may generate a composite image offline, and the output control unit 719 may output the offline-generated composite image to multiple terminals 70.
[0085] Furthermore, if the acquisition unit 711 acquires information about a person, including a video of the person, the region estimation unit 713 does not need to perform estimation processing for each frame. In other words, the region estimation unit 713 may perform estimation processing every predetermined number of frames. That is, the facial expression estimation unit 714 may generate estimated facial expression images corresponding to the same facial expression for a predetermined number of frames.
[0086] Furthermore, in the seventh embodiment, the online meeting control device 7 may have a learning unit 717 in the arithmetic unit 71. That is, similar to the learning unit 417 in the fourth embodiment, the learning unit 717 may cause the facial expression estimation unit 714 to learn a method for estimating a person's facial expression based on facial expression labels and the results of facial expression estimation of a sample person by the facial expression estimation unit 714.
[0087] Furthermore, in the seventh embodiment, the online meeting control device 7 may have a display control unit 718 in the arithmetic unit 71. That is, similar to the display control unit 618 in the sixth embodiment, the display control unit 718 may display the composite image instead of the image when the composite image generation unit 716 generates a composite image, and may superimpose information indicating that the image was generated by the composite image generation unit 716 onto the composite image. [7-4: Technical Effects of Online Meeting Control Device 7]
[0088] In the seventh embodiment, the online meeting control device 7 generates a composite image based on an image and an image of the mask region corresponding to the estimated facial expression of the person. Therefore, even when a person is wearing a mask, it is possible to obtain an image corresponding to the person's facial expression in which the person's mouth is not obscured.
[0089] Recently, due to changes in hygiene awareness, wearing masks is recommended, especially in crowded places. While some people prefer to participate in online communication without wearing a mask, wearing one is recommended when participating from shared spaces such as satellite offices. In other words, there is a demand for the distribution of natural-looking facial images without masks, even in places where removing a mask is discouraged, such as crowded places.
[0090] In contrast, the online meeting control device 7 in the seventh embodiment generates a composite image of a person without a mask based on an image of the area corresponding to the mask area, corresponding to the estimated facial expression of the person, when the person is wearing a mask. Therefore, it can provide a natural-looking face image without a mask. Consequently, even when participating from a shared location such as a satellite office, a natural-looking face image without a mask can be delivered. [8: Addendum]
[0091] The following additional information is disclosed regarding the embodiments described above. [Note 1] An acquisition means for acquiring information about a person, including at least an image of the person, A detection means for detecting a facial region including the face of the person from the aforementioned image, If at least a portion of the facial region is occluded, an estimation means for estimating the occluded region is provided. A facial expression estimation means for estimating the facial expression of the person based on information about the person, Estimated facial expression image generation means for generating an estimated facial expression image for a region corresponding to the occluded region, corresponding to the facial expression estimated by the facial expression estimation means, A composite image generation means for generating a composite image based on the aforementioned image and the estimated facial expression image. An information processing device equipped with the following features. [Note 2] The shielded area, in which at least a portion of the facial area is shielded, is the masked area that is shielded by the mask worn by the person. The information processing device described in Appendix 1. [Note 3] The facial expression estimation means estimates the facial expression of the person based on the area around the person's eyes in the facial region. The information processing device described in Appendix 2. [Note 4] The acquisition means acquires learning information including sample information relating to a sample person with a predetermined facial expression and an expression label indicating the predetermined facial expression. The facial expression estimation means estimates the facial expression of the sample person based on the sample information, The system further comprises a learning means for causing the facial expression estimation means to learn a method for estimating the facial expressions of the person, based on the facial expression labels and the results of the facial expression estimation means for the sample person. An information processing device as described in any one of the appendices 1 to 3. [Note 5] The estimated facial expression image generation means generates the estimated facial expression image based on a pre-registered image of the person in which at least the occluded area is not occluded. An information processing device as described in any one of the appendices 1 to 3. [Note 6] The estimated facial expression image generation means generates the estimated facial expression image based on pre-registered images of the person corresponding to the facial expression estimated by the facial expression estimation means. The information processing device described in Appendix 5. [Note 7] When the composite image generation means generates the composite image, the display control means further includes a means for displaying the composite image in place of the original image and superimposing information indicating that the image was generated by the composite image generation means onto the composite image. An information processing device as described in any one of the appendices 1 to 3. [Note 8] An acquisition means for acquiring information about a person, including at least an image of the person, from at least one terminal among multiple terminals used for conducting a meeting, A detection means for detecting a facial region including the face of the person from the aforementioned image, If at least a portion of the facial region is occluded, an estimation means for estimating the occluded region is provided. A facial expression estimation means for estimating the facial expression of the person based on information about the person, Estimated facial expression image generation means for generating an estimated facial expression image for a region corresponding to the occluded region, corresponding to the facial expression estimated by the facial expression estimation means, A composite image generation means that generates a composite image based on the aforementioned image and the estimated facial expression image, When the composite image generation means generates the composite image, the output control means outputs the composite image to the multiple terminals in place of the original image. An online meeting system equipped with [features / equipment]. [Note 9] We obtain information about the person in question, including at least an image of the person. From the aforementioned image, the facial region including the person's face is detected. If at least a portion of the facial region is occluded, the occluded region is estimated. Based on the information about the person, estimate the person's facial expression. An estimated facial expression image is generated for the region corresponding to the occluded area, based on the estimated facial expression. A composite image is generated based on the aforementioned image and the estimated facial expression image. Information processing methods. [Note 10] On the computer, We obtain information about the person in question, including at least an image of the person. From the aforementioned image, the facial region including the person's face is detected. If at least a portion of the facial region is occluded, the occluded region is estimated. Based on the information about the person, estimate the person's facial expression. An estimated facial expression image is generated for the region corresponding to the occluded area, based on the estimated facial expression. A composite image is generated based on the aforementioned image and the estimated facial expression image. A recording medium on which computer programs for executing information processing methods are stored.
[0092] This disclosure may be modified as appropriate, insofar as it does not contradict the gist or idea of the invention as can be inferred from the claims and the specification as a whole, and information processing devices, information processing methods, and recording media with such modifications are also included in the technical idea of this disclosure. [Explanation of symbols]
[0093] 1,2,3,4,5,6 Information Processing Device 11,211,711 Acquisition Department 12,212,712 detection unit 13,213,713 Area estimation part 14,214,714 Facial expression estimation part 15,215,715 Estimated facial expression image generation unit 16,216,716 Composite Image Generation Unit 417,717 Learning Department 618,718 Display Control Unit 700 Online Meeting Systems 7 Online meeting control device 70 devices 719 Output Control Unit
Claims
1. An acquisition means for acquiring information about a person, including at least an image of the person, A detection means for detecting a facial region including the face of the person from the aforementioned image, If at least a portion of the facial region is occluded, an estimation means for estimating the occluded region is provided. A facial expression estimation means for estimating the facial expression of the person based on information about the person, Estimated facial expression image generation means for generating an estimated facial expression image for a region corresponding to the occluded region, corresponding to the facial expression estimated by the facial expression estimation means, A composite image generation means for generating a composite image based on the aforementioned image and the estimated facial expression image. Equipped with, The shielded area, in which at least a portion of the facial area is shielded, is the masked area that is shielded by the mask worn by the person. The facial expression estimation means estimates the facial expression of the person based on the area around the person's eyes in the facial region. Information processing device.
2. The acquisition means acquires learning information including sample information relating to a sample person with a predetermined facial expression and an expression label indicating the predetermined facial expression. The facial expression estimation means estimates the facial expression of the sample person based on the sample information, The system further comprises a learning means for causing the facial expression estimation means to learn a method for estimating the facial expressions of the person, based on the facial expression labels and the results of the facial expression estimation means for the sample person. The information processing apparatus according to claim 1.
3. The estimated facial expression image generation means generates the estimated facial expression image based on a pre-registered image of the person in which at least the occluded area is not occluded. The information processing apparatus according to claim 1.
4. The estimated facial expression image generation means generates the estimated facial expression image based on pre-registered images of the person corresponding to the facial expression estimated by the facial expression estimation means. The information processing apparatus according to claim 3.
5. When the composite image generation means generates the composite image, the display control means further includes a means for displaying the composite image in place of the original image and superimposing information indicating that the image was generated by the composite image generation means onto the composite image. The information processing apparatus according to claim 1.
6. An acquisition means for acquiring information about a person, including at least an image of the person, from at least one terminal among multiple terminals used for conducting a meeting, A detection means for detecting a facial region including the face of the person from the aforementioned image, If at least a portion of the facial region is occluded, an estimation means for estimating the occluded region is provided. A facial expression estimation means for estimating the facial expression of the person based on information about the person, Estimated facial expression image generation means for generating an estimated facial expression image for a region corresponding to the occluded region, corresponding to the facial expression estimated by the facial expression estimation means, A composite image generation means that generates a composite image based on the aforementioned image and the estimated facial expression image, When the composite image generation means generates the composite image, the output control means outputs the composite image to the multiple terminals in place of the original image. Equipped with, The shielded area, in which at least a portion of the facial area is shielded, is the masked area that is shielded by the mask worn by the person. The facial expression estimation means estimates the facial expression of the person based on the area around the person's eyes in the facial region. Online meeting system.
7. This involves obtaining information about the person in question, including at least an image of the person, To detect the facial region including the person's face from the aforementioned image, If at least a portion of the facial region is occluded, estimate the occluded region. To estimate the facial expression of the person based on the information about the person, An estimated facial expression image is generated for the region corresponding to the occluded area, based on the estimated facial expression. A composite image is generated based on the aforementioned image and the estimated facial expression image. Includes, The shielded area, in which at least a portion of the facial area is shielded, is the masked area that is shielded by the mask worn by the person. Estimating the aforementioned facial expression includes estimating the facial expression of the person based on the area around the person's eyes in the facial region. The information processing method performed by computers.
8. On the computer, This involves obtaining information about the person in question, including at least an image of the person, To detect the facial region including the person's face from the aforementioned image, If at least a portion of the facial region is occluded, estimate the occluded region. To estimate the facial expression of the person based on the information about the person, An estimated facial expression image is generated for the region corresponding to the occluded area, based on the estimated facial expression. A composite image is generated based on the aforementioned image and the estimated facial expression image. Includes, The shielded area, in which at least a portion of the facial area is shielded, is the masked area that is shielded by the mask worn by the person. Estimating the aforementioned facial expression includes estimating the facial expression of the person based on the area around the person's eyes in the facial region. A computer program that executes an information processing method.
Citation Information
Patent Citations
Method and device for synthesizing facial image of person wearing head mount display
JP1999096366A
Apparatus, method and program for image processing, and recording medium on which image processing program is recorded
JP2002352258A
Method and system for reconstructing occluded face parts in a virtual reality environment
JP2017534096A
Image analysis apparatus, image analysis method, and image analysis program
JP2018151919A
Image processing apparatus, camera apparatus, and image processing method
JP2020048149A