Image processing device, image processing method, and image processing system

The image processing system uses computer-generated avatars to replace audience areas with estimated emotions and attributes, addressing privacy concerns and enhancing the event atmosphere for viewers.

JP7787238B2Active Publication Date: 2025-12-16FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024102988
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2025-12-16
Estimated Expiration
2039-11-13

AI Technical Summary

Technical Problem

As high-definition image distribution systems risk violating the privacy of individual spectators by allowing their identification, there is a need to convey the atmosphere of an event while protecting spectator privacy.

Method used

An image processing system that generates images by replacing audience areas with computer-generated avatars reflecting estimated emotions and attributes, synthesizing these with real-life images, and allows dynamic viewing through a head-mounted display.

Benefits of technology

Protects spectator privacy while effectively conveying the venue's atmosphere and emotions, enhancing the experience for viewers, especially in large venues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007787238000001
    Figure 0007787238000001
  • Figure 0007787238000002
    Figure 0007787238000002
  • Figure 0007787238000003
    Figure 0007787238000003
Patent Text Reader

Abstract

To provide an image processing device, an image processing method, and an image processing system capable of transmitting an atmosphere in a site while protecting privacy of a person at the site.SOLUTION: An image processing device 100 includes: a captured image input unit 111 for inputting a captured image; an emotion estimation unit 112 for estimating emotion of an audience based on the captured image; a representative emotion determination unit 113 for dividing an audience area into a plurality of areas to determine emotion representative of each divided area; a CG image generating unit 114 for generating a CG image that is the CG image of the audience area and generates the CG image in which the audience is represented by an avatar; and a composite image generation unit 115 for synthesizing a CG image in a section of the audience area of the captured image to generate a composite image. The CG image of the audience area is generated by arranging one avatar in each divided area. The avatar arranged in each divided area reflects emotion representing each divided area.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an image processing method, and an image processing system. [Background technology]

[0002] Patent Document 1 describes a system that distributes images captured at an event such as a concert in real time, allowing users (viewers) receiving the content to freely change their field of view to view the images.

[0003] Furthermore, Patent Document 2 describes a system that distributes images captured at an event in real time, in which avatars of viewers are displayed on a display installed at the venue in the manner of spectators. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2016 / 009865 [Patent Document 2] Japanese Patent Application Laid-Open No. 2013-020389 Summary of the Invention

[0005] One embodiment of the technique of the present disclosure provides an image processing device, an image processing method, and an image processing system that can convey the atmosphere of a scene while protecting the privacy of people present at the scene. [Means for solving the problem]

[0006] (1) An image processing device comprising: a first image input unit that inputs a first image including a specific area; a first estimation unit that estimates the facial expression and / or emotion of a person in the specific area based on the first image; a second image generation unit that generates as a second image an image of the specific area in which the person is represented by an avatar, in which at least the facial expression and / or emotion estimated by the first estimation unit is reflected in the avatar; and a third image generation unit that generates a third image by synthesizing the second image with the specific area of ​​the first image.

[0007] (2) The image processing device of (1), wherein the second image generation unit divides the specific area into multiple areas, places one avatar in each divided area, and generates the second image.

[0008] (3) An image processing device according to (2), further comprising a first determination unit that determines a facial expression and / or emotion representing the divided area based on the estimation result by the first estimation unit, and a second image generation unit that generates a second image by reflecting the facial expression and / or emotion determined by the first determination unit in the avatar of each divided area.

[0009] (4) The image processing device of (3), wherein the first determination unit determines a facial expression and / or emotion that represents the divided area based on standard values ​​of facial expressions and / or emotions of people belonging to the divided area.

[0010] (5) An image processing device according to any one of (2) to (4), further comprising a second estimation unit that estimates attributes of people in a specific area based on the first image, and a second determination unit that determines attributes representing the divided area based on the estimation results by the second estimation unit, wherein the second image generation unit generates the second image by reflecting the attributes determined by the second determination unit in the avatars of each divided area.

[0011] (6) The image processing device of (5), wherein the attributes include at least one of age and gender.

[0012] (7) An image processing device according to any one of (2) to (6), in which the specific area is divided into divided areas according to the number of people.

[0013] (8) The image processing device of (1), wherein the second image generation unit generates a second image by reflecting the facial expressions and / or emotions of each person estimated by the first estimation unit onto multiple avatars for each person.

[0014] (9) An image processing device according to any one of (1) to (8), wherein the first image is an image captured over a 360° range.

[0015] (10) The image processing device according to any one of (1) to (9), wherein the first image is an image taken of an event venue, and the specific area is an area in the event venue where spectators are present.

[0016] (11) The image processing device according to any one of (1) to (10), wherein the first estimation unit quantifies the degree of each of a plurality of types of facial expressions and / or emotions to estimate the facial expressions and / or emotions.

[0017] (12) An image processing system comprising an image processing device according to any one of (1) to (11) and a playback device for playing back a third image generated by the image processing device, wherein the playback device comprises a third image input unit for inputting the third image, a fourth image generation unit for cutting out a part of the third image to generate a fourth image for display, an instruction unit for instructing switching of the display range, and a fourth image output unit for outputting the fourth image, wherein the fourth image generation unit switches the range for cutting out an image from the third image in accordance with instructions from the instruction unit to generate the fourth image.

[0018] (13) An image processing system according to (12), wherein the playback device is a head-mounted display and includes a detection unit that detects the movement of the main body, and the instruction unit instructs switching of the display range in accordance with the movement of the main body detected by the detection unit.

[0019] (14) An image processing method including the steps of: inputting a first image including a specific area; estimating the facial expression and / or emotion of a person in the specific area based on the first image; generating a second image, which is an image of the specific area in which the person is represented by an avatar, in which at least the estimated facial expression and / or emotion is reflected in the avatar; and generating a third image by synthesizing the second image with the specific area of ​​the first image.

[0020] (15) An image processing method according to (14), further comprising the steps of: cutting out a part of the third image to generate a fourth image for display; and outputting the fourth image, wherein the step of generating the fourth image includes receiving an instruction to switch the display range, and switching the range from which the image is cut out from the third image in accordance with the received instruction to generate the fourth image. [Brief explanation of the drawings]

[0021] [Figure 1] A diagram showing the outline of the system configuration of an image processing system. [Figure 2] A diagram showing an example of installation of a photographing device. [Figure 3] FIG. 10 is a diagram showing an example of the imaging range of the imaging device; [Figure 4] FIG. 1 is a block diagram showing an example of a hardware configuration of an image processing apparatus; [Figure 5] Block diagram of functions realized by the image processing device [Figure 6] Functional block diagram of emotion estimation unit [Figure 7] Conceptual diagram of face detection [Figure 8] Conceptual diagram of emotion recognition based on facial images [Figure 9] A diagram showing an example of dividing the spectator area [Figure 10] Conceptual diagram of how to find emotions that represent divided areas [Figure 11] An example of an avatar that reflects emotions [Figure 12] FIG. 10 is a diagram showing an example of a portion of a captured image. [Figure 13] An example of an image layer for the spectator area [Figure 14] An example of a CG image [Figure 15] Block diagram showing an example of the configuration of a playback device [Figure 16] Block diagram of functions realized by the control unit of the playback device [Figure 17] 1 is a flowchart showing the processing flow of an image processing system. [Figure 18] Functional block diagram of image processing device [Figure 19] Block diagram of functions realized by the image processing device [Figure 20] Floor plan showing an example of an event venue [Figure 21] FIG. 10 is a diagram showing an example of a portion of a captured image. [Figure 22] An example of a CG image of the spectator area DETAILED DESCRIPTION OF THE INVENTION

[0022] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] [First embodiment] [overview] In systems that deliver images of events such as concerts in real time, there are known systems that allow users (viewers) receiving the content to freely change their field of view while viewing the images. In these types of systems, the images delivered are becoming higher definition in order to provide clearer images.

[0024] However, as the images being distributed become more and more high-definition, it becomes possible to identify individual spectators, which raises the risk of violating their privacy.

[0025] In this embodiment, a system for delivering images captured at an event such as a concert in real time is provided, which can convey the atmosphere of the venue while protecting the privacy of spectators at the venue.

[0026] [Image processing system configuration] FIG. 1 is a diagram showing an outline of the system configuration of an image processing system according to this embodiment.

[0027] The image processing system 1 of this embodiment is a system that distributes images captured at events such as concerts, plays, performing arts, operas, ballets, and sports.

[0028] As shown in FIG. 1, the image processing system 1 includes a photographing device 10 that photographs an event, an image processing device 100 that generates images for distribution from images photographed by the photographing device 10, a distribution device 200 that distributes the images generated by the image processing device 100, and a playback device 300 that plays back the images distributed from the distribution device 200.

[0029] [Photographing equipment] The camera device 10 captures images of an event at the event venue. The camera device 10 captures images at a fixed position. The camera device 10 also captures an area that includes at least some of the audience. In this embodiment, the camera device 10 captures images in a 360° range.

[0030] 2 is a diagram showing an example of the installation of a camera. The figure shows an example of a concert (an example of an event) being filmed at a concert venue (an example of an event venue). The figure also shows a floor plan of the concert venue.

[0031] The concert venue 2 has a stage area 3 and an audience area 4. The stage area 3 is an area where performers perform. The stage area 3 is equipped with a stage 5. The audience area 4 is an area where audience members are located. The audience area is an example of a specific area. The audience members are an example of people within a specific area. The audience area 4 is equipped with multiple seats 6. The seats 6 are arranged in a tiered manner. The audience members watch the performance from the seats 6.

[0032] 2 shows an example in which a shooting position P (the installation position of the shooting device 10) is set between the stage area 3 and the audience area 4.

[0033] FIG. 3 is a diagram showing an example of the imaging range of the imaging device.

[0034] The image capturing device 10 captures an image of a 360° range from the image capturing position P. More specifically, it captures an image of a hemispherical range (a range of 360° horizontally and 180° vertically). Therefore, both the stage area 3 and the audience area 4 are captured simultaneously. Note that this type of image capturing device (an image capturing device capable of capturing an image of a 360° range) is well known, and therefore a description of its specific configuration will be omitted (for example, a configuration in which a single device is configured to capture an image of a 360° range using a wide-angle lens, or a configuration in which multiple cameras are arranged radially and the images captured by each camera are combined to obtain an image capturing an image of a 360° range, etc.).

[0035] The photographing device 10 photographs images at a predetermined frame rate. That is, it photographs images as a moving image. The photographing device 10 sequentially outputs the photographed images to the image processing device 100. The connection form (communication form) between the photographing device 10 and the image processing device 100 is not particularly limited.

[0036] [Image processing device] The image processing device 100 receives an image captured by the image capturing device 10 and generates an image for distribution. Since the image capturing device 10 captures an image as a moving image, the image processing device 100 processes the image in frame units and generates an image (moving image) for distribution.

[0037] The image to be distributed is generated by replacing a portion of it with a CG image (an image created using computer graphics (CG)). More specifically, the image is generated by replacing the portion of the audience area 4 in the live-action image (an image captured by the image capture device 10) with a CG image. The CG image is composed of images that represent the audience as avatars (characters that represent the audience).

[0038] FIG. 4 is a block diagram showing an example of a hardware configuration of an image processing device.

[0039] The image processing device 100 is composed of a computer equipped with a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a HDD (Hard Disk Drive) 104, an operation unit (e.g., a keyboard and a mouse) 105, a display unit (e.g., a liquid crystal display (LCD), an organic electro-luminescence display (OLED), etc.) 106, an input interface (I / F) 107, and an output interface 108. An image captured by the imaging device 10 is input to the image processing device 100 via the input interface 107. An image for distribution generated by the image processing device 100 is output to a distribution device 200 via the output interface 108.

[0040] FIG. 5 is a block diagram of functions realized by the image processing device.

[0041] As shown in the figure, image processing device 100 has the functions of a captured image input unit 111, an emotion estimation unit 112, a representative emotion determination unit 113, a CG image generation unit 114, a composite image generation unit 115, and an image output unit 116. Each function is realized by a processor, CPU 101, executing a predetermined program. This program is stored in, for example, ROM 102 or HDD 104.

[0042] The captured image input unit 111 inputs an image (captured image) output from the image capturing device 10. As described above, the image capturing device 10 of this embodiment captures an image in a 360° range. Therefore, the captured image includes an image of the spectator area 4. The captured image is an example of a first image. The captured image input unit 111 is an example of a first image input unit.

[0043] The emotion estimation unit 112 analyzes the captured images and estimates the emotion of each spectator in the spectator area 4. In this embodiment, the emotion is estimated from images of the spectators' faces. The emotion estimation unit 112 is an example of a first estimation unit.

[0044] FIG. 6 is a functional block diagram of the emotion estimation unit according to the present embodiment.

[0045] As shown in the figure, the emotion estimation unit 112 has a face detection unit 112A that detects a person's face, and an emotion recognition unit 112B that recognizes emotions from a face image.

[0046] The face detection unit 112A detects the faces of spectators in the spectator area 4 from the captured image. FIG. 7 is a conceptual diagram of face detection. The figure shows a portion of the captured image (a portion of the image captured in the direction of the spectator area). The face is detected by identifying its position within the captured image (a position on the coordinate system set for the captured image). The position of the face is identified, for example, by the center position (coordinate position (xF, yF)) of a rectangular frame F that surrounds the detected face. Publicly known technology is used for face detection.

[0047] The emotion recognition unit 112B recognizes the emotion of each spectator based on the facial image of each spectator detected by the face detection unit 112A. FIG. 8 is a conceptual diagram of emotion recognition based on facial images. In this embodiment, emotions are classified into six types: "anger," "disgust," "fear," "happiness," "sadness," and "surprise." The degree of each emotion is calculated from the facial image to recognize the emotion of the spectator. More specifically, the degree of each emotion (also referred to as "emotionality") is quantified to recognize the emotion. The degree of each emotion is quantified as an emotion score. The emotion score is expressed, for example, as a percentage. This type of emotion recognition technology is well known. The emotion recognition unit 112B of this embodiment also employs well-known technology (for example, a method of recognizing emotions using an image recognition model generated by machine learning, deep learning, etc.) to recognize emotions from facial images. The emotion recognition result can be represented, for example, using an emotion vector E (Eanger, Edisgust, Efear, Ehappiness, Esadness, Esurprise) in which the emotion scores of each emotion are used as element values. Emotion recognition unit 112B outputs emotion vector E as the emotion recognition result.

[0048] It is not always possible to detect the faces of all spectators in all frames. The emotion recognition unit 112B performs emotion recognition processing on spectators whose faces have been detected.

[0049] The emotion estimation unit 112 outputs the emotion recognition result (emotion vector E) by the emotion recognition unit 112B as an emotion estimation result. The emotion estimation result of each spectator is output in association with the position of the face of each spectator.

[0050] The representative emotion determination unit 113 divides the audience area 4 into a plurality of areas, and determines an emotion to represent each divided area (divided area). The representative emotion determination unit 113 is an example of a first determination unit.

[0051] 9 is a diagram showing an example of dividing the spectator area. The figure shows an example in which the spectator area 4 is divided into 12 divided areas 4A to 4L. The figure also shows an example in which the divided areas 4A to 4L are all divided so that they have the same number of seats (an example in which the divided areas 4A to 4L are all divided so that the number of spectators is approximately the same).

[0052] FIG. 10 is a conceptual diagram of how to find an emotion that represents a divided area. As shown in the figure, an emotion that represents a divided area is found from a collection of emotion vectors E of the spectators in that divided area. In this embodiment, the average value EAV of the emotion vectors E (average value of each element value that makes up the emotion vector E) is found from the collection of emotion vectors E of the spectators in that divided area, and the emotion that represents that divided area is found based on the found average EAV. Specifically, the emotion with the highest element value is identified and the emotion that represents that divided area is found. In the example shown in FIG. 10, the average emotion vector EAV is EAV(EAVanger, EAVdisgust, EAVfear, EAVhappiness, EAVsadness, EAVsurprise) = EAV(3,1,3,88,2,10), and the highest element value is EAVhappiness = 88. Therefore, as shown in FIG. 10 The emotion that represents this divided area is "joy." The average value of the emotion vector E is an example of a standard value.

[0053] The CG image generation unit 114 generates a CG image of the audience area 4. This image is composed of images of spectators in the audience area 4 represented by avatars (characters that represent the spectators). In this embodiment, the audience area 4 is divided into multiple areas, and one avatar is placed in each divided area to generate a CG image of the audience area 4. The division pattern is the same as the division pattern used by the representative emotion determination unit 113 (see FIG. 9). The CG image generation unit 114 reflects the emotion of each divided area (the emotion representing each divided area) determined by the representative emotion determination unit 113 in the avatar placed in each divided area to generate a CG image of the audience area 4. Reflecting emotions in avatars means reflecting the emotions in the avatar's expression. FIG. 11 is a diagram showing an example of an avatar that reflects emotions. As shown in the figure, in this embodiment, emotions are reflected in the avatar's facial expression.

[0054] The CG image of the spectator area 4 is generated, for example, by placing an avatar on an image that resembles the spectator area (image layer of the spectator area). An outline of the generation of this CG image will be described below.

[0055] FIG. 12 is a diagram showing an example of a portion of a captured image (live-action image). This diagram shows an image obtained by capturing an image in the direction of the spectator area 4 (the direction indicated by arrow R in FIG. 2 (directly behind)). This image portion is replaced with a CG image when generating a composite image. Note that the diagram is deformed for ease of understanding. As described above, the image capture device 10 captures images at a fixed position. Therefore, it is possible to know in advance how the spectator area 4 will be captured by the image capture device 10. First, a base image layer of the spectator area is generated from an image captured in advance. This image does not necessarily have to be the same as the actual spectator area. For example, a deformed image of the actual spectator area can be generated as the spectator area image layer. FIG. 13 is a diagram showing an example of the spectator area image layer. Data for the spectator area image layer is stored, for example, in HDD 104. The CG image of the spectator area 4 is generated by placing an avatar on the spectator area image layer. FIG. 14 is a diagram showing an example of a CG image. The avatars are placed at positions corresponding to the positions of each divided area, with one avatar placed in each divided area. The avatars placed in each divided area reflect the emotions that represent that divided area. Figure 14 shows an example in which the emotion representing divided area 4A is "surprise," the emotion representing divided area 4B is "joy," the emotion representing divided area 4C is "joy," the emotion representing divided area 4D is "joy," the emotion representing divided area 4E is "joy," the emotion representing divided area 4F is "joy," the emotion representing divided area 4G is "joy," the emotion representing divided area 4H is "joy," the emotion representing divided area 4I is "surprise," the emotion representing divided area 4J is "surprise," the emotion representing divided area 4K is "joy," and the emotion representing divided area 4L is "joy." In this case, as shown in the figure, avatars with the emotion of "surprise" are placed in divided areas 4A, 4I, and 4J, and avatars with the emotion of "joy" are placed in divided areas 4B, 4C, 4D, 4E, 4F, 4G, 4H, and 4K, and a CG image of the audience area is generated. Note that if multiple avatars corresponding to each emotion are prepared (see FIG. 11), the avatar to be used is selected randomly.Alternatively, they are used in a predetermined order. The avatars arranged in each of the divided areas 4A to 4L are displayed adjusted to a predetermined size. In the example shown in FIG. 14, the size of the avatars displayed in each of the divided areas 4A to 4L is changed depending on the distance from the shooting position. In other words, the perspective is adjusted when the avatars are displayed.

[0056] The CG image generated by the CG image generation unit 114 is applied to the composite image generation unit 115. The CG image generation unit 114 is an example of a second image generation unit. The CG image of the audience area generated by the CG image generation unit 114 is an example of a second image.

[0057] The composite image generation unit 115 generates a composite image by combining the CG image generated by the CG image generation unit 114 with the captured image (real-life image). The composite image generation unit 115 generates a composite image by combining the CG image with the image portion of the spectator area in the captured image. As a result, an image (composite image) is generated in which the portion of the spectator area is composed of the CG image.

[0058] The composite image generated by composite image generation unit 115 is supplied to image output unit 116. Composite image generation unit 115 is an example of a third image generation unit. The composite image generated by composite image generation unit 115 is an example of a third image.

[0059] The image output unit 116 outputs the composite image generated by the composite image generation unit 115 as an image for distribution to the distribution device 200. The connection form (communication form) between the image processing device 100 and the distribution device 200 is not particularly limited.

[0060] [Distribution device] The distribution device 200 transmits images (videos) for distribution generated by the image processing device 100 to the playback device 300. The distribution device 200 is a so-called video distribution server, and transmits images for distribution to the playback device 300 in response to a request from the playback device 300, which is a client. The distribution device 200 is configured by a computer, and functions as the distribution device 200 by the computer executing a predetermined program. In other words, it functions as a video distribution server. This type of distribution device is well known, so a description of its specific configuration will be omitted. Note that the connection (communication) between the distribution device 200 and the playback device 300 is not particularly limited. For example, a form in which they communicate with each other via a network such as the Internet can be adopted.

[0061] [Playback device] The playback device 300 plays back the images (videos) transmitted from the distribution device 200. The playback device 300 cuts out a portion of the image transmitted from the distribution device 200 and plays it back. Therefore, the user (viewer) sees a portion of the image captured in 360 degrees. For example, in FIG. 3, the area VA indicated by the dashed line is the display range of the image. The display range of the image (= the range from which the image is cut out) can be changed according to an instruction from the user.

[0062] In this embodiment, the playback device 300 is configured as a head-mounted display. The head-mounted display detects the head posture of the user wearing the display, and switches the display range of the image according to the detected head posture. More specifically, the display range of the image is switched according to the line of sight inferred from the detected head posture.

[0063] FIG. 15 is a block diagram showing an example of the configuration of a playback device.

[0064] As shown in the figure, a playback device 300 of this embodiment, which is configured as a head-mounted display, includes a communication unit 301, a detection unit 302, an operation unit 303, a display unit 304, a control unit 306, and the like.

[0065] The communication unit 301 communicates with the distribution device 200. Images (videos) transmitted from the distribution device 200 are received via the communication unit 301.

[0066] The detection unit 302 detects the movement (posture) of the playback device main body (the part worn on the head) and detects the movement (posture) of the head of the user wearing the playback device 300. The detection unit 302 is configured to include sensors for head tracking, such as an acceleration sensor and an angular velocity sensor.

[0067] The operation unit 303 is composed of a plurality of operation buttons etc. provided on the playback device body. Operations on the playback device 300 are performed via this operation unit 303.

[0068] The display unit 304 is configured by a liquid crystal display, an organic electroluminescence display, etc. The display unit 304 is an example of a fourth image output unit. An image is displayed (output) on this display unit 304.

[0069] The control unit 306 controls the overall operation of the playback device 300. The control unit 306 is configured, for example, by a microcomputer equipped with a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and realizes various functions by executing predetermined programs.

[0070] FIG. 16 is a block diagram of functions realized by the control unit of the playback device.

[0071] As shown in the figure, the control unit 306 realizes the functions of a playback image input unit 306A, a visual field specifying unit 306B, a display control unit 306C, and the like.

[0072] The playback image input unit 306A controls the communication unit 301 to receive an image transmitted from the distribution device 200 and input an image (playback image) to be played back by the playback device 300. The image transmitted from the distribution device 200 is a composite image (third image) generated by the image processing device 100. Therefore, the playback image is also an example of the third image. The playback image input unit 306A is an example of a third image input unit. The input image (playback image) is applied to the display control unit 306C.

[0073] The visual field specifying unit 306B specifies the user's visual field based on the detection result of the detection unit 302. The "visual field" is the range the user is looking at, and corresponds to the range (display range) of the image displayed on the display unit 304. The visual field specifying unit 306B specifies the visual field from the movement (posture) of the user's head detected by the detection unit 302. Information about the visual field specified by the visual field specifying unit 306B is sent to the display control unit 306C.

[0074] The display control unit 306C generates a display image from the image input to the playback image input unit 306A, and displays it on the display unit 304. The display image is an image obtained by cutting out a part of the image (playback image) input to the playback image input unit 306A, and corresponds to the field of view. The display control unit 306C determines the range from which the image is cut out in accordance with the field of view specified by the field of view specifying unit 306B. The display control unit 306C switches the range of image extraction according to the field of view specified by the field of view specifying unit 306B, and generates an image for display. The image for display is an example of a fourth image. The display control unit 306C is also an example of a fourth image generating unit. In this embodiment, the image extraction range is switched according to the field of view specified by the field of view specifying unit 306B, and therefore the field of view specifying unit 306B is an example of an instruction unit.

[0075] [Image processing system operation] FIG. 17 is a flowchart showing the flow of processing in the image processing system of this embodiment.

[0076] First, photography is performed by the photography device 10 installed at the event venue (step S11). The photography device 10 photographs the event at a fixed position. The image (photographed image) photographed by the photography device 10 is output to the image processing device 100 (step S12). This image is an image photographed over a 360° range, and includes the audience area.

[0077] The image processing device 100 receives an image (captured image) output from the image capture device 10 (step S21) and performs a predetermined process to generate an image for distribution. First, a process is performed to estimate the emotion of each spectator from the input captured image (step S22). The emotion of the spectator is estimated based on the image of the spectator's face. Next, an emotion representing each divided area is determined based on the estimation result of the emotion of each spectator (step S23). Next, a CG image of the spectator area is generated (step S24). This CG image is an image in which spectators in the spectator area are represented by avatars, and is generated by placing the avatars on an image that simulates the spectator area (image layer of the spectator area). One avatar is placed in each divided area, and is positioned corresponding to the position of each divided area. Furthermore, the avatar placed in each divided area reflects the emotion representing that divided area. Once the CG image is generated, a composite image is generated (step S25). The composite image is generated by combining the CG image with a part of the captured image (real-life image). The CG image is combined with the part of the spectator area in the captured image. As a result, the portion of the photographed image (real-life image) in which the audience is captured is masked with the CG image. The generated composite image is output to distribution device 200 as an image to be distributed (step S26).

[0078] The distribution device 200 receives an image (image for distribution) output from the image processing device 100 (step S31) and transmits the image to the playback device 300 (step S32).

[0079] The playback device 300 receives an image for distribution transmitted from the distribution device 200 (step S41), performs predetermined processing, and outputs the image to the display unit 304. That is, the playback device 300 displays the image on the display unit 304. The image displayed on the display unit 304 is an image obtained by cutting out a portion of the received image. The playback device 300 generates an image for display from the received image (step S42), and displays the image on the display unit 304 (step S43).

[0080] The range from which the image is cut out (display range) can be changed in response to an instruction from the user. In this embodiment, since the playback device 300 is configured as a head-mounted display, the range from which the image is cut out can be changed in response to the movement (posture) of the head.

[0081] As described above, according to the image processing system 1 of this embodiment, a portion of a captured image is replaced with a CG image and distributed. The portion replaced with the CG image is the area in which the audience is photographed. This allows the audience's privacy to be appropriately protected. Furthermore, the CG images that are replaced are images in which the audience is represented by avatars, and each avatar reflects the emotions of the audience at the corresponding position. This allows the reactions and atmosphere of the audience at the venue to be conveyed to the user (viewer). This also allows the atmosphere of the venue to be shared on-site.

[0082] Furthermore, in the image processing system 1 of this embodiment, a CG image of the audience area is generated by replacing the actual number of audience members with fewer avatars. This makes it possible to clearly express the emotions of audience members who appear far away and small in the actual image. Therefore, it is possible to more easily convey the atmosphere of the venue to the user (viewer) than by looking at an actual image (live-action image). In particular, for events held in large venues, the atmosphere of the venue can be more appropriately conveyed than by looking at an actual image.

[0083] [Variations] [Variations of the method for determining the emotion that represents each divided area] In the above embodiment, the average value of the emotion vectors of the spectators in each divided area is calculated to determine the emotion that represents each divided area, but the method of determining the emotion that represents each divided area is not limited to this. For example, one spectator may be selected from each divided area, and the emotion of the selected spectator may be used to represent the emotion of each divided area. In this case, spectators may be selected randomly, or may be selected from predetermined locations (audience seats). Furthermore, the emotion that represents each divided area may be determined by calculating the median or mode of the emotion vectors.

[0084] Furthermore, when determining the emotion that represents each divided area from the average value, it is not necessary to calculate the average value for all spectators. For example, the emotion that represents each divided area may be calculated by calculating the average value of the emotion vectors of spectators randomly selected in each divided area. Alternatively, the emotion that represents each divided area may be calculated by calculating the average value of the emotion vectors of spectators located in a predetermined position (audience seating) in each divided area.

[0085] [Variations of Avatar Emotion Reflection] In the above embodiment, one of six emotions ("anger," "disgust," "fear," "joy," "sadness," and "surprise") is reflected in the avatar, but the emotions reflected in the avatar are not limited to these.

[0086] Furthermore, if the degree of emotion (emotion score) representing a divided area is equal to or less than a threshold, the avatar may be displayed without reflecting the emotion. In this case, the avatar may be displayed with a neutral expression (a state in which no emotion is expressed (also known as emotionless or expressionless)). Alternatively, the avatar may be displayed with a predetermined facial expression.

[0087] In the above embodiment, one type of emotion is reflected in the avatar, but a combination of multiple types of emotions may be reflected in one avatar. Also, the degree of emotion may be reflected in the avatar.

[0088] In addition, in the above embodiment, the audience's emotions are estimated from the captured image and the estimated emotions (in the above embodiment, emotions representing the divided areas) are reflected in the avatars, but the audience's facial expressions may be estimated from the captured image and the estimated facial expressions may be reflected in the avatars (including the case where emotions representing the divided areas are found and reflected in the avatars). Furthermore, the audience's emotions and facial expressions may be estimated from the captured image and the estimated emotions and facial expressions may be reflected in the avatars (including the case where emotions and facial expressions representing the divided areas are found and reflected in the avatars).

[0089] The type of facial expression can also be expressed by a word indicating an emotion, in which case both the facial expression and the emotion are specified.

[0090] [Avatar display variations] The avatar may reflect the attributes (gender, age, race (physical characteristics such as bone structure, skin, hair, etc.)) that represent each divided area.

[0091] 18 is a functional block diagram of an image processing device for generating an image for distribution by reflecting attributes representative of each divided area in an avatar. As shown in the figure, an attribute estimation unit 121 and a representative attribute determination unit 122 are further provided.

[0092] The attribute estimation unit 121 acquires a captured image from the captured image input unit 111, analyzes the acquired captured image, and estimates the attributes of each spectator in the spectator area 4. For example, it estimates the gender and age of each spectator. The technology for estimating a person's attributes from an image is a well-known technology. The attribute estimation unit 121 of this embodiment also employs a well-known technology (for example, a method of recognizing a person's attributes using a trained image recognition model). The attribute estimation unit 121 is an example of a second estimation unit.

[0093] The representative attribute determination unit 122 determines an attribute that represents each divided area. For example, for gender, the gender with the largest number of people is identified as the representative gender (so-called majority vote). For age, the representative age is identified by calculating the average. In addition, the representative attribute is determined by calculating the median, mode, etc. The representative attribute determination unit 122 is an example of a second determination unit.

[0094] When generating a CG image of the audience area 4, the CG image generation unit 114 generates the CG image of the audience area 4 by reflecting the emotions and attributes that represent each divided area in the avatars placed in each divided area.

[0095] In this way, by reflecting attributes in the avatar, the atmosphere of the venue can be conveyed to the user more realistically.

[0096] In the above example, the process of estimating the emotions of the audience from the captured image and the process of estimating the attributes are configured to be performed by separate processing units (the emotion estimation unit 112 and the attribute estimation unit 121), but they can also be configured to be performed by a single processing unit. For example, a trained image recognition model can be used to estimate the emotions and attributes of the audience from the captured image.

[0097] [A variation of the CG image of the audience area] In the above embodiment, an avatar is placed on an image that resembles the spectator area (spectator area image layer) to generate a CG image of the spectator area, but the configuration of the CG image of the spectator area is not limited to this. The base image of the spectator area (spectator area image layer) does not necessarily have to be an image that resembles an actual spectator area. For example, an image of a fictitious spectator area may be prepared and used as the base image layer.

[0098] [Variations for dividing the spectator area] In the above embodiment, the spectator area is divided into 12 areas, but the division is not limited to this. The area can be divided according to the number of people.

[0099] In the above embodiment, the spectator area is divided so that all divided areas have the same number of seats, but the manner in which the spectator area is divided is not limited to this. For example, the area may be divided so that the number of seats increases depending on the distance from the shooting position. In other words, the farther away from the shooting position, the more spectators are included.

[0100] In addition, after detecting faces in the entire audience area, it is also possible to group nearby faces and divide the audience area.

[0101] [Modifications of the imaging device] In the above embodiment, the 360° range is configured to capture a hemispherical range (a range of 360° horizontally and 180° vertically), but it may also be configured to capture a full spherical range (a range of 360° horizontally and vertically).

[0102] On the other hand, the shooting range does not necessarily have to be a 360° range around the camera, as long as part of the shooting range includes the audience area (the area where people whose privacy should be protected are present).

[0103] [Modifications of playback device] In the above embodiment, the playback device is configured as a head-mounted display, but the configuration of the playback device is not limited to this. Alternatively, the playback device can be configured as an electronic device such as a smartphone, a tablet terminal, or a personal computer. These electronic devices have a touch panel on the screen, and a user can instruct the display range to be changed by touching the touch panel. Furthermore, operation instructions can be given using an acceleration sensor, a gyro sensor, a compass, or the like provided in these electronic devices. Operation instructions can also be given by gestures using a front camera, a face recognition camera, or the like.

[0104] [Image distribution variations] In the above embodiment, an example has been described in which images captured at an event are distributed in real time, but the present invention can also be applied to cases in which already-captured images are distributed.

[0105] [Other variations] Although the above embodiment does not specifically mention the distribution of audio, it is also possible to configure the system so that audio within the venue is collected and distributed at the same time as images are captured.

[0106] [Second embodiment] [overview] When distributing images of an event, if the number of spectators at the venue is small, the excitement of users (viewers) receiving the content will also decrease.

[0107] In the image processing system of this embodiment, multiple avatars are generated from one spectator, a CG image of the spectator area is generated, and an image for distribution (a composite image) is generated.

[0108] This allows the reactions (emotions) of the audience at the venue to be amplified and conveyed to users, even at events with a small audience, increasing entertainment value.

[0109] The image processing system of this embodiment differs from the image processing system 1 of the first embodiment only in the process of generating images for distribution, so the following description will only focus on the configuration of the image processing device.

[0110] [Configuration of image processing device] The hardware configuration is the same as that of the image processing device 100 of the image processing system 1 of the first embodiment (see FIG. 4). That is, it is configured as a computer, and functions as an image processing device when the CPU executes a predetermined program.

[0111] FIG. 19 is a block diagram of functions realized by the image processing device of this embodiment.

[0112] As shown in the figure, the image processing device 100A of this embodiment has the functions of a captured image input unit 131, a feeling estimation unit 132, a CG image generation unit 134, a composite image generation unit 135, and an image output unit 136. Each function is realized by a processor, that is, a CPU 101, executing a predetermined program. This program is stored in, for example, a ROM 102 or an HDD 104.

[0113] The photographed image input unit 131 inputs an image (photographed image) output from the photographing device 10. This image is an image that includes the spectator area as a part thereof (for example, an image photographed from a fixed position over a 360° range).

[0114] The emotion estimation unit 132 analyzes the captured image and estimates the emotion of each spectator in the spectator area. The emotion estimation unit 132 detects the face of each spectator in the spectator area from the captured image and estimates the emotion of each spectator in the spectator area from the detected face image. This is the same as the emotion estimation unit 112 in the first embodiment.

[0115] The CG image generation unit 134 generates a CG image of the spectator area. This image is composed of images of spectators in the spectator area represented by avatars. In this embodiment, the CG image of the spectator area is generated using a number of avatars greater than the actual number of spectators. Specifically, multiple avatars are generated from one spectator, and the generated multiple avatars are placed on a base image layer (for example, an image that resembles the spectator area) to generate the CG image of the spectator area. The generation of the CG image of the spectator area will be described below using an example.

[0116] FIG. 20 is a plan view showing an example of an event venue.

[0117] Here, an example will be described in which images captured at a lecture (an example of an event) are distributed. A lecture hall (an example of an event hall) 400 has a podium area 410 and an audience area 420. The podium area 410 is an area where a lecturer gives a lecture. The podium area 410 is equipped with a podium 411 and a lecture desk 412, etc. The audience area 420 is an area where audience members (listeners) are located. The audience area 420 is equipped with a plurality of seats 421 and desks 422. The seats 421 and desks 422 are arranged in a stepped pattern. The image capturing device 10 captures images at an image capturing position P set between the podium area 410 and the audience area 420.

[0118] FIG. 21 is a diagram showing an example of a portion of a captured image (actual image). The figure shows an image obtained when capturing an image in the direction of the spectator area 420 (the direction indicated by arrow R in FIG. 20 (directly behind)). This image portion is the part that will be replaced with a CG image when generating a composite image. Note that the figure is shown in a deformed form to make it easier to understand. As shown in the figure, the captured image shows spectators 500 who are actually in the spectator area. The figure shows an example in which there are eight spectators 500.

[0119] FIG. 22 is a diagram showing an example of a CG image of the audience area.

[0120] As shown in the figure, in the CG image, spectators are displayed as avatars 600. As described above, in this embodiment, one spectator is replaced with multiple avatars 600 and displayed. FIG. 22 shows an example in which ten avatars 600 are generated and displayed from one spectator. Since there are eight actual spectators, 80 avatars 600 are generated and displayed. Each avatar 600 reflects the emotion of the spectator from whom it was generated. Therefore, avatars 600 generated from the same spectator will express the same emotion. The avatars 600 are randomly arranged on a base image layer.

[0121] In this way, the CG image generation unit 134 generates multiple avatars from one spectator, and places the multiple generated avatars on a base image layer to generate a CG image of the spectator area.

[0122] The composite image generation unit 135 generates a composite image by combining the CG image generated by the CG image generation unit 134 with the photographed image (real-life image). As a result, an image (composite image) in which the audience area portion is composed of the CG image is generated.

[0123] The image output unit 136 outputs the composite image generated by the composite image generation unit 135 to the distribution device 200 as an image to be distributed.

[0124] [Operation of image processing device] The image processing device 100 receives an image (captured image) output from the image capturing device 10, performs predetermined processing, and generates an image for distribution.

[0125] First, the face of each spectator is detected from the input photographed image, and the emotion of each spectator is estimated based on the image of the detected face. Next, a CG image of the spectator area is generated. This CG image is an image in which the spectators in the spectator area are represented by avatars. Multiple avatars are generated from one spectator. Each avatar reflects the emotion of the spectator from which it was generated. The CG image is generated by placing avatars on a base image layer. Once the CG image is generated, a composite image is generated. The composite image is generated by combining the CG image with a portion of the photographed image (live-action image). The CG image is combined with the portion of the photographed image in the spectator area. As a result, the portion of the photographed image (live-action image) in which the spectators are captured is masked with the CG image. The generated composite image is output to distribution device 200 as an image for distribution.

[0126] As described above, the image processing device of this embodiment generates multiple avatars from a single audience member and generates a CG image for compositing. This makes it possible to amplify the reactions of, for example, several dozen audience members to those of several hundred audience members, thereby increasing the entertainment value.

[0127] [Variations] [Avatar placement variations] In the above embodiment, a CG image is generated by randomly arranging multiple avatars generated from a single audience member, but the arrangement of the avatars is not limited to this. For example, the avatars may be arranged according to predetermined rules.

[0128] [Number of avatars to generate] The number of avatars generated from each audience member may be set by the user or may be set automatically. In the case of automatic setting, for example, the number of avatars to be displayed in the CG image may be set in advance, and the number of avatars to be generated from each audience member may be determined by working backwards from that number. For example, the number of avatars to be displayed in the generated CG image may be 100. In this case, if the number of audience members detected from the photographed image (live-action image) is 10, the number of avatars generated from each audience member will be 10. Also, if the number of audience members detected from the photographed image (live-action image) is 9, the number of avatars generated from each audience member will be 11 (rounded down to the nearest whole number).

[0129] [Avatar generated from a single spectator] It is more preferable that the avatars generated from a single spectator be composed of different characters.

[0130] [A variation of the CG image of the audience area] The base image of the spectator area (spectator area image layer) does not necessarily have to be an image that imitates the actual spectator area. For example, you can prepare an image of a fictitious spectator area and use this image as the base image layer.

[0131] [Other embodiments] [Variations of CG image generation] The CG image of the audience area may be generated by replacing all audience members whose faces have been detected with individual avatars. In other words, the CG image of the audience area is generated by replacing each audience member whose face has been detected with an avatar on a one-to-one basis. In this case, avatars are displayed for each audience member whose face has been detected.

[0132] In the above embodiment, emotions are reflected in the facial expressions of the avatar, but emotions may also be reflected in the movements of the avatar (gestures, hand movements, etc.). Furthermore, emotions may also be reflected in both facial expressions and movements.

[0133] In the above embodiment, the captured image is processed frame by frame and a CG image is generated frame by frame, but the CG image may be generated at predetermined frame intervals. In this case, the image for distribution switches to the image portion of the audience area, i.e., the CG image portion, at predetermined frame intervals.

[0134] [System Configuration] In the above embodiment, the function of generating an image for distribution from an image captured by the photographing device 10 (the function of the image processing device 100) and the function of distributing the image (the function of the distribution device 200) are realized by separate devices, but they can also be configured to be realized by a single device.

[0135] [Regarding image processing devices] Some or all of the functions of the image processing device can be realized by various processors. These include a CPU (Central Processing Unit), a general-purpose processor that executes programs and functions as various processing units; a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture; and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically to execute specific processing. A program is synonymous with software.

[0136] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types. For example, a single processing unit may be configured with multiple FPGAs, or a combination of a CPU and an FPGA. Alternatively, multiple processing units may be configured with a single processor. A first example of configuring multiple processing units with a single processor is a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as multiple processing units. A second example is a configuration in which a processor is used that realizes the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure. [Explanation of symbols]

[0137] 1. Image processing system 2. Concert venue 3 Stage Area 4 Spectator Area 4A~4L divided area 5 Stages 6 seats 10 Imaging equipment 100 Image processing device 100A Image Processing Device 101 CPU 102 ROM 104 HDD 107 Input Interface 108 Output Interface 111 Photographed image input unit 112 Emotion estimation part 112A Face detection unit 112B Emotion recognition section 113 Representative emotion determination unit 114 CG image generation unit 115 Composite image generation unit 116 Image output unit 121 Attribute estimation part 122 Representative attribute determination section 131 Photographed image input unit 132 Emotion estimation part 134 CG image generation unit 135 Composite image generation unit 136 Image output unit 200 Distribution Device 300 Playback device 301 Communications Department 302 Detection unit 303 Operation section 304 Display section 306 Control Unit 306A Replay image input unit 306B Field of View Specification Unit 306C Display control unit 410 Podium Area 411 Podium 412 Lectern 420 Spectator Area 421 seats 422 desk 500 spectators 600 Avatars F Frame surrounding detected face P Shooting position R direction arrow S11~S43 Image processing system processing procedure Image display range on VA playback device

Claims

1. a first image input unit that inputs a first image including a specific area; a first estimation unit that estimates facial expressions and / or emotions of people in the specific area based on the first image; a second image generation unit that generates, as a second image, an image in which an avatar is placed on an image that imitates the specific area, and in which at least the facial expression and / or emotion estimated by the first estimation unit is reflected in the avatar; and a third image generating unit that generates a third image by combining the second image with the specific area of ​​the first image; An image processing device comprising:

2. The first image is an image captured in a 360° range. The image processing device according to claim 1 .

3. the first image is an image of an event venue, and the specific area is an area of ​​the event venue where spectators are present; 3. The image processing device according to claim 1 or 2.

4. the first estimation unit quantifies the degree of each of a plurality of types of facial expressions and / or emotions to estimate the facial expressions and / or emotions; The image processing device according to claim 1 .

5. The avatar is randomly placed on an image that resembles the specific area. The image processing device according to claim 1 .

6. The size of the avatar is adjusted according to the placement position, and the sense of perspective is adjusted. The image processing device according to claim 5 .

7. The image processing device according to any one of claims 1 to 6, a playback device that plays back the third image generated by the image processing device; Equipped with The playback device a third image input unit that inputs the third image; a fourth image generating unit that cuts out a part of the third image and generates a fourth image for display; an instruction unit for instructing switching of the display range; a fourth image output unit that outputs the fourth image; wherein the fourth image generation unit switches a range for cutting out an image from the third image in response to an instruction from the instruction unit, and generates the fourth image. Image processing system.

8. the playback device is a head-mounted display, A detection unit is provided to detect the movement of the main body, the instruction unit instructs switching of the display range in response to the movement of the main body detected by the detection unit. The image processing system according to claim 7 .

9. inputting a first image including the specific area; estimating facial expressions and / or emotions of people in the specific area based on the first image; generating, as a second image, an image in which an avatar is placed on an image simulating the specific area, and in which at least the estimated facial expression and / or emotion is reflected in the avatar; generating a third image by combining the second image with the specific area of ​​the first image; An image processing method comprising:

Citation Information

Patent Citations

  • Image processing apparatus, camera device, communication system, image processing method, and program

    JP2009194687A

  • Display system for installation in venue

    JP2013020389A

  • Video display system, video display device and video display method

    JP2017022665A

  • Information processing device, information processing system, and method of outputting facial expression images

    JP2019087226A

  • Information processing device and method, display control device and method, reproduction device and method, programs, and information processing system

    WO2016009865A1