Information processing apparatus, information processing method, and program

JP2025029203A5Pending Publication Date: 2025-12-25GEOCREATES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024215225
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-28
Filing Date
2024-12-10
Publication Date
2025-12-25

AI Technical Summary

Benefits of technology

【0018】 本発明によれば、画像データが示す空間を視認した視認者が抱く感情を把握できるようになるという効果を奏する。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an information processing apparatus, method, and program for providing information indicating emotions that arise in a viewer who has viewed a space indicated by image data.SOLUTION: An information processing apparatus has: an image acquisition unit that acquires an analysis target image showing an analysis target space; an image analysis unit that analyzes the analysis target image to specify the hue and texture in association with each of a plurality of areas in the analysis target image; a storage unit that stores first data for analysis associated with a position in a reference space image, the combination of the hue and texture at the position, and a viewer's emotions on the reference space image; an emotion specification unit that specifies emotions of the viewer viewing the analysis target space on the basis of positions in the reference space image where the degree of correlation between the positions in the specified areas and the combination of the hue and texture corresponding to the areas is equal to or more than a threshold in the first data for analysis, and the combination of the hue and texture at the position; and an output unit that outputs the emotions at the viewing in association with the analysis target image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device, an information processing method, and a program for outputting information relating to an image showing a state of a space. [Background technology]

[0002] There is a known technique for calculating the green view factor, which is the ratio of the area of ​​plants such as trees within the visual field seen by the human eye. Patent Document 1 discloses a technique for calculating the ratio of the amount of green seen to image data showing a specified space as the green view factor. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2010-272097 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, even if the green view rate of image data was calculated, it was not possible to understand what impression a viewer who viewed the space represented by the image data had of the space.

[0005] The present invention has been made in consideration of these points, and has an object to make it possible to grasp the emotions felt by a viewer who views a space represented by image data. [Means for solving the problem]

[0006] In a first aspect of the present invention, there is provided an information processing device having an image acquisition unit that acquires image data to be analyzed that represents a state of a space to be analyzed; an image analysis unit that analyzes the image data to be analyzed to identify a hue and texture associated with each of a plurality of regions in the image to be analyzed that is represented by the image data to be analyzed; a memory unit that stores first analysis data in which a position in a reference space image, a combination of hue and texture at that position, and the emotion of a viewer of a reference space corresponding to the reference space image are associated; an emotion identification unit that identifies, in the first analysis data, a position in the reference space image where a correlation between the position of each of the plurality of regions identified by the image analysis unit and the combination of hue and texture corresponding to the plurality of regions is equal to or greater than a threshold, and a combination of hue and texture at that position, thereby identifying the emotion of a viewer of the space to be analyzed when viewing; and an output unit that outputs information indicating the emotion when viewing in association with the image to be analyzed.

[0007] In the first analysis data, the position in the reference space that the viewer is viewing is associated with the emotion of the viewer, and the emotion identification unit may refer to the first analysis data to identify the emotion at the time of viewing in association with the viewing position of the viewer in the space to be analyzed, and the output unit may output information indicating the emotion at the time of viewing in association with a position in the image to be analyzed.

[0008] The memory unit may further store second analysis data associating a position in the reference space image, a hue at that position, and the emotion of a viewer of the reference space image, and the emotion identification unit may, for a first region in the image to be analyzed, refer to the first analysis data to identify a first emotion of the viewer of the image to be analyzed based on a relationship between the position of the first region and a combination of hue and texture corresponding to the first region, and for a second region different from the first region, refer to the second analysis data to identify a second emotion of the viewer of the image to be analyzed based on the relationship between the position of the second region and the hue corresponding to the second region, and identify the emotion at the time of viewing by averaging the first emotion and the second emotion.

[0009] The image data to be analyzed may include depth information associated with pixels, and the emotion identification unit may determine an area where the depth indicated by the depth information is less than a threshold as the first area, and a area where the depth indicated by the depth information is equal to or greater than the threshold as the second area.

[0010] The emotion identifying unit may identify the viewing emotion by weighting the first emotion more heavily than the second emotion and averaging the weighted first emotion.

[0011] The emotion identification unit may identify the viewing emotion further based on a proportion of a green area in the analysis target image and a proportion of a brown area in the analysis target image.

[0012] The image acquisition unit may acquire the analysis target image data corresponding to each of the multiple analysis target spaces, the emotion identification unit may identify the emotion at the time of viewing of the analysis target space corresponding to each of the multiple analysis target image data, and the output unit may output information indicating the emotion at the time of viewing corresponding to each of the multiple analysis target images in an area corresponding to each of the multiple analysis target images in a floor map corresponding to a space including the multiple analysis target spaces corresponding to the multiple analysis target images.

[0013] The output unit may output information indicating the viewing emotion to an area on the floor map corresponding to the analysis target image in which the viewing emotion indicates a predetermined emotion.

[0014] The image acquisition unit may acquire a plurality of pieces of analysis target image data each showing a state in which the analysis target space is viewed from a different position, and the emotion identification unit may identify the viewing emotion of the analysis target space corresponding to each of the plurality of analysis target image data, and identify the viewing emotion corresponding to the analysis target space by averaging the identified plurality of viewing emotions.

[0015] The memory unit may store as the first analysis data a machine learning model that outputs the emotions of a viewer of a reference spatial image when a position in the reference spatial image and a combination of hue and texture at that position are input.

[0016] In a second aspect of the present invention, there is provided an information processing method executed by a computer, comprising the steps of acquiring image data to be analyzed that represents the state of a space to be analyzed; analyzing the image data to be analyzed to identify a hue and texture associated with each of a plurality of regions in the image to be analyzed that is represented by the image data to be analyzed by analyzing the image data to be analyzed; identifying the emotions of a viewer of the space to be analyzed when viewing by identifying positions in the reference space image, combinations of hue and texture at those positions, and combinations of hue and texture at those positions in the reference space image where the correlation between the positions of each of the identified plurality of regions and the combinations of hue and texture corresponding to the plurality of regions is equal to or greater than a threshold value, in first analysis data in which the positions in the reference space image, the combinations of hue and texture at those positions, and the emotions of a viewer of the space to be analyzed corresponding to the reference space image are associated with each other; and outputting information indicating the emotions when viewing in association with the image to be analyzed.

[0017] In a third aspect of the present invention, a program is provided to cause a computer to realize functions as an image acquisition unit that acquires image data to be analyzed that represents the state of a space to be analyzed; an image analysis unit that analyzes the image data to be analyzed to identify hue and texture in association with each of a plurality of regions in the image to be analyzed that is represented by the image data to be analyzed; an emotion identification unit that identifies the emotion of a viewer of the space to be analyzed when viewing by identifying a position in a reference space image, a combination of hue and texture at that position, and a combination of hue and texture at that position in the reference space image where the correlation between the position of each of the plurality of regions identified by the image analysis unit and the combination of hue and texture corresponding to the plurality of regions is equal to or greater than a threshold value in first analysis data in which a position in the reference space image, a combination of hue and texture at that position, and an emotion of a viewer of the reference space corresponding to the reference space image are associated with each other; and an output unit that outputs information indicating the emotion at the time of viewing in association with the image to be analyzed. Effect of the Invention

[0018] According to the present invention, it is possible to grasp the feelings of a viewer who views a space represented by image data. [Brief description of the drawings]

[0019] [Figure 1] FIG. 1 is a diagram for explaining an overview of an information processing system. [Diagram 2] FIG. 1 is a diagram illustrating a configuration of an information processing device. [Diagram 3] 1 is an example of a reference space image. [Figure 4] 1 is an example of first analysis data. [Diagram 5] 11 is an example of second analysis data. [Figure 6] 2 is an example of an image to be analyzed. [Figure 7] 11A and 11B are diagrams for explaining a process of identifying the hues and textures of a plurality of regions. [Figure 8] FIG. 2 is a diagram for explaining a first region and a second region in an image to be analyzed. [Figure 9] FIG. 13 is a schematic diagram in which a heat map of comfort levels is superimposed on an image to be analyzed. [Figure 10] 2 is an example of an image to be analyzed. [Figure 11] This is an example of a floor map with visual emotions superimposed. [Figure 12] 1 is an example of a sequence executed by the information processing system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] [Outline of Information Processing System S] FIG. 1 is a diagram for explaining an overview of an information processing system S. The information processing system S is a system that provides a user with information indicating emotions that are estimated to be felt by a viewer who views a specified space. The user is, for example, a designer or architect who designs a house, office, store, park, garden, theme park, etc. The viewer is, for example, a person who is expected to use a residence or facility designed by the user. Hereinafter, the specified space is referred to as an analysis target space.

[0021] The emotions felt by a viewer when viewing the analysis target space (hereinafter referred to as "viewing emotions") differ depending on whether the analysis target space contains plants or wood or whether it contains metal or concrete. In other words, the viewing emotions change depending on the hues of the objects contained in the analysis target space and the surface textures (i.e., textures) of the objects. Therefore, the information processing system S estimates the viewing emotions based on the combination of hues and textures in the analysis target space represented by the analysis target image viewed by the viewer, or the position at which the analysis target space was viewed, etc.

[0022] The information processing system S includes an image providing device 1, an administrator terminal 2, a viewing terminal 3, and an information processing device 4. The image providing device 1, the administrator terminal 2, the viewing terminal 3, and the information processing device 4 transmit and receive various data via a network N such as the Internet.

[0023] The analysis target image showing the state of the analysis target space is, for example, an image for viewing real estate when purchasing or renting real estate such as a house or office, an image showing a space for purchasing a product such as a store or showroom that sells products, or an image showing a space for receiving an experiential service such as a tourist spot, a theme park, or a museum. The analysis target image may be an image of an indoor space or an image of an outdoor space. The analysis target image may be, for example, a virtual reality image (VR image) created by actually photographing an outdoor space or a space inside a building, or an image created using computer graphics technology.

[0024] The image providing device 1 is a server that stores one or more analysis target image data to be viewed by a viewer who uses the information processing system S. The image providing device 1 is managed, for example, by a business operator that operates a design office, a housing exhibition center, a real estate company, or a store that sells products.

[0025] The image providing device 1 provides one or more analysis target images to the viewing terminal 3 in a form that can be viewed by the viewer. The image providing device 1 may specify the analysis target image selected by the viewer after displaying thumbnail images of the multiple analysis target images on the viewing terminal 3. The image providing device 1 may obtain, from the viewing terminal 3, information indicating a position specified on the viewing terminal 3 by the viewer.

[0026] The administrator terminal 2 is a computer used by an administrator who manages the image providing device 1 or multiple images to be analyzed. The administrator terminal 2 uploads multiple images to be analyzed stored by the administrator to the image providing device 1 via the network N. The viewing terminal 3 may function as the administrator terminal 2, and a user who uses the administrator terminal 2 may upload images to be analyzed to the image providing device 1.

[0027] The viewing terminal 3 is a terminal used by the viewer to view one or more analysis target images and to display emotions when viewing the analysis target space, and is, for example, a computer, a smartphone, or a tablet. The viewing terminal 3 may be a terminal owned by the viewer, or may be a terminal installed in a design office, a real estate company, a housing exhibition center, or a store. The viewing terminal 3 receives analysis target image data indicating the analysis target image from the image providing device 1, and displays the analysis target image based on the received analysis target image data on a display. A user who has the viewer view the analysis target image may use the viewing terminal 3 together with the viewer.

[0028] When the viewer or user performs a predetermined operation on the viewing terminal 3, the viewing terminal 3 displays one or more analysis target images in a state selectable by the viewer. For example, when the viewer selects one analysis target image from among the multiple analysis target images, the viewing terminal 3 displays the selected analysis target image 6. The viewing terminal 3 may display the analysis target images one by one, and switch the analysis target image to be displayed in response to the viewer selecting an icon for switching the analysis target image to be displayed. When the analysis target image is an image captured by an omnidirectional camera, for example, the viewing terminal 3 may change the viewing direction of the analysis target image to be displayed based on the operation of the viewer.

[0029] The viewing terminal 3 transmits, for example, analysis target image identification information (hereinafter referred to as "analysis target image ID") for identifying one or more analysis target images selected by the viewer from a plurality of analysis target images to the information processing device 4 (see FIG. 1(1)). The viewing terminal 3 may transmit, as the analysis target image, a captured image captured based on the operation of the viewer. The viewing terminal 3 may also transmit, as the analysis target image, a CG image created using computer graphics technology of the interior of a house, office, store, etc., a park, a garden, a facility such as a theme park, etc.

[0030] The information processing device 4 is a computer that identifies the viewing emotion of a viewer when viewing an analysis target space represented by an analysis target image 6 identified by an analysis target image ID received from the viewing terminal 3. The information processing device 4 estimates the viewing emotion based on a combination of hues and textures in the analysis target space represented by the analysis target image 6 (see FIG. 1(2)).

[0031] Then, the information processing device 4 transmits, for example, a heat map 7 indicating the strength of the viewing emotion to the viewing terminal 3 as information indicating the identified viewing emotion (see FIG. 1(3)). The viewing terminal 3 displays the heat map 7 on the analysis target image 6 on its display. In this way, the user or viewer can understand the viewing emotion when the viewer views the analysis target space. In the following explanation, the operation of the information processing device 4 when transmitting the viewing emotion to the viewing terminal 3 is mainly illustrated.

[0032] [Configuration of information processing device 4] 2 is a diagram showing a configuration of the information processing device 4. The information processing device 4 has a communication unit 41, a storage unit 42, and a control unit 43. The control unit 43 has an image acquisition unit 431, an image analysis unit 432, an emotion identification unit 433, and an output unit 434.

[0033] The communication unit 41 has a communication interface for transmitting and receiving data between the image providing device 1, the administrator terminal 2, or the viewing terminal 3 via the network N. The communication unit 41 inputs the analysis target image data received from the image providing device 1, and the analysis target image ID, viewing state information, or biometric information received from the viewing terminal 3 to the image acquisition unit 431. The communication unit 41 also transmits information indicating the viewing emotion input from the output unit 434 to the viewing terminal 3.

[0034] The storage unit 42 includes storage media such as a ROM (Read Only Memory), a RAM (Random Access Memory), and a hard disk. The storage unit 42 stores the analysis target image data received from the image providing device 1 in association with an analysis target image ID for identifying the analysis target image data. The storage unit 42 also stores a program executed by the control unit 43.

[0035] The storage unit 42 stores first analysis data used to identify the viewing emotion of a viewer of the analysis target space. In the first analysis data, a position in a reference space image, a combination of a hue and a texture at the position, and an emotion (viewing emotion) of a viewer of the reference space corresponding to the reference space image are associated with each other.

[0036] FIG. 3 is an example of the reference space image 5. The reference space image 5 in FIG. 3 is divided into a plurality of regions. The reference space image 5 in this embodiment is divided into nine regions of three rows and three columns. The position of each region is expressed, for example, by two-dimensional coordinates indicating the row (X) position and the column (Y) position. Specifically, the positions of each region are set as follows, from the upper left to the lower right: (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), and (3,3). Note that the method for dividing the reference space image 5 is not limited to the above, and the image can be divided into any number of regions with any arrangement. The size of the regions is also arbitrary.

[0037] Fig. 4 is an example of the first analysis data. In the first analysis data, each area of ​​the reference space is associated with the hue and texture of the area, and the emotion when viewing the area. In the first analysis data, the position viewed by the viewer in the reference space may be associated with the emotion when viewing the position.

[0038] Hue is an index that indicates differences in the appearance of a color. Hue is expressed, for example, in the RGB color space that uses the three primary colors of Red (R), Green (G), and Blue (B). Hue may also be expressed using the HSV color space or the HLS color space.

[0039] Texture is an index that indicates the texture, feel, etc., of the surface of an object. Examples of indices that indicate the texture of an object's surface include, but are not limited to, "vegetable," "woodgrain," and "metallic." An index that indicates the texture of a surface may be smoothness or surface roughness, which are indices that indicate the unevenness of the surface of an object, or may indicate a pattern or design that appears on the surface.

[0040] The viewing emotion is an index of the emotion felt by the viewer who viewed the reference space represented by the reference space image 5. The emotion felt by the viewer who viewed the reference space has been measured in advance by an experiment. The emotion of the viewer is measured based on, for example, a physiological state including at least one of the brain waves and the heart rate of the viewer who viewed a specific position in the reference space. The emotion of the viewer is expressed, for example, as a comfort level and an arousal level using the Russell circle model. For example, the emotion of the viewer is estimated to be that the greater the amount of alpha waves in the brain waves of the viewer who viewed the reference space, the more relaxed and comfortable the viewer who viewed the reference space is (i.e., the higher the comfort level of the viewer is). In addition, the emotion of the viewer is estimated to be that the greater the amount of beta waves in the brain waves of the viewer who viewed the reference space, the more focused the viewer who viewed the reference space is (i.e., the higher the arousal level of the viewer is).

[0041] As a specific example, the first analysis data indicates that the hue at the position [2,1] of the reference space image is [green (59,175,177)] and the texture is [vegetative]. The first analysis data indicates that the viewer who viewed the position [2,1] of the reference space image has a comfort level of [7] and an arousal level of [4] as emotions. In other words, the viewer who viewed the space at the position [2,1], which is [green (59,175,117)] and [vegetative], has an emotion of comfort level of [7] and an arousal level of [4] (relaxed state).

[0042] The storage unit 42 may store a viewer's emotion in association with each of a plurality of reference space images. In this case, the storage unit 42 stores the viewer's emotion in association with each of a plurality of positions of each reference space image.

[0043] The storage unit 42 may store a machine learning model as the first analysis data. For example, the first learning model, which is the first analysis data, is a model generated by executing a known machine learning process using a plurality of reference space images. Specifically, the first learning model is a model that outputs a viewing emotion of the reference space image when a position in the reference space image and a combination of a hue and a texture at the position are input. That is, in the first learning model, the position in the reference space image and the combination of a hue and a texture at the position are explanatory variables, and the viewing emotion is a target variable. The information processing device 4 predicts the viewing emotion from the position and the combination of a hue and a texture at the position using the first learning model. The learning model is, for example, a neural network. The machine learning process is, for example, an error backpropagation method.

[0044] The emotion felt by the viewer when viewing an object with the color and surface texture visible is different from the emotion felt by the viewer when viewing only the color without the texture. Therefore, the storage unit 42 stores second analysis data in which a position in the reference space image, a hue at that position, and the emotion of the viewer of the reference space image are associated with each other.

[0045] FIG. 5 is an example of the second analysis data. Unlike the first analysis data, the second analysis data does not associate textures with positions in the reference space image and viewing emotions. To give a specific example, the viewing emotions at position (2,1) in the second analysis data are comfort level [7] and arousal level [7], which are different from the viewing emotions (comfort level [7] and arousal level [4] (see FIG. 4)) in the first analysis data.

[0046] The second analysis data may be a learning model, similar to the first analysis data. In this case, the second learning model of the second analysis data is a model that outputs the emotion when viewing the reference space image when a position in the reference space image and a hue at the position are input. That is, in the second learning model, the position in the reference space image and the hue at the position are explanatory variables, and the emotion when viewing is a target variable.

[0047] The control unit 43 is, for example, a CPU (Central Processing Unit). The control unit 43 executes the programs stored in the storage unit 42 to function as an image acquisition unit 431, an image analysis unit 432, an emotion identification unit 433, and an output unit 434.

[0048] The image acquisition unit 431 acquires various types of information received by the communication unit 41. The image acquisition unit 431 acquires analysis target image data that represents the state of one or more analysis target spaces selected by the user or viewer. Specifically, the image acquisition unit 431 acquires analysis target image data that corresponds to the analysis target image ID transmitted by the viewing terminal 3. FIG. 6 is an example of the analysis target image 6.

[0049] The image acquisition unit 431 inputs the analysis target image data identified by the acquired analysis target image ID to the image analysis unit 432. Specifically, the image acquisition unit 431 acquires the analysis target image data identified by the acquired analysis target image ID from the image providing device 1, and inputs the acquired analysis target image data to the image analysis unit 432. When the viewing terminal 3 transmits multiple analysis target image IDs, the image acquisition unit 431 acquires the analysis target image data corresponding to each analysis target image ID from the image providing device 1, and inputs the acquired multiple analysis target image data to the image analysis unit 432.

[0050] The image analysis unit 432 analyzes the analysis target image data to identify the hue and texture in the analysis target space represented by the analysis target image 6. For example, the image analysis unit 432 identifies the hue and texture in association with each of a plurality of regions in the analysis target image 6 represented by the analysis target image data.

[0051] Fig. 7 is a diagram for explaining a process for identifying the hue and texture of a plurality of regions. As shown in Fig. 7, an analysis target image 6 is divided into nine regions of three rows and three columns. Identification information for identifying positions from the upper left to the lower right is set for the nine regions. The identification information for identifying positions is set, for example, as (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), and (3,3) from the upper left to the lower right.

[0052] The image analysis unit 432 identifies the hue of each region. For example, the image analysis unit 432 identifies the hue of each of the pixels included in each region, and identifies the average or median of the hues of the pixels as the hue of the region. As an example, the image analysis unit 432 identifies the average of the hues of the pixels included in the region at position (1,1) as the hue of the region.

[0053] The image analysis unit 432 identifies the texture of each region. The image analysis unit 432 identifies the texture of each region by extracting features of each region using, for example, a plurality of mask patterns. Note that the process by which the image analysis unit 432 identifies the texture is not limited to the above process.

[0054] The emotion identification unit 433 identifies the viewing emotion of the viewer viewing the analysis target space. Specifically, the emotion identification unit 433 first identifies the degree of correlation between the position of each area of ​​the analysis target space image, the combination of the hue and texture at that position, and the position of each area of ​​the reference space image 5, and the combination of the hue and texture at that position, in the first analysis data. Then, the emotion identification unit 433 identifies the viewing emotion corresponding to the position in the reference space image where the correlation degree is equal to or greater than a threshold, and the combination of the hue and texture at that position, as the viewing emotion of the viewer of the analysis target space. The correlation degree threshold is determined in advance by an experiment or the like.

[0055] As a specific example, the emotion identification unit 433 identifies an area of ​​the analysis target image 6 where the degree of correlation with a combination of hue [brown] and texture [wood grain] at position (1, 3) of the reference space image 5 is equal to or greater than a threshold value. If the hue of position (2, 3) of the analysis target image 6 is [brown] and the texture is [wood grain], the emotion identification unit 433 determines that the degree of correlation between the area at position (1, 3) of the reference space image 5 and the area at position (2, 3) of the analysis target image 6 is equal to or greater than a threshold value. Then, the emotion identification unit 433 identifies the viewing emotion of the area at position (2, 3) of the analysis target image 6 as the viewing emotion (comfort [8], arousal [8]) corresponding to position (1, 3) of the reference space image 5 (see FIG. 4).

[0056] When the first analysis data is a learning model, the emotion identification unit 433 inputs the analysis target space image into the learning model to identify the emotion at the time of viewing the analysis target space corresponding to the analysis target space image. Specifically, the emotion identification unit 433 inputs the position of each area of ​​the analysis target space image and a combination of hue and texture at the position into the learning model to identify the emotion at the time of viewing the analysis target space corresponding to the analysis target space image. In this way, the emotion identification unit 433 can more accurately identify the emotion at the time of viewing by using the learning model that identifies the emotion at the time of viewing by inputting a combination of hue and texture.

[0057] The emotion identification unit 433 may identify an emotion for each viewing position of the viewer in the analysis target space. For example, the emotion identification unit 433 identifies the viewing emotion in association with the viewing position of the viewer in the analysis target space by referring to the first analysis data. Specifically, the emotion identification unit 433 identifies the viewing emotion at each viewing position, assuming that the viewer views each position in the analysis target space. As a specific example, the emotion identification unit 433 identifies the viewing emotion when it is assumed that the viewer views the position (1,1) of the analysis target image 6, in association with the position (1,1) of the analysis target image. In this way, the emotion identification unit 433 can identify the emotion that a viewer who plans to use a house or facility designed by the user is expected to have when viewing a predetermined position in the analysis target space.

[0058] Furthermore, the emotion identification unit 433 may identify the viewing position of the viewer in the analysis space image corresponding to the analysis target space, and identify the viewing emotion in association with the identified viewing position. For example, the emotion identification unit 433 identifies the viewing position based on the analysis target image displayed on the display and the position of the viewer's pupils when the viewer wears goggles having a camera for photographing the viewer's eyes and a display for displaying the analysis target image. Then, the emotion identification unit 433 identifies the viewing emotion at the identified viewing position. In this way, the emotion identification unit 433 can identify the emotion that the viewer is expected to have when viewing the image at the viewing position.

[0059] The emotion identification unit 433 identifies the viewing emotion according to the depth of the analysis target space. For example, the emotion identification unit 433 identifies the viewing emotion based on a combination of hue and texture in the front area of ​​the analysis target space, and identifies the viewing emotion based only on hue in the back area. In this case, the analysis target image data includes depth information associated with pixels. In other words, each pixel included in the analysis target image data is associated with depth information. The depth is, for example, a numerical value indicating a position in the depth direction of the analysis target space, and the larger the numerical value, the deeper the pixel is in the analysis space.

[0060] The emotion identification unit 433 determines a region where the depth indicated by the depth information is less than a depth threshold for distinguishing between the foreground and the background as a first region, which is a region in the foreground of the space to be analyzed. The emotion identification unit 433 determines a region where the depth indicated by the depth information is equal to or greater than the depth threshold as a second region different from the first region. The depth threshold is, for example, an intermediate value between the maximum and minimum values ​​of the depth indicated by the depth information of each pixel included in the image data to be analyzed, but is not limited thereto.

[0061] Fig. 8 is a diagram for explaining a first region A and a second region B in an image 6 to be analyzed. The first region A shown in Fig. 8 is a region that includes only pixels that are less than the depth threshold. The second region B is a region that includes only pixels that are equal to or greater than the depth threshold. In other words, the first region A of the image 6 to be analyzed is located in front of the second region B.

[0062] With respect to the first region A, the emotion identification unit 433 refers to the first analysis data to identify the first emotion of the viewer of the analysis target image based on the relationship between the position of the first region A and the combination of the hue and texture corresponding to the first region A. For example, the emotion identification unit 433 inputs the position of the first region A and the combination of the hue and texture of the first region A into the first learning model to identify the first emotion of the viewer of the analysis target image.

[0063] For the second region B, the emotion identification unit 433 refers to the second analysis data to identify the second emotion of the viewer of the image to be analyzed 6 based on the relationship between the position of the second region B and the hue corresponding to the second region B.

[0064] Then, the emotion identifying unit 433 identifies the viewing emotion based on the first emotion and the second emotion. For example, the emotion identifying unit 433 identifies the viewing emotion by averaging the first emotion and the second emotion. The emotion identifying unit 433 may also identify the median of the first emotion and the second emotion as the viewing emotion. In this way, the emotion identifying unit 433 can identify the viewing emotion that reflects the emotion felt by the viewer when viewing a front area where the color and surface texture of the object can be seen, and the emotion felt by the viewer when viewing a back area where the texture is difficult to see and only the color can be seen.

[0065] Since the foreground area in the analysis target space has a greater influence on the viewing emotion than the background area, the emotion identifying unit 433 may identify the viewing emotion by weighting the first emotion more heavily than the second emotion and averaging them. In this way, the emotion identifying unit 433 can weight the emotion when the viewer views the foreground area of ​​the analysis target space more heavily than the emotion when the viewer views the background area of ​​the analysis target space. As a result, the emotion identifying unit 433 can identify the viewing emotion including the influence of distance.

[0066] The emotion identification unit 433 may identify the viewing emotion based on the ratio of green areas in the analysis target image 6 (hereinafter referred to as "green view ratio") and the ratio of brown areas in the analysis target image 6 (hereinafter referred to as "wood view ratio"). For example, the emotion identification unit 433 identifies a higher comfort level as the green view ratio increases between 0% and a predetermined value, and identifies a lower comfort level as the green view ratio increases after the predetermined value. In other words, the comfort level for the green view ratio is plotted as an upwardly convex graph. The emotion identification unit 433 identifies at least one of the comfort level and the awakening level for the wood view ratio, similar to the green view ratio. The predetermined value of the wood view ratio is different, but may be the same as the predetermined value of the green view ratio. In this way, the emotion identification unit 433 can identify the viewing emotion according to the ratio of plants and the ratio of wood to the entire analysis target space.

[0067] The emotion identification unit 433 identifies the emotion at the time of viewing of one analysis target space by identifying the emotion at the time of viewing of each of the multiple analysis target images corresponding to one analysis target space. One analysis target space is, for example, an indoor living room, kitchen, dining room, bedroom, etc. In this case, the image acquisition unit 431 acquires multiple analysis target image data showing a state in which one analysis target space is visually recognized from different positions. In other words, the image acquisition unit 431 acquires, as analysis target image data, multiple captured images in which at least one of the position and the direction in which one analysis target space is captured is different.

[0068] Next, the image analysis unit 432 identifies the hue and texture of each region of the analysis target image corresponding to each analysis target image data. Next, the emotion identification unit 433 identifies the viewing emotion of the analysis target space corresponding to each of the multiple analysis target image data. Specifically, the emotion identification unit 433 inputs the hue and texture of each region of each analysis target image data to the first learning model, thereby identifying multiple viewing emotions corresponding to each analysis target image data. Then, the emotion identification unit 433 identifies the viewing emotion corresponding to one analysis target space based on the identified multiple viewing emotions. For example, the emotion identification unit 433 identifies an average value of multiple viewing emotions or a median value of multiple viewing emotions as the viewing emotion corresponding to one analysis target space.

[0069] Furthermore, the emotion identification unit 433 may identify a viewing emotion corresponding to one analysis target space by averaging a plurality of viewing emotions after weighting each viewing emotion according to the importance of the analysis target image data corresponding to the viewing emotion. Specifically, the emotion identification unit 433 averages the viewing emotions by weighting the analysis target image data showing a state in which the viewer views the one analysis target space from a position where the viewer frequently views more heavily than the analysis target image data showing a state in which the viewer views the one analysis target space from a position where the viewer rarely views. In this way, the emotion identification unit 433 can appropriately identify the emotion that the viewer has when spending time in one analysis target space.

[0070] The output unit 434 outputs information indicating the viewing emotion in association with the analysis target image 6. For example, the output unit 434 outputs information for displaying the viewing emotion superimposed on the analysis target image 6. Specifically, the output unit 434 outputs information for displaying a heat map of the comfortableness of the viewing emotion superimposed on the analysis target image 6 to the viewing terminal 3.

[0071] FIG. 9 is a schematic diagram in which a heat map 7 of comfort level is superimposed on the analysis target image 6. The white areas in FIG. 9 are areas where the comfort level is higher than other areas. The black areas in FIG. 9 are areas where the comfort level is lower than other areas. This makes it easier for the user to understand where in the analysis target space the viewer looks to get a high comfort level.

[0072] The output unit 434 may output a radar chart in which a plurality of items indicating the viewing emotion are expressed on a regular polygon as information indicating the viewing emotion. The plurality of items indicating the viewing emotion are, for example, types of emotion. The types of emotion are, for example, relaxation, concentration, satisfaction, fatigue, and productivity. The output unit 434 transmits information for displaying a radar chart indicating the viewing emotion together with the analysis target image 6. In this way, the user can easily understand what emotions the viewer who views the analysis target space corresponding to the analysis target image 6 has.

[0073] (Processing to display emotions when viewed superimposed on floor map) In a space including a plurality of spaces to be analyzed, a user may want to know the viewing emotions corresponding to each of the plurality of spaces to be analyzed. The space may be, for example, a commercial facility such as a department store or a shopping mall, or a public facility such as a hospital or a government office. In a department store, the plurality of spaces to be analyzed may be, for example, a plurality of stores and a rest area in the department store. The information processing device 4 identifies the viewing emotions of the plurality of spaces to be analyzed included in the space, and displays the identified viewing emotions on the manager terminal 2 in association with a map of the space (hereinafter referred to as a "floor map"). The process of displaying the viewing emotions superimposed on the floor map will be specifically described below.

[0074] The image acquisition unit 431 acquires analysis target image data corresponding to each of a plurality of analysis target spaces included in one space. For example, the image acquisition unit 431 acquires a plurality of captured images of each of a plurality of analysis target spaces as analysis target image data. The image acquisition unit 431 may also acquire one analysis target image data showing a plurality of analysis target spaces. For example, the image acquisition unit 431 acquires a 360-degree image created by synthesizing a plurality of captured images captured in all directions of a space from a predetermined position on a computer. The image acquisition unit 431 may also acquire a plurality of images obtained by dividing a 360-degree image by a predetermined angle as a plurality of analysis target image data. FIG. 10 is an example of an analysis target image 8. The analysis target image 8 in FIG. 10 corresponds to each of a plurality of analysis target spaces included in one space.

[0075] The emotion identification unit 433 identifies the emotion at the time of viewing of the analysis target space corresponding to each of the multiple analysis target image data. Specifically, the emotion identification unit 433 inputs the multiple analysis target image data to the first learning model, thereby identifying the emotion at the time of viewing of each analysis target image data.

[0076] The output unit 434 outputs information for displaying a floor map corresponding to a space including a plurality of analysis target spaces corresponding to a plurality of analysis target images by superimposing the viewing emotion on the viewing terminal 3. Specifically, the output unit 434 outputs information indicating the viewing emotion corresponding to each of the plurality of analysis target images to an area on the floor map corresponding to each of the plurality of analysis target images. FIG. 11 is an example of a floor map on which the viewing emotion is superimposed. The shaded area in FIG. 11 indicates the comfort level when the viewer views the area. The darker the shaded area, the higher the comfort level. For example, the comfort level of the areas 71 and 72 is higher than the areas 70, 73, 74, and 75. The comfort level of the area 70 is higher than the areas 73, 74, and 75.

[0077] In this way, a user who sees the floor map on which the viewing emotion is superimposed can easily understand which part of the floor the viewer should look at to feel comfortable. In addition, the user can understand which part of the floor the viewer should look at to feel uncomfortable, which can be utilized to improve the floor environment.

[0078] [Sequence executed by information processing system S] Fig. 12 is an example of a sequence executed by the information processing system S. The sequence of Fig. 12 is executed when an instruction to start a process of identifying a viewing emotion is transmitted from the viewing terminal 3, for example.

[0079] First, the image providing device 1 transmits a plurality of image data to the viewing terminal 3 that has transmitted the instruction to start the process of identifying the viewing emotion (step S1). For example, the image providing device 1 transmits a plurality of image data, each of which is captured at different positions and in different directions inside a store selling products, to the viewing terminal 3.

[0080] The viewing terminal 3 displays on a display a plurality of images corresponding to the plurality of image data transmitted from the image providing device 1. The viewer selects an image for which a viewing emotion is to be identified from among the plurality of images displayed on the viewing terminal 3. The viewing terminal 3 accepts the selection of an image for which a viewing emotion is to be identified from among the plurality of images (step S2). The viewing terminal 3 transmits an analysis target image ID for identifying the analysis target image to the information processing device 4 (step S3).

[0081] The information processing device 4 identifies the hue and texture of the analysis target image corresponding to the analysis target image data identified by the analysis target image ID transmitted from the viewing terminal 3 (step S4). Specifically, the information processing device 4 identifies the positions of multiple regions of the analysis target image and the combination of hue and texture in each region.

[0082] Next, the information processing device 4 identifies the viewing emotion of the analysis target image whose hue and texture have been identified (step S5). Specifically, the information processing device 4 inputs the position of each region and the combination of the hue and texture of each region into the first learning model to identify the viewing emotion of each region.

[0083] The information processing device 4 outputs the viewing emotion to the viewing terminal 3 in association with the analysis target image 6 (step S6). For example, the output unit 434 outputs information for displaying the viewing emotion superimposed on the analysis target image. Specifically, the output unit 434 outputs information for displaying a heat map of the comfortableness of the viewing emotion superimposed on the analysis target image to the viewing terminal 3.

[0084] The viewing terminal 3 displays the viewing emotion identified by the information processing device 4 on the display (step S7). Specifically, the viewing terminal 3 displays a heat map of the viewing emotion comfort level on the display, superimposed on the analysis target image (see FIG. 9).

[0085] In addition, the information processing device 4 may not only transmit the viewing emotion to the viewing terminal 3, but may also display the viewing emotion on the display of the administrator terminal 2, print it on a print medium such as paper, or transmit it to an external device other than the viewing terminal 3 (for example, the administrator terminal 2).

[0086] [Effects of information processing device 4] As described above, the information processing device 4 stores the first analysis data in which a position in the reference space image, a combination of hue and texture at the position, and the emotion of the viewer of the reference space corresponding to the reference space image are associated. First, the information processing device 4 analyzes the analysis target image data representing the state of the analysis target space to identify the hue and texture of each area in the analysis target image represented by the analysis target image data. Next, the information processing device 4 identifies the positions in the reference space image corresponding to the positions of multiple areas and the combinations of hue and texture corresponding to each area, and the combinations of hue and texture at the positions in the first analysis data to identify the viewing emotion of the viewer of the analysis target space. Then, the information processing device 4 outputs information indicating the viewing emotion in association with the analysis target image.

[0087] In this way, the information processing device 4 can identify the viewing emotion of the viewer that is estimated to be felt when viewing the space of the hue and texture of the analysis target image whose correlation with the combination of the hue and texture of the reference space image is equal to or greater than a threshold value. As a result, the information processing device 4 can provide information indicating the viewing emotion of the viewer who viewed the analysis target space represented by the analysis target image data to a user who is a designer or architect. Then, the user who receives the viewing emotion can understand the emotion that the viewer is expected to have when viewing, for example, the analysis target space designed by the user himself. Furthermore, the user can design a space that is comfortable and allows the user to concentrate by checking the viewing emotion of the analysis target space whose layout has been changed with reference to the viewing emotion.

[0088] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by distributing or integrating functionally or physically in any unit. In addition, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effect of the new embodiment resulting from the combination combines the effect of the original embodiment. [Explanation of symbols]

[0089] S Information Processing System 1 Image providing device 2. Administrator terminal 3 Viewing device 4. Information processing equipment 41 Communications Department 42 Storage section 43 Control Unit 431 Image Acquisition Unit 432 Image Analysis Unit 433 Emotion identification part 434 Output section

Claims

1. an image acquisition unit that acquires analysis target image data that represents the state of the analysis target space; an image analysis unit that analyzes the analysis target image data to identify a hue and a texture in association with each of a plurality of regions in the analysis target image represented by the analysis target image data; a storage unit that stores a first machine learning model that outputs the emotion of a viewer of a reference space image when a position in the reference space image and a combination of hue and texture at the position are input; an emotion identification unit that identifies an emotion of a viewer of the analysis target space when viewing the analysis target space by inputting the positions of the plurality of regions identified by the image analysis unit and a combination of a hue and a texture corresponding to the plurality of regions into the first machine learning model; an output unit that outputs information indicating the viewing emotion in association with the analysis target image; An information processing device having the above.

2. the storage unit further stores a second machine learning model that outputs an emotion of the viewer of the reference space image when a position in the reference space image and a hue at the position are input; and the emotion identification unit identifies a first emotion of a viewer of the analysis target image using the first machine learning model for a first region in the analysis target image, identifies a second emotion of the viewer of the analysis target image using the second machine learning model for a second region different from the first region, and identifies the viewing emotion by averaging the first emotion and the second emotion; The information processing device according to claim 1 .

3. the analysis target image data includes depth information associated with pixels; the emotion identification unit determines a region where the depth indicated by the depth information is less than a threshold as the first region, and determines a region where the depth indicated by the depth information is equal to or greater than the threshold as the second region. The information processing device according to claim 2 .

4. the emotion identification unit identifies the viewing emotion by weighting the first emotion more heavily than the second emotion and averaging the weights. The information processing device according to claim 3 .

5. the emotion identification unit identifies the viewing emotion further based on a proportion of a green area in the analysis target image and a proportion of a brown area in the analysis target image.

3. The information processing device according to claim 1 or 2.

6. the image acquisition unit acquires the analysis target image data corresponding to each of the plurality of analysis target spaces, the emotion identification unit identifies the viewing emotion of the analysis target space corresponding to each of the plurality of analysis target image data; the output unit outputs information indicating the viewing emotion corresponding to each of the plurality of analysis target images to an area corresponding to each of the plurality of analysis target images in a floor map corresponding to a space including the plurality of analysis target spaces corresponding to the plurality of analysis target images.

3. The information processing device according to claim 1 or 2.

7. The computer executes acquiring analysis target image data representing a state of an analysis target space; analyzing the analysis target image data to identify a hue and a texture associated with each of a plurality of regions in the analysis target image represented by the analysis target image data; a step of inputting the positions of the identified regions and combinations of hues and textures corresponding to the regions into a first machine learning model that outputs the emotion of a viewer of the reference space image when a position in the reference space image and a combination of hues and textures at the position are input, thereby identifying the emotion of a viewer of the analysis target space when viewing the image; outputting information indicating the viewing emotion in association with the analysis target image; An information processing method comprising:

8. On the computer, an image acquisition unit that acquires analysis target image data that represents the state of the analysis target space; an image analysis unit that analyzes the analysis target image data to identify a hue and a texture in association with each of a plurality of regions in the analysis target image represented by the analysis target image data; an emotion identification unit that identifies the emotion of a viewer of the analysis target space when viewing the analysis target space by inputting the positions of the identified regions and combinations of hues and textures corresponding to the regions into a first machine learning model that outputs the emotion of a viewer of the reference space image when a position in the reference space image and a combination of hues and textures at the position are input; an output unit that outputs information indicating the viewing emotion in association with the analysis target image; A program to realize the function as follows.