Information processing system and program
The information processing system addresses the challenge of overlapping speech content and facial expressions by controlling display information in non-overlapping areas, ensuring clear speech comprehension and facial observation.
Patent Information
- Application Number
- JP2024095312
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-12-24
AI Technical Summary
Existing systems fail to allow understanding of a speaker's speech content while observing their facial expressions due to overlapping display configurations.
An information processing system that acquires a speech image, determines a specific area within the image that does not overlap the speaker's face, and controls display information to be shown in this area, considering size, visibility, and characteristics.
Enables understanding of the speaker's utterance while observing their facial expressions by optimizing the display area and visibility of the speech content.
Smart Images

Figure 2025186884000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system and a program. [Background technology]
[0002] Patent document 1 describes an information processing device that has an environmental information processing unit that predicts changes in the environment outside the vehicle while content is being viewed, and an optimization processing unit that determines how the content will be drawn based on the predicted changes in the environment outside the vehicle, in which the optimization processing unit detects a non-transparent screen on which the content will be drawn, or the scenery outside the vehicle as seen through a transparent screen on which the content will be drawn, as the background of the content, and determines the drawing color based on the color of the background. Patent Document 2 describes a color compensation method that sets a preset object position of a virtual object relative to a real scene, captures an image of the real scene by using an image sensor, maps the image of the real scene to a field of view (FOV) of a display, generates a background image relative to the FOV of the display, performs color compensation on the virtual object according to a background overlap region corresponding to the preset object position in the background image, generates an adjusted virtual object, and displays the adjusted virtual object on a display according to the preset object position. Patent document 3 describes an electronic device in which a display unit displays a first virtual image together with a background visible in the user's line of sight, and the display unit displays a second virtual image to enhance the visibility of the first virtual image together with the background visible in the user's line of sight prior to the timing of displaying the first virtual image, and the second virtual image includes a first virtual image displayed in a color that is complementary to the color of the background. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2023 / 074126 [Patent Document 2] Japanese Patent Publication No. 2020-17252 [Patent Document 3] Japanese Patent Publication No. 2022-89884 Summary of the Invention [Problem to be solved by the invention]
[0004] In some cases, the content of a speaker's speech is displayed in an area of the speech image that includes the speaker. In this case, it may be necessary to adopt a configuration in which the content of the speech is displayed in an area that overlaps the speaker's face in the speech image. However, if such a configuration is adopted, it becomes impossible to understand the content of the speaker's speech while viewing the speaker's facial expression.
[0005] An object of the present invention is to make it possible to understand the content of a speaker's utterance while observing the speaker's facial expression. [Means for solving the problem]
[0006] The invention described in claim 1 is an information processing system comprising one or more processors, which acquire a speech image including a speaker, acquire display information for displaying the speech content of the speaker, and control the display information to be displayed in a specific area within the speech image that does not overlap the speaker's face. The invention described in claim 2 is an information processing system described in claim 1, in which the one or more processors determine the specific area to be an area having a size determined based on the display information. A third aspect of the present invention is the information processing system according to the second aspect, wherein the size determined based on the display information is a size in which the display information can be arranged. The invention described in claim 4 is the information processing system described in claim 2, in which the size determined based on the display information is a size that makes it possible to arrange the display information by reducing the size of the display information within a visible range. The invention described in claim 5 is the information processing system described in claim 1, wherein the one or more processors determine an area having characteristics defined based on the display information as the specific area. The invention described in claim 6 is an information processing system described in claim 5, in which the area having characteristics determined based on the display information is an area in which the visibility of the display information when placed therein is higher than a predetermined visibility standard. The invention described in claim 7 is an information processing system described in claim 5, in which the area having characteristics determined based on the display information is an area in which the amount of change of the display information to increase the visibility of the display information above a predetermined visibility standard is less than a predetermined standard amount. The invention described in claim 8 is an information processing system described in claim 1, in which the one or more processors acquire the speech image further including other speakers, further acquire other display information for displaying the speech content of the other speakers, and control the display information and the other display information to be displayed in multiple areas within the speech image that do not overlap either the speaker's face or the faces of the other speakers. The invention described in claim 9 is the information processing system described in claim 8, wherein the multiple areas include a first area that is within a predetermined reference distance from the speaker and the other speakers, and a second area that is within the reference distance from the speaker but not within the reference distance from the other speakers, and the one or more processors determine the first area to be the specific area under the condition that the visibility of the display information when placed in the first area is higher than a predetermined visibility standard, and that the visibility of the display information when placed in the second area is lower than the visibility standard. The invention described in claim 10 is the information processing system described in claim 9, wherein the multiple areas further include a third area that is not within the reference distance from the speaker but is within the reference distance from the other speaker, and the one or more processors determine the first area to be the specific area under the further condition that the visibility of the other display information when placed in the first area and the visibility of the other display information when placed in the third area are higher than the visibility standard. The invention described in claim 11 is an information processing system described in claim 1, in which the one or more processors, prior to displaying the display information in the specific area within the speech image, modify the display information so as to improve the visibility of the display information when the display information is placed in the specific area. The invention described in claim 12 is a program for enabling a computer to realize the following functions: acquiring a speech image including a speaker; acquiring display information for displaying the speech content of the speaker; and controlling the display information to be displayed in a specific area within the speech image that does not overlap the speaker's face. [Effects of the Invention]
[0007] According to the invention of claim 1, it is possible to understand the content of a speaker's utterance while observing the speaker's facial expression. According to the invention of claim 2, the display information for displaying the speech content of the speaker can be displayed in the display area, taking into consideration the size of the display area. According to the invention of claim 3, the display information can be grasped within one display area without reducing the size of the display information. According to the invention of claim 4, even if there is no area large enough to arrange the display information, the display information can be grasped within one display area. According to the invention of claim 5, display information for displaying the content of a speaker's speech can be displayed in the display area, taking into consideration the characteristics of the display area. According to the invention of claim 6, the visibility of the displayed information can be improved without changing the displayed information. According to the seventh aspect of the present invention, even if there is no area where the visibility of the display information can be improved, the visibility of the display information can be improved. According to the invention of claim 8, it is possible to understand the contents of speeches from a plurality of speakers while viewing their facial expressions. According to the invention of claim 9, it is possible to determine whether to display information for displaying the speech content of one speaker in an area within a reference distance from one speaker and another speaker, or in an area within a reference distance from one speaker but not within a reference distance from another speaker. According to the invention of claim 10, it is possible to decide to display display information for displaying the speech content of another speaker in an area that is not within a reference distance from one speaker but is within a reference distance from another speaker. According to the invention of claim 11, it is possible to improve the visibility of the display information for displaying the content of the speaker's speech. According to the invention of claim 12, it is possible to understand the content of what the speaker is saying while looking at the speaker's facial expression. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of an AR system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating an example of the hardware configuration of AR glasses according to the present embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a conceptual configuration of an AR module according to the present embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of the hardware configuration of an AR server according to the present embodiment. [Figure 5] FIG. 1 is a diagram illustrating a schematic operation of an AR system according to a first embodiment. [Figure 6] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a second embodiment. [Figure 7] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a third embodiment. [Figure 8] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a fourth embodiment. [Figure 9] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a fifth embodiment. [Figure 10] FIG. 2 is a block diagram illustrating an example of a functional configuration of an AR server according to the present embodiment. [Figure 11] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the first aspect. [Figure 12] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the second aspect. [Figure 13]10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the third aspect. [Figure 14] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the fourth aspect. [Figure 15] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the fifth aspect. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, the present embodiment will be described in detail with reference to the accompanying drawings.
[0010] (Outline of this embodiment) This embodiment provides an information processing system that acquires a speech image including a speaker, acquires display information for displaying the speaker's speech content, and controls the display information to be displayed in a specific area within the speech image that does not overlap the speaker's face.
[0011] Here, the "system" may be configured by a single device or multiple devices. In the following, an information processing system configured by a single device will be taken as an example. The single device will be described as an AR server in an AR (Augmented Reality) system.
[0012] (Overall configuration of the AR system) 1 is a diagram showing an example of the overall configuration of an AR system 1 according to the present embodiment. As shown in the figure, the AR system 1 includes AR glasses 10, an AR server 30, and a communication line 80. Although only one AR glass 10 is shown in the figure, there may be multiple AR glasses 10.
[0013] The AR glasses 10 are a glasses-type wearable terminal device. Here, "wearable" means that the user can wear them. Therefore, a glasses-type wearable terminal device is a computer device that has the shape of glasses and can actually be worn on the user's head. The AR glasses 10 are a device that enables the user to see AR. Here, "AR" stands for "Augmented Reality," and refers to displaying a virtual screen superimposed on real space for the user. In other words, the user can view the virtual screen through the AR glasses 10, and can also view real space through the AR glasses 10. In this case, the "virtual screen" refers to an image that is created by a computer and can be viewed through the AR glasses 10. And the "real space" refers to a space that actually exists.
[0014] The AR glasses 10 have two cameras 11 attached to both ends of the front part of the frame. In this embodiment, a two-dimensional image is assumed as the augmented reality image (hereinafter also referred to as "AR image"), but a three-dimensional image may also be used. A three-dimensional image refers to an image in which distance information is recorded for each pixel, and is also called a "distance image." To acquire a three-dimensional image, a stereo camera may be used as the camera 11. Alternatively, a LiDAR (Light Detection and Ranging) may be used to acquire a three-dimensional image.
[0015] Although the AR glasses 10 are shown here as an eyeglass-type device, the present invention is not limited to this. Any shape or format of device that displays AR may be used. Specifically, a broader optically transmissive display may be used. For example, MR (Mixed Reality) glasses may be used instead of AR glasses.
[0016] The AR server 30 is a server computer that performs processing to display information on the AR glasses 10. Specifically, the AR server 30 generates information to be displayed on the AR glasses 10 and outputs this information to a microdisplay 122 (described later) of the AR glasses 10.
[0017] The communication line 80 is a line used for information communication between the AR glasses 10 and the AR server 30. For example, a wireless local area network (LAN) or the Internet may be used as the communication line 80. Alternatively, for example, a mobile communication system such as 4G or 5G, or Bluetooth (registered trademark) may be used as the communication line 80.
[0018] (AR glasses hardware configuration) 2 is a diagram showing an example of the hardware configuration of the AR glasses 10 according to the present embodiment. As shown in the figure, the AR glasses 10 include a data processing unit 100. The AR glasses 10 also include a camera 11, an AR module 120, a microphone 130, and a speaker 140. The AR glasses 10 also include a communication module 150.
[0019] The data processing unit 100 includes a processor 101. The data processing unit 100 further includes a read only memory (ROM) 102 and a random access memory (RAM) 103. The data processing unit 100 further includes a flash memory 104. The processor 101 is configured by, for example, a CPU (Central Processing Unit), and realizes various functions by executing programs. The ROM 102, RAM 103, and flash memory 104 are all semiconductor memories. The ROM 102 stores a BIOS (Basic Input Output System) and the like. The RAM 103 is a main storage device used to execute programs. For example, a DRAM (Dynamic RAM) is used as the RAM 103. The flash memory 104 is used to record firmware, programs, data files, etc. The flash memory 104 is used as an auxiliary storage device.
[0020] The camera 11 captures an image in front of the user's field of view. The viewing angle of the camera 11 may be approximately the same as or greater than the viewing angle of a person. For example, a CMOS image sensor or a CCD image sensor is used as the camera 11. The number of cameras 11 may be one or more. In the example of FIG. 1, there are two cameras 11. In this case, the two cameras 11 may be arranged, for example, at both ends of the front part of the frame. Using two cameras 11 enables stereo imaging. This makes it possible to measure the distance to the subject and estimate the front-to-back relationship between the subjects.
[0021] The AR module 120 is a module that realizes the visual recognition of augmented reality by combining an AR image with a real landscape. The AR module 120 is composed of optical components and electronic components. Representative methods for the AR module 120 include the following: First, a method in which a half mirror is placed in front of the user's eye. Second, a method in which a volume hologram is placed in front of the user's eye. Third, a method in which a blazed diffraction grating is placed in front of the user's eye.
[0022] The microphone 130 is a device that converts the user's voice and surrounding sounds into electrical signals. The speaker 140 is a device that converts an electrical signal into sound and outputs the sound. The speaker 140 may be a bone conduction speaker or a cartilage conduction speaker. The speaker 140 may be a device independent of the AR glasses 10, such as a wireless earphone. In this case, the speaker 140 is connected to the AR glasses 10 via Bluetooth (registered trademark) or the like.
[0023] The communication module 150 is a device that complies with a protocol used for communication via the communication line 80. The communication module 150 may also be a device that complies with a protocol used for communication with other external devices. Examples of protocols used for communication with external devices include Wifi (registered trademark) and Bluetooth (registered trademark).
[0024] Although not shown, the AR glasses 10 may be additionally provided with an inertial sensor, a positioning sensor, a vibrator, and the like.
[0025] 3 is a diagram illustrating a conceptual configuration example of AR module 120 according to this embodiment. AR module 120 shown in FIG. 3 corresponds to a method in which a blazed diffraction grating is placed in front of the user's eyes. AR module 120 shown in Fig. 3 includes a light guide plate 121 and a microdisplay 122. AR module 120 shown in Fig. 3 also includes a diffraction grating 123A to which image light L2 is input. AR module 120 shown in Fig. 3 also includes a diffraction grating 123B from which image light L2 is output.
[0026] The light guide plate 121 corresponds to a lens of glasses. The light guide plate 121 has a transmittance of, for example, 85% or more. Therefore, the user can directly view the scenery ahead through the light guide plate 121. The external light L1 travels straight through the light guide plate 121 and the diffraction grating 123B and enters the eye E of the user.
[0027] The microdisplay 122 is a display device that displays an AR image that is visually recognized by the user. Light of the AR image displayed on the microdisplay 122 is projected onto the light guide plate 121 as image light L2. The image light L2 is refracted by the diffraction grating 123A and reaches the diffraction grating 123B while reflecting inside the light guide plate 121. The diffraction grating 123B refracts the image light L2 toward the eye E of the user. As a result, external light L1 and image light L2 are simultaneously incident on the user's eye E. As a result, the user recognizes the presence of an AR image in front of the user's line of sight.
[0028] (AR server hardware configuration) 4 is a diagram illustrating an example of the hardware configuration of the AR server 30 according to this embodiment. As illustrated, the AR server 30 includes a data processing unit 300. The AR server 30 further includes an HDD (hard disk drive) 310 and a communication module 320.
[0029] The data processing unit 300 includes a processor 301. The data processing unit 300 further includes a ROM 302 and a RAM 303. The processor 301 is configured by, for example, a CPU. The processor 301 realizes various functions by executing programs. The ROM 302 and the RAM 303 are both semiconductor memories. The ROM 302 stores the BIOS and the like. The RAM 303 is used as a main storage device used for executing programs. The RAM 303 may be, for example, a DRAM.
[0030] The HDD 310 is an auxiliary storage device that uses a magnetic disk as a recording medium. In this embodiment, the HDD 310 is used as the auxiliary storage device. However, a non-volatile rewritable semiconductor memory may also be used as the auxiliary storage device. An operating system and application programs are installed on the HDD 310.
[0031] The communication module 320 is a device that complies with the protocol used for communication over the communication line 80 .
[0032] Although not shown, the AR server 30 may be additionally provided with a display, a keyboard, a mouse, and the like.
[0033] (Outline of AR system operation) ((First Aspect)) FIG. 5 is a diagram showing a schematic operation of the AR system 1 according to the first embodiment. In FIG. 5, a background image 200 including a speaker U is viewed through the AR glasses 10. The background image 200 includes regions 201 to 203, which are uniformly colored regions that do not overlap the face of the speaker U. The AR server 30 also acquires display information 205 representing the speech of the speaker U. Here, all of the regions 201 to 203 are regions that do not reduce the visibility of the display information 205 when it is displayed. Among these regions, the region 203 is particularly set to increase the visibility of the display information 205 when it is displayed. Therefore, the user selects the region 203 as indicated by the mouse cursor 209. As a result, the AR server 30 displays the display information 205 in the region 203 of the background image 200.
[0034] ((Second Aspect)) FIG. 6 is a diagram showing a schematic operation of the AR system 1 according to the second embodiment. In FIG. 6, a background image 220 including a speaker U is viewed through the AR glasses 10. The background image 220 includes regions 221 to 223, which are uniformly colored regions that do not overlap the face of the speaker U. The AR server 30 also acquires display information 225 representing the speech of the speaker U. Here, the regions 221 and 222 are regions where the visibility of the display information 225 becomes low when the display information 225 is displayed. That is, the regions 221 and 222 are regions where the visibility of the display information 225 becomes low unless the color of the display information 225 is changed when the display information 225 is displayed. On the other hand, the region 223 is a region where the visibility of the display information 225 becomes high when the display information 225 is displayed. That is, the color of the region 223 does not need to be changed when the display information 225 is displayed. Therefore, the AR server 30 displays the display information 225 in the region 223 of the background image 220.
[0035] ((Third Aspect)) 7(a) and (b) are diagrams showing the general operation of the AR system 1 of the third embodiment. In FIG. 7(a), a background image 240 including a speaker U is viewed through the AR glasses 10. The background image 240 includes an area 241, which is an area of uniform color that does not overlap the face of the speaker U. The AR server 30 also acquires display information 245 that represents the speech of the speaker U. Here, the area 241 is an area large enough to display the display information 245. Therefore, the AR server 30 displays the display information 245 in the area 241 of the background image 240.
[0036] In FIG. 7(b), a background image 250 including a speaker U is viewed through the AR glasses 10. The background image 250 includes a region 251, which is a uniformly colored region that does not overlap the face of the speaker U. The AR server 30 also acquires display information 255 representing the speech of the speaker U. Here, the region 251 is not large enough to display the display information 255. In this case, the AR server 30 extracts a region 252 that is large enough to display the display information 255 and has a small amount of color change. The region 252 is an area where two colors change only once, as an example of an area where a small amount of color change occurs. Therefore, the AR server 30 displays the display information 255 in the region 252 of the background image 250. In this case, the region 252 may include a region 253, which may reduce the visibility of the display information 255 when it is displayed. Therefore, the AR server 30 changes the color of the portion of the display information 255 that overlaps the region 253. For example, the AR server 30 may change the color of the portion of the dialogue area 257 of the display information 255 that overlaps with the area 253. Note that the color change is omitted in the drawing.
[0037] ((Fourth Aspect)) 8(a) and 8(b) are diagrams showing the general operation of the AR system 1 of the fourth embodiment. In FIG. 8(a), a background image 260 including a speaker U is viewed through the AR glasses 10. The background image 260 includes regions 261 and 262, which are uniformly colored regions that do not overlap the face of the speaker U. The AR server 30 also acquires display information 265 representing the speech of the speaker U. Here, the region 261 is set to be a region large enough to display the display information 265. On the other hand, the region 262 is set to be a region that is not large enough to display the display information 265. In other words, the region 262 is set to be a region that cannot be displayed even if the display information 265 is transformed. Therefore, the AR server 30 displays the display information 265 in the region 261 of the background image 260.
[0038] In FIG. 8(b), a background image 270 including a speaker U is viewed through the AR glasses 10. The background image 270 includes a region 271, which is a uniformly colored region that does not overlap the face of the speaker U. The AR server 30 also acquires display information 275 representing the speech of the speaker U. Here, the region 271 is not large enough to display the display information 275. In other words, the region 271 is an region that cannot be displayed even if the display information 275 is transformed. Therefore, the AR server 30 changes the size of the display information 275 and displays it in the region 271 of the background image 270. For example, the AR server 30 may change the sizes of the text 276 and the dialogue region 277 of the display information 275.
[0039] ((Fifth Aspect)) FIG. 9 is a diagram showing a schematic operation of the AR system 1 according to the fifth embodiment. In Fig. 9, a background image 280 including speakers U1 and U2 is viewed through the AR glasses 10. The background image 280 includes areas 281 to 283, which are uniformly colored areas that do not overlap the faces of the speakers U1 and U2. The AR server 30 also acquires display information 285 representing the speech of speaker U1 and display information 286 representing the speech of speaker U2. Whether the display information 285 and 286 represent an utterance by speakers U1 or U2 can be indicated, for example, by the color of the display information 285 and 286. However, in the figure, the colors are indicated by thick and thin lines.
[0040] Here, areas 281 and 282 are located near speaker U1, and are therefore candidates for areas in which display information 285 is to be displayed. Of these, area 281 is an area in which the visibility of display information 285 is high when it is displayed. On the other hand, area 282 is an area in which the visibility of display information 285 is low when it is displayed. Therefore, area 281 is the only candidate for an area in which display information 285 is to be displayed. On the other hand, areas 281 and 283 are located near speaker U2 and are therefore candidates for areas for displaying display information 286. Both areas 281 and 283 are areas that will increase the visibility of display information 286 when it is displayed. Therefore, areas 281 and 283 are candidates for areas for displaying display information 285.
[0041] In this case, the AR server 30 determines the area to display the display information in ascending order of the number of display area candidates. First, the number of candidates for the area in which the display information 285 is to be displayed is 1. Therefore, the AR server 30 displays the display information 285 in the area 281 of the background image 280. Next, the number of candidates for the area in which the display information 286 is to be displayed is two. Therefore, the AR server 30 displays the display information 286 in the area 283 of the background image 280. In other words, the area 281 in which the display information 285 is already displayed is excluded from the area in which the display information 286 is to be displayed. In FIG. 9, this exclusion is indicated by a dashed line from the display information 286 to the area 281.
[0042] (AR server functional configuration) 10 is a block diagram showing an example of the functional configuration of the AR server 30 according to this embodiment. As shown in the figure, the AR server 30 includes a captured image acquisition unit 41, an audio information acquisition unit 42, and a display information acquisition unit 43. The AR server 30 also includes a display area determination unit 44, a display information change unit 45, and a display control unit 46.
[0043] The captured image acquisition unit 41 acquires a captured image including a speaker captured by the camera 11 of the AR glasses 10. In the first to fourth aspects, the captured image acquisition unit 41 acquires a captured image including the speaker U. In this case, the captured image including the speaker U is an example of an utterance image including a speaker. Furthermore, the processing of the captured image acquisition unit 41 is an example of acquiring an utterance image. In the fifth aspect, the captured image acquisition unit 41 acquires a captured image including speakers U1 and U2. In this case, speaker U1 is an example of a speaker, and speaker U2 is an example of another speaker. Furthermore, the captured image including speakers U1 and U2 is an example of a speech image that includes the speakers and further includes the other speakers. Furthermore, the processing of the captured image acquisition unit 41 is an example of acquiring a speech image.
[0044] The audio information acquisition unit 42 acquires audio information including the speaker's voice collected by the microphone 130 of the AR glasses 10. At this time, the audio information acquisition unit 42 may acquire the speaker's identification information by including it in the audio information and linking it to the speaker's voice. In the first to fourth aspects, the audio information acquisition unit 42 acquires audio information including the audio of the speaker U. In this case, since only the speaker U is included in the captured image, the audio information acquisition unit 42 does not need to acquire identification information of the speaker U. In the fifth aspect, the voice information acquisition unit 42 acquires voice information including the voice of the speaker U1 and the voice of the speaker U2. At this time, the voice information acquisition unit 42 may acquire identification information of the speaker U1 by including it in the voice information in a state where it is linked to the voice of the speaker U1. Furthermore, the voice information acquisition unit 42 may acquire identification information of the speaker U2 by including it in the voice information in a state where it is linked to the voice of the speaker U2.
[0045] The display information acquisition unit 43 acquires display information for displaying the content of the speaker's speech based on the audio information acquired by the audio information acquisition unit 42. For example, the display information acquisition unit 43 may acquire, as display information, character information obtained by speech recognition of the speaker's speech included in the audio information. The display information acquisition unit 43 may also acquire, as display information, an illustration image drawn based on this character information. Furthermore, the display information acquisition unit 43 may acquire, as display information, a video of a sign language interpreter based on the speaker's speech included in the audio information. In the first to fourth aspects, the display information acquisition unit 43 acquires display information for displaying the speech content of the speaker U. In this case, the processing of the display information acquisition unit 43 is an example of acquiring display information for displaying the speech content of the speaker. In the fifth aspect, the display information acquisition unit 43 acquires display information for displaying the speech content of the speakers U1 and U2. At this time, the display information acquisition unit 43 acquires the display information for displaying the speech content of the speakers U1 and U2 separately. At this time, the display information acquisition unit 43 may make such a distinction based on identification information of the speakers U1 and U2 included in the audio information. Furthermore, the display information acquisition unit 43 may represent the result of such a distinction using a color in the display information. In this case, the processing of the display information acquisition unit 43 is an example of acquiring display information for displaying the speech content of one speaker and further acquiring other display information for displaying the speech content of another speaker.
[0046] The display area determination unit 44 extracts an area that does not overlap the face of the speaker from the captured image acquired by the captured image acquisition unit 41. For example, the display area determination unit 44 may detect a face area using an existing technology and extract the other area as an area that does not overlap the face.
[0047] Furthermore, the display area determination unit 44 extracts areas of uniform color from areas that do not overlap the extracted speaker's face. For example, if the distance between colors within an area is equal to or less than a threshold, the display area determination unit 44 may extract the area as an area of uniform color.
[0048] Furthermore, the display area determination unit 44 determines a display area for displaying display information from the extracted uniformly colored areas. At this time, the display area determination unit 44 may determine the display area based on the display information acquired by the display information acquisition unit 43. In this case, the display area is an example of a specific area that does not overlap with the face of the speaker in the speech image.
[0049] Specifically, the display area determination unit 44 may determine an area having a size determined based on the display information as the display area. In this case, the processing of the display area determination unit 44 is an example of determining an area having a size determined based on the display information as the specific area. Here, the size determined based on the display information may be, for example, a size that is a predetermined proportion of the display information. The size determined based on the display information may be, for example, a size sufficient to display the display information. In this case, the processing of the display area determination unit 44 is an example of determining an area large enough to accommodate the display information as a specific area. The size determined based on the display information may be, for example, a size that can be displayed by reducing the size of the display information. Reducing the size of the display information means reducing the size of the display information within a visible range. In this case, the processing of the display area determination unit 44 is an example of determining, as a specific area, an area having a size in which the display information can be arranged by reducing the size of the display information within a visible range.
[0050] Furthermore, the display area determination unit 44 may determine an area having characteristics determined based on the display information as the display area. In this case, the processing of the display area determination unit 44 is an example of determining an area having characteristics determined based on the display information as the specific area. The feature determined based on the display information may be, for example, a feature related to the visibility of the display information when the display information is displayed. Here, the feature related to the visibility of the display information may be, for example, a feature that the visibility of the display information is high. Note that high visibility may mean that the visibility is higher than a predetermined visibility standard. In this case, the processing of the display area determination unit 44 is an example of determining, as the specific area, an area in which the visibility of the display information when the display information is arranged is higher than a predetermined visibility standard. The feature related to the visibility of the display information may be, for example, a feature that the amount of change required to increase the visibility of the display information is small. Note that high visibility may mean that the visibility is higher than a predetermined visibility standard. Furthermore, a small amount of change may mean that the amount of change is smaller than a predetermined standard amount. In this case, the processing of the display area determination unit 44 is an example of determining, as the specific area, an area in which the amount of change required to increase the visibility of the display information to a level higher than the predetermined visibility standard is smaller than a predetermined standard amount.
[0051] Furthermore, when the captured image includes multiple speakers, the display area determination unit 44 may determine the display area as follows. That is, the display area determination unit 44 first identifies the area near each speaker. Next, the display area determination unit 44 determines the number of areas within the area that will be highly visible when display information representing the speech content of each speaker is displayed. Next, the display area determination unit 44 determines the display area starting from the display information representing the speech content of each speaker that has the fewest number of such areas.
[0052] In the first mode, the display area determination unit 44 extracts an area from the captured image that does not overlap the face of the speaker U. The display area determination unit 44 also extracts an area of uniform color from the area that does not overlap the face of the speaker U. The display area determination unit 44 then extracts, as a candidate, an area from the uniform color area that requires minimal changes when displaying display information. The user then selects an area that they wish to use as the display area from these candidates. The display area determination unit 44 then determines the area selected by the user as the display area.
[0053] In the second mode, the display area determination unit 44 extracts an area from the captured image that does not overlap the face of the speaker U. The display area determination unit 44 also extracts an area of uniform color from the area that does not overlap the face of the speaker U. Furthermore, the display area determination unit 44 determines, as the display area, an area of uniform color that requires minimal changes when displaying display information.
[0054] In the third mode, the display area determination unit 44 extracts an area from the captured image that does not overlap the face of the speaker U. In addition, the display area determination unit 44 extracts an area of uniform color from the area that does not overlap the face of the speaker U. Furthermore, the display area determination unit 44 determines whether or not there is an area large enough to display display information within the uniformly colored area. If there is such an area, the display area determination unit 44 determines that area as the display area. If there are multiple such areas, the display area determination unit 44 determines the area that will have the highest visibility when the display information is displayed as the display area. If there is no such area, the display area determination unit 44 determines an area with a small amount of color change as the display area.
[0055] In the fourth mode, the display area determination unit 44 extracts an area from the captured image that does not overlap the face of the speaker U. In addition, the display area determination unit 44 extracts an area of uniform color from the area that does not overlap the face of the speaker U. Furthermore, the display area determination unit 44 determines whether or not there is an area large enough to display display information within the uniformly colored area. If there is such an area, the display area determination unit 44 determines that area as the display area. If there are multiple such areas, the display area determination unit 44 determines the area that will have the highest visibility when the display information is displayed as the display area. If there is no such area, the display area determination unit 44 determines whether there is an area where the display information can be displayed if it is made smaller. If there is such an area, the display area determination unit 44 determines that area as the display area. If there are multiple such areas, the display area determination unit 44 determines the area that will have the highest visibility when the display information is displayed as the display area. If there is no such area, the display area determination unit 44 determines the area with the least amount of color change as the display area.
[0056] In the fifth mode, the display area determination unit 44 extracts an area from the captured image that does not overlap the faces of the speakers U1 and U2. The display area determination unit 44 also extracts an area of uniform color from the area that does not overlap the faces of the speakers U1 and U2. Furthermore, the display area determination unit 44 extracts an area near the speaker U1 and an area near the speaker U2 from the uniformly colored areas. In the example of Fig. 9, the areas near the speaker U1 are areas 281 and 282, and the areas near the speaker U2 are areas 281 and 283. Next, the display area determination unit 44 counts the number of candidate areas for displaying display information representing the speech content of the speaker U1 among the areas near the speaker U1. The display area determination unit 44 counts the number of candidate areas for displaying display information representing the speech content of the speaker U2 among the areas near the speaker U2. Here, the candidate areas for displaying display information may be areas that have high visibility when the display information is displayed. In the example of FIG. 9, the only candidate area for displaying display information representing the speech content of the speaker U1 is area 281, and the number of such areas is one. The candidate areas for displaying display information representing the speech content of the speaker U2 are areas 281 and 283, and the number of such areas is two. Next, the display area determination unit 44 determines a display area for displaying display information representing the speech content of the speakers U1 and U2 based on the number of area candidates. Specifically, the display area is determined from the display information with the fewest number of area candidates. In the example of FIG. 9, the number of area candidates for displaying display information 285, 286 representing the speech content of the speakers U1 and U2 is 1 and 2, respectively. Therefore, first, the display area for displaying display information 285 representing the speech content of the speaker U1 is determined to be area 281. Next, the display area for displaying display information 286 representing the speech content of the speaker U2 is determined to be area 283.
[0057] In this case, region 281 is an example of a first region that is within a predetermined reference distance from the speaker and other speakers. Region 282 is an example of a second region that is within the reference distance from the speaker but not within the reference distance from other speakers. Region 283 is an example of a third region that is not within the reference distance from the speaker but is within the reference distance from other speakers. Furthermore, the processing of the display area determination unit 44 is an example of determining the first area as a specific area under the condition that the visibility of the display information when placed in the first area is higher than a predetermined visibility standard, and the visibility of the display information when placed in the second area is lower than the visibility standard. Furthermore, the processing of the display area determination unit 44 is an example of determining the first area as a specific area with the additional condition that the visibility of other display information when the other display information is placed in the first area and the visibility of other display information when the other display information is placed in the third area are higher than a visibility standard.
[0058] The display information modification unit 45 modifies the display information acquired by the display information acquisition unit 43 in accordance with the display area determined by the display area determination unit 44. At this time, the display information modification unit 45 modifies the display information so that the visibility of the display information is improved when the display information is displayed in the display area. Here, the modification of the display information may be a modification of the inner color, border color, size, etc. of the text included in the display information. Alternatively, the modification of the display information may be a modification of the inner color, border color, size, etc. of a dialogue area included in the display information. This is an example of the display information modification unit 45 modifying the display information so as to improve the visibility of the display information when the display information is placed in a specific area prior to displaying the display information in a specific area within the speech image.
[0059] The display control unit 46 controls the display information acquired by the display information acquisition unit 43 to be displayed on the AR glasses 10. Here, the display information acquired by the display information acquisition unit 43 may be changed by the display information change unit 45. For example, the display control unit 46 transmits the display information and the position at which to display the display information to the AR glasses 10. As a result, the display information is displayed with high visibility in a display area on the AR glasses 10 that does not overlap the face of the speaker. In the first to fourth aspects, the display control unit 46 controls to display display information for displaying the speech content of the speaker U. In this case, the display control unit 46 controls to display the display information in a display area that does not overlap the face of the speaker U and that provides high visibility. In this case, the processing of the display control unit 46 is an example of controlling to display the display information in a specific area that does not overlap the face of the speaker in the speech image. In the fifth aspect, the display control unit 46 controls to display display information for displaying the speech content of the speakers U1 and U2. At that time, the display control unit 46 controls to display the display information in a display area that does not overlap the faces of the speakers U1 and U2 and that increases visibility. In this case, the processing of the display information acquisition unit 43 is an example of controlling to display the display information and other display information in multiple areas that do not overlap the faces of the speakers and the faces of other speakers in the speech image.
[0060] (AR server operation) ((First Aspect)) FIG. 11 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the first aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 401). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 402). Next, the display information acquisition unit 43 acquires display information representing the content of the speaker's speech based on this voice information (step 403).
[0061] Next, the display area determination unit 44 extracts an area from the captured image that does not overlap the speaker's face (step 404). The display area determination unit 44 also extracts an area of uniform color from the area that does not overlap the speaker's face (step 405). Consider a case where display information is to be displayed in this uniform color area. In this case, the display area determination unit 44 extracts candidate areas that require minimal changes to improve the visibility of the display information (step 406). Suppose the user then selects an area in which the user wants to display the display information from the candidates. The display area determination unit 44 then determines the selected area as the display area (step 407).
[0062] Next, the display information change unit 45 changes the display information based on the captured image in the determined display area (step 408). Specifically, the display information change unit 45 changes the display information so that the visibility of the display information relative to the captured image in the display area is increased. Note that if the visibility of the display information relative to the captured image in the display area is already sufficiently high, this step does not need to be executed.
[0063] Thereafter, the display control unit 46 controls the AR glasses 10 to display the display information in the determined display area (step 409). When the AR server 30 acquires a plurality of pieces of display information, it only needs to execute the processes of steps 404 to 409 the number of times corresponding to the number of pieces of display information.
[0064] ((Second Aspect)) FIG. 12 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the second embodiment. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 421). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 422). Next, the display information acquisition unit 43 acquires display information representing the content of the speaker's speech based on this voice information (step 423).
[0065] Next, the display area determination unit 44 extracts an area from the captured image that does not overlap the speaker's face (step 424). The display area determination unit 44 also extracts an area of uniform color from this area that does not overlap the speaker's face (step 425). Furthermore, consider a case where display information is displayed in this uniform color area. In this case, the display area determination unit 44 determines an area that will undergo minimal changes in order to increase the visibility of the display information as the display area (step 426).
[0066] Next, the display information change unit 45 changes the display information based on the captured image in the determined display area (step 427). Specifically, the display information change unit 45 changes the display information so that the visibility of the display information relative to the captured image in the display area is increased. Note that if the visibility of the display information relative to the captured image in the display area is already sufficiently high, this step does not need to be executed.
[0067] After that, the display control unit 46 controls the AR glasses 10 to display the display information in the determined display area (step 428). When the AR server 30 acquires a plurality of pieces of display information, it only needs to execute the processes of steps 424 to 428 the number of times corresponding to the number of pieces of display information.
[0068] ((Third Aspect)) FIG. 13 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the third aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 441). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 442). Next, the display information acquisition unit 43 acquires display information representing the content of the speaker's speech based on this voice information (step 443).
[0069] Next, the display area determination unit 44 extracts an area from the captured image that does not overlap the speaker's face (step 444). The display area determination unit 44 also extracts an area of uniform color from this area that does not overlap the speaker's face (step 445). The display area determination unit 44 then determines whether or not there is an area of sufficient size within this uniform color area (step 446). Here, a sufficiently large area refers to an area large enough to display display information.
[0070] As a result, it is determined that there is a sufficiently large area within this uniformly colored area. In this case, the display area determination unit 44 determines the display area from within that sufficiently large area. Specifically, the display area determination unit 44 determines an area that will undergo minimal changes to improve the visibility of the displayed information as the display area (step 447). On the other hand, if it is determined that there is no sufficiently large area within this uniform color area, the display area determination unit 44 determines an area with minimal color variation as the display area from among areas that do not overlap with the speaker's face (step 448).
[0071] Next, the display information change unit 45 changes the display information based on the captured image in any of the determined display areas (step 449). Specifically, the display information change unit 45 changes the display information so that the visibility of the display information relative to the captured image in the display area is increased. Note that if the visibility of the display information relative to the captured image in the display area is already sufficiently high, this step does not need to be executed.
[0072] After that, the display control unit 46 controls the AR glasses 10 to display the display information in the determined display area (step 450). If the AR server 30 acquires multiple pieces of display information, it only needs to execute the processing of steps 444 to 450 the number of times corresponding to the number of pieces of display information.
[0073] ((Fourth Aspect)) FIG. 14 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the fourth aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 461). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 462). Next, the display information acquisition unit 43 acquires display information representing the content of the speaker's speech based on this voice information (step 463).
[0074] Next, the display area determination unit 44 extracts an area from the captured image that does not overlap the speaker's face (step 464). The display area determination unit 44 also extracts an area of uniform color from this area that does not overlap the speaker's face (step 465). Furthermore, the display area determination unit 44 determines whether or not there is an area of sufficient size within this area of uniform color (step 466). Here, an area of sufficient size refers to an area large enough to display display information.
[0075] As a result, it is determined that there is a sufficiently large area within this uniformly colored area. In this case, the display area determination unit 44 determines the display area from within that sufficiently large area. Specifically, the display area determination unit 44 determines an area that will undergo minimal changes to improve the visibility of the displayed information as the display area (step 467). On the other hand, if it is determined that there is no area of sufficient size within this uniform color area, the display area determination unit 44 determines whether there is an area where the display information can be displayed if it is made smaller (step 468). Specifically, the display area determination unit 44 determines whether there is such an area within the uniform color area.
[0076] As a result, it is determined that there is an area where the display information can be displayed if it is reduced in size. In this case, the display area determination unit 44 determines the display area from the area where the display information can be displayed if it is reduced in size. Specifically, the display area determination unit 44 determines an area that requires minimal changes to improve the visibility of the display information as the display area (step 469). Then, the display information modification unit 45 reduces the size of the display information to fit the determined display area (step 470). On the other hand, if it is determined that there is no area in which the display information can be displayed if it is made smaller, the display area determination unit 44 determines, as the display area, an area with a small amount of color change from among areas that do not overlap with the speaker's face (step 471).
[0077] Next, the display information change unit 45 changes the display information based on the captured image in any of the determined display areas (step 472). Specifically, the display information change unit 45 changes the display information so that the visibility of the display information relative to the captured image in the display area is increased. Note that if the visibility of the display information relative to the captured image in the display area is already sufficiently high, this step does not need to be executed.
[0078] After that, the display control unit 46 controls the AR glasses 10 to display the display information in the determined display area (step 473). When the AR server 30 acquires a plurality of pieces of display information, it only needs to execute the processes of steps 464 to 473 the number of times corresponding to the number of pieces of display information.
[0079] ((Fifth Aspect)) FIG. 15 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the fifth aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a plurality of speakers from the AR glasses 10 (step 481). Next, the voice information acquisition unit 42 acquires voice information including the voices of multiple speakers from the AR glasses 10 (step 482). Next, the display information acquisition unit 43 acquires a plurality of pieces of display information each representing the speech content of a plurality of speakers based on this voice information (step 483).
[0080] Next, the display area determination unit 44 extracts areas from the captured image that do not overlap the speaker's face (step 484). The display area determination unit 44 also extracts areas of uniform color from the areas that do not overlap the speaker's face (step 485). Furthermore, the display area determination unit 44 counts the number of candidate areas for displaying display information representing the speech content of each speaker (step 486). Specifically, for each speaker, the display area determination unit 44 calculates the number of areas near the speaker that will increase the visibility of the display information. As a result, the display area determination unit 44 determines the display area starting from the speaker with the fewest number of candidate areas (step 487).
[0081] Next, the display information change unit 45 changes the display information based on the captured image in the determined display area (step 488). Specifically, the display information change unit 45 changes the display information so that the visibility of the display information relative to the captured image in the display area is increased. Note that if the visibility of the display information relative to the captured image in the display area is already sufficiently high, this step does not need to be executed.
[0082] After that, the display control unit 46 controls the AR glasses 10 to display the display information in the determined display area (step 489). When the AR server 30 acquires a plurality of pieces of display information, it only needs to execute the processes of steps 484 to 489 the number of times corresponding to the number of pieces of display information.
[0083] (Processor) In this embodiment, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Furthermore, the operations of the processor in this embodiment may not only be performed by one processor, but may also be performed by multiple processors located at physically separate locations working together. Furthermore, the order of the operations of the processor is not limited to the order described in this embodiment, and may be changed.
[0084] (program) The present embodiment can be applied to a program and a program product. For example, a program to which this embodiment is applied can be understood as a program that enables a computer to realize the following functions: acquiring a speech image including a speaker; acquiring display information for displaying the content of the speaker's speech; and controlling the display information to be displayed in a specific area within the speech image that does not overlap the speaker's face. The program to which this embodiment is applied can be provided not only by communication means but also by being stored in a recording medium such as a CD-ROM.
[0085] (Addendum) (((1))) one or more processors; the one or more processors: Acquire a speech image including the speaker; acquiring display information for displaying the speech content of the speaker; Controlling the display information so that it is displayed in a specific area within the speech image that does not overlap with the face of the speaker; Information processing system. (((2))) the one or more processors: The information processing system according to (((1))), wherein an area having a size determined based on the display information is determined as the specific area. (((3))) The information processing system according to (((2))), wherein the size determined based on the display information is a size in which the display information can be arranged. (((4))) The information processing system described in (((2))), wherein the size determined based on the display information is a size that allows the display information to be arranged by reducing the size of the display information within a visible range. (((5))) the one or more processors: The information processing system according to any one of ((1))) to ((4))), wherein an area having a characteristic determined based on the display information is determined to be the specific area. (((6))) The information processing system described in (((5))), wherein the area having characteristics determined based on the display information is an area in which the visibility of the display information when placed therein is higher than a predetermined visibility standard. (((7))) The information processing system described in (((5))), wherein the area having characteristics determined based on the display information is an area in which the amount of change of the display information to increase the visibility of the display information above a predetermined visibility standard is less than a predetermined standard amount. (((8))) the one or more processors: acquiring the speech image further including another speaker; Further acquiring other display information for displaying the speech content of the other speaker; controlling the display information and the other display information to be displayed in a plurality of areas in the speech image that do not overlap with either the face of the speaker or the faces of the other speakers; An information processing system according to any one of ((1))) to (((7))). (((9))) the plurality of regions include a first region that is within a predetermined reference distance from the speaker and the other speakers, and a second region that is within the reference distance from the speaker but not within the reference distance from the other speakers; the one or more processors: The information processing system described in (((8))) determines the first area to be the specific area under the condition that the visibility of the display information when placed in the first area is higher than a predetermined visibility standard, and the visibility of the display information when placed in the second area is lower than the visibility standard. (((10))) the plurality of regions further includes a third region that is not within the reference distance from the speaker but is within the reference distance from the other speaker; the one or more processors: The information processing system described in (((9))) determines the first area to be the specific area under the further condition that the visibility of the other display information when placed in the first area and the visibility of the other display information when placed in the third area are higher than the visibility standard. (((11))) the one or more processors: An information processing system described in any of ((1))) to (((10))), wherein, prior to displaying the display information in the specific area within the speech image, the display information is modified so as to improve the visibility of the display information when the display information is placed in the specific area. (((12))) On the computer, A function of acquiring a speech image including the speaker; a function of acquiring display information for displaying the speech content of the speaker; a function of controlling the display information to be displayed in a specific area within the speech image that does not overlap with the face of the speaker; A program to achieve this.
[0086] According to the invention of (((1))), it is possible to understand the content of what the speaker is saying while looking at the speaker's facial expression. According to the invention (((2))), the display information for displaying the content of the speaker's speech can be displayed in the display area, taking into consideration the size of the display area. According to the invention (((3))), the display information can be grasped within one display area without reducing the size of the display information. According to the invention (((4))), even if there is no area large enough to arrange the display information, it is possible to grasp the display information within one display area. According to the invention (((5))), display information for displaying the content of a speaker's speech can be displayed in the display area, taking into consideration the characteristics of the display area. According to the invention (((6))), the visibility of the displayed information can be improved without changing the displayed information. According to the invention (((7))), even if there is no area where the visibility of the displayed information can be improved, the visibility of the displayed information can be improved. According to the invention (((8))), it is possible to understand the contents of speech of multiple speakers while viewing their facial expressions. According to the invention (((9))), it is possible to determine whether to display information for displaying the speech content of one speaker in an area that is within a reference distance from one speaker and another speaker, or in an area that is within the reference distance from one speaker but not within the reference distance from another speaker. According to the invention of (((10))), it is possible to determine to display display information for displaying the speech content of another speaker in an area that is not within a reference distance from one speaker but is within a reference distance from another speaker. According to the invention (((11))), it is possible to improve the visibility of the display information for displaying the content of the speaker's speech. According to the invention of (((12))), it is possible to understand the content of what the speaker is saying while looking at the speaker's facial expression. [Explanation of symbols]
[0087] 1...AR system, 10...AR glasses, 30...AR server, 41...captured image acquisition unit, 42...audio information acquisition unit, 43...display information acquisition unit, 44...display area determination unit, 45...display information change unit, 46...display control unit
Claims
1. one or more processors; the one or more processors: Acquire a speech image including the speaker; acquiring display information for displaying the speech content of the speaker; Controlling the display information so that it is displayed in a specific area within the speech image that does not overlap with the face of the speaker; Information processing system.
2. the one or more processors: The information processing system according to claim 1 , wherein the specific area is determined to be an area having a size determined based on the display information.
3. The information processing system according to claim 2 , wherein the size determined based on the display information is a size in which the display information can be arranged.
4. The information processing system according to claim 2 , wherein the size determined based on the display information is a size that allows the display information to be arranged by reducing the size of the display information within a visible range.
5. the one or more processors: The information processing system according to claim 1 , wherein an area having a characteristic defined based on the display information is determined as the specific area.
6. The information processing system according to claim 5 , wherein the area having the characteristic determined based on the display information is an area in which the visibility of the display information when the display information is arranged is higher than a predetermined visibility standard.
7. The information processing system of claim 5, wherein the area having the characteristics determined based on the display information is an area in which the amount of change of the display information to increase the visibility of the display information above a predetermined visibility standard is less than a predetermined standard amount.
8. the one or more processors: acquiring the speech image further including another speaker; Further acquiring other display information for displaying the speech content of the other speaker; controlling the display information and the other display information to be displayed in a plurality of areas in the speech image that do not overlap with either the face of the speaker or the faces of the other speakers; The information processing system according to claim 1 .
9. the plurality of regions include a first region that is within a predetermined reference distance from the speaker and the other speakers, and a second region that is within the reference distance from the speaker but not within the reference distance from the other speakers, the one or more processors:
9. The information processing system of claim 8, wherein the first area is determined to be the specific area under the condition that the visibility of the display information when placed in the first area is higher than a predetermined visibility standard, and the visibility of the display information when placed in the second area is lower than the visibility standard.
10. the plurality of regions further includes a third region that is not within the reference distance from the speaker but is within the reference distance from the other speaker; the one or more processors: The information processing system of claim 9, wherein the first area is determined to be the specific area under the further condition that the visibility of the other display information when placed in the first area and the visibility of the other display information when placed in the third area are higher than the visibility standard.
11. the one or more processors: The information processing system of claim 1, wherein prior to displaying the display information in the specific area within the speech image, the display information is modified to improve visibility of the display information when placed in the specific area.
12. On the computer, A function of acquiring a speech image including the speaker; a function of acquiring display information for displaying the speech content of the speaker; a function of controlling the display information to be displayed in a specific area within the speech image that does not overlap with the face of the speaker; A program to achieve this.
Citation Information
Patent Citations
Augmented sense of reality system and color compensation method thereof
JP2020017252A
Electronic devices and display methods
JP2022089884A
Information processing device, information processing method, program, and information processing system
WO2023074126A1