Information processing system and program
The information processing system addresses the challenge of intuitively displaying supplementary images by aligning their two-dimensional positions with their three-dimensional locations, improving recognition of the target being supplemented.
Patent Information
- Application Number
- JP2024094679
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-23
AI Technical Summary
Existing systems struggle to intuitively display supplementary images in three-dimensional space within an utterance image, making it difficult to recognize which target the supplementary image is supplementing.
An information processing system that acquires a speech image and controls the display of supplementary images, such as sign language or explanatory images, at two-dimensional positions corresponding to their three-dimensional locations within the speech image.
Enhances intuitive recognition of which target the supplementary image is supplementing by aligning the display position with the three-dimensional position of the object, allowing for accurate and intuitive supplementation of speech content.
Smart Images

Figure 2025186092000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system and a program. [Background technology]
[0002] US Patent No. 6,299,949 describes a sensory eyewear system that can recognize and interpret sign language and present translated information to a user of a mixed reality device. Patent document 2 describes a video device that includes a sign language image generation means that converts the semantic content recognized by a voice recognition means into an animated image in sign language, and a display means that displays the image information generated by the sign language image generation means on a screen. Patent document 3 describes an augmented reality system in which an augmented reality terminal operates in a first mode for generating virtual information and a second mode for presenting an augmented reality image, and includes a position information acquisition means for acquiring position information of the augmented reality terminal, a virtual information generation means for generating virtual information based on information input by a user, a transmission means for associating the generated virtual information with position information at the time of generation and transmitting the information to an information processing device, a receiving means for receiving virtual information from the information processing device based on the position information in the second mode, a virtual image generation means for generating an image of a floating body based on the virtual information, and an augmented reality image display control means for controlling the display of the augmented reality image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2019-535059 [Patent Document 2] Japanese Patent Application Publication No. 07-191599 [Patent Document 3] Japanese Patent Publication No. 2020-077187 Summary of the Invention [Problem to be solved by the invention]
[0004] In some cases, a supplementary image is displayed in an utterance image including the speaker to supplement the content of the speaker's speech. In this case, it is possible to adopt a configuration in which the supplementary image is displayed at a two-dimensional position within the utterance image without taking into account the three-dimensional position of the target to be supplemented within the utterance image. However, if such a configuration is adopted, it may be difficult to intuitively recognize which target the supplementary image is supplementing.
[0005] An object of the present invention is to increase the possibility that a supplementary image for supplementing the content of a speaker's speech can intuitively recognize which supplementary target the image is supplementing. [Means for solving the problem]
[0006] The invention described in claim 1 is an information processing system that includes one or more processors, which acquire a speech image including a speaker, acquire a supplementary image for supplementing the speech content of the speaker, which is based on the speech content but is not a text that directly represents the speech content, and control the supplementary image to be displayed at a two-dimensional position within the speech image that corresponds to the three-dimensional position of the supplementary object within the speech image. The invention described in claim 2 is the information processing system described in claim 1, wherein the supplementary target is the speaker, and the supplementary image is a sign language image representing sign language based on the speech content. The invention described in claim 3 is the information processing system described in claim 2, wherein the sign language image is a video image of a sign language interpreter performing sign language actions based on the spoken content. The invention described in claim 4 is an information processing system described in claim 2, in which the sign language image is a moving image of sign language actions generated from the content of the speech without requiring a sign language interpreter to perform sign language actions based on the content of the speech. The invention described in claim 5 is an information processing system described in claim 2, in which the sign language image is a moving image of sign language actions performed by a virtual character generated from sign language actions performed by a sign language interpreter based on the speech content. The invention described in claim 6 is the information processing system described in claim 2, wherein the two-dimensional position within the speech image is a position where the hand of the speaker is displayed within the speech image. The invention described in claim 7 is an information processing system described in claim 1, wherein the supplementary target is an object included in the utterance content, and the supplementary image is an explanatory image that represents an explanation of the object included in the utterance content. The invention described in claim 8 is an information processing system described in claim 7, in which the explanatory image is an image indicating the object generated based on a demonstrative term indicating the object included in the utterance content. The invention described in claim 9 is an information processing system described in claim 7, wherein the explanatory image is an image representing the state of the object generated based on a state word representing the state of the object included in the utterance content. The invention described in claim 10 is the information processing system described in claim 7, wherein the two-dimensional position within the speech image is a position within a predetermined range around the supplementary target within the speech image. The invention described in claim 11 is a program for enabling a computer to realize the following functions: a function for acquiring a speech image including a speaker; a function for acquiring a supplementary image for supplementing the speech content of the speaker, the supplementary image being based on the speech content but not being a text that directly represents the speech content; and a function for controlling the display of the supplementary image at a two-dimensional position within the speech image that corresponds to the three-dimensional position where the supplementary object exists within the speech image. [Effects of the Invention]
[0007] According to the invention of claim 1, it is more likely that a user will be able to intuitively recognize which supplementary target a supplementary image for supplementing the content of a speaker's speech is supplementing. According to the invention of claim 2, it is more likely that a user will be able to intuitively recognize which speaker's speech content a sign language image is supplementing. According to the invention of claim 3, it is possible to display sign language images even without a function for generating moving images of sign language actions made by a virtual person from the sign language actions made by a sign language interpreter. According to the invention of claim 4, it is possible to display a sign language image without the sign language interpreter actually making sign language movements based on the content of the speech. According to the invention of claim 5, by using the video of the sign language actions performed by the sign language interpreter as is, it is possible to display the sign language image without placing a psychological burden on the sign language interpreter. According to the invention of claim 6, it is possible to make it appear as if the speaker is making sign language movements even if the speaker is not actually making sign language movements. According to the invention of claim 7, it is more likely that the speaker will be able to intuitively recognize which object included in the speaker's speech content the explanatory image complements. According to the invention of claim 8, the possibility of intuitively recognizing which object the explanatory image complements is further increased. According to the invention of claim 9, the possibility of intuitively recognizing what state of the object the explanatory image represents increases. According to the invention of claim 10, the relationship between the explanatory image and the object can be made easier to see. According to the invention of claim 11, it is more likely that a user will be able to intuitively recognize which supplementary target a supplementary image for supplementing the content of a speaker's speech is supplementing. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of an AR system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating an example of the hardware configuration of AR glasses according to the present embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a conceptual configuration of an AR module according to the present embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of the hardware configuration of an AR server according to the present embodiment. [Figure 5] FIG. 1 is a diagram illustrating a schematic operation of an AR system according to a first embodiment. [Figure 6] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a second embodiment. [Figure 7]FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a third embodiment. [Figure 8] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a fourth embodiment. [Figure 9] FIG. 10 is a diagram illustrating a schematic operation of an AR system according to a fifth embodiment. [Figure 10] FIG. 2 is a block diagram illustrating an example of a functional configuration of an AR server according to the present embodiment. [Figure 11] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the first aspect. [Figure 12] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the second aspect. [Figure 13] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the third aspect. [Figure 14] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the fourth aspect. [Figure 15] 10 is a flowchart illustrating an example of the operation of the AR server in the AR system according to the fifth aspect. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, the present embodiment will be described in detail with reference to the accompanying drawings.
[0010] (Outline of this embodiment) This embodiment provides an information processing system that acquires a speech image including a speaker, acquires a supplementary image to supplement the speaker's speech content, which is based on the speech content but is not a text that directly represents the speech content, and controls the display of the supplementary image at a two-dimensional position within the speech image that corresponds to the three-dimensional position where the supplementary object is located within the speech image.
[0011] Here, the "system" may be configured by a single device or multiple devices. In the following, an information processing system configured by a single device will be taken as an example. The single device will be described as an AR server in an AR (Augmented Reality) system.
[0012] (Overall configuration of the AR system) 1 is a diagram showing an example of the overall configuration of an AR system 1 according to the present embodiment. As shown in the figure, the AR system 1 includes AR glasses 10, an AR server 30, and a communication line 80. Although only one AR glass 10 is shown in the figure, there may be multiple AR glasses 10.
[0013] The AR glasses 10 are a glasses-type wearable terminal device. Here, "wearable" means that the user can wear them. Therefore, a glasses-type wearable terminal device is a computer device that has the shape of glasses and can actually be worn on the user's head. The AR glasses 10 are a device that enables the user to see AR. Here, "AR" stands for "Augmented Reality," and refers to displaying a virtual screen superimposed on real space for the user. In other words, the user can view the virtual screen through the AR glasses 10, and can also view real space through the AR glasses 10. In this case, the "virtual screen" refers to an image that is created by a computer and can be viewed through the AR glasses 10. And the "real space" refers to a space that actually exists.
[0014] The AR glasses 10 have two cameras 11 attached to both ends of the front part of the frame. In this embodiment, a two-dimensional image is assumed as the augmented reality image (hereinafter also referred to as "AR image"), but a three-dimensional image may also be used. A three-dimensional image refers to an image in which distance information is recorded for each pixel, and is also called a "distance image." To acquire a three-dimensional image, a stereo camera may be used as the camera 11. Alternatively, a LiDAR (Light Detection and Ranging) may be used to acquire a three-dimensional image.
[0015] Although the AR glasses 10 are shown here as an eyeglass-type device, the present invention is not limited to this. Any shape or format of device that displays AR may be used. Specifically, a broader optically transmissive display may be used. For example, MR (Mixed Reality) glasses may be used instead of AR glasses.
[0016] The AR server 30 is a server computer that performs processing to display information on the AR glasses 10. Specifically, the AR server 30 generates information to be displayed on the AR glasses 10 and outputs this information to a microdisplay 122 (described later) of the AR glasses 10.
[0017] The communication line 80 is a line used for information communication between the AR glasses 10 and the AR server 30. For example, a wireless local area network (LAN) or the Internet may be used as the communication line 80. Alternatively, for example, a mobile communication system such as 4G or 5G, or Bluetooth (registered trademark) may be used as the communication line 80.
[0018] (AR glasses hardware configuration) 2 is a diagram showing an example of the hardware configuration of the AR glasses 10 according to the present embodiment. As shown in the figure, the AR glasses 10 include a data processing unit 100. The AR glasses 10 also include a camera 11, an AR module 120, a microphone 130, and a speaker 140. The AR glasses 10 also include a communication module 150.
[0019] The data processing unit 100 includes a processor 101. The data processing unit 100 further includes a read only memory (ROM) 102 and a random access memory (RAM) 103. The data processing unit 100 further includes a flash memory 104. The processor 101 is configured by, for example, a CPU (Central Processing Unit), and realizes various functions by executing programs. The ROM 102, RAM 103, and flash memory 104 are all semiconductor memories. The ROM 102 stores a BIOS (Basic Input Output System) and the like. The RAM 103 is a main storage device used to execute programs. For example, a DRAM (Dynamic RAM) is used as the RAM 103. The flash memory 104 is used to record firmware, programs, data files, etc. The flash memory 104 is used as an auxiliary storage device.
[0020] The camera 11 captures an image in front of the user's field of view. The viewing angle of the camera 11 may be approximately the same as or greater than the viewing angle of a person. For example, a CMOS image sensor or a CCD image sensor is used as the camera 11. The number of cameras 11 may be one or more. In the example of FIG. 1, there are two cameras 11. In this case, the two cameras 11 may be arranged, for example, at both ends of the front part of the frame. Using two cameras 11 enables stereo imaging. This makes it possible to measure the distance to the subject and estimate the front-to-back relationship between the subjects.
[0021] The AR module 120 is a module that realizes the visual recognition of augmented reality by combining an AR image with a real landscape. The AR module 120 is composed of optical components and electronic components. Representative methods for the AR module 120 include the following: First, a method in which a half mirror is placed in front of the user's eye. Second, a method in which a volume hologram is placed in front of the user's eye. Third, a method in which a blazed diffraction grating is placed in front of the user's eye.
[0022] The microphone 130 is a device that converts the user's voice and surrounding sounds into electrical signals. The speaker 140 is a device that converts an electrical signal into sound and outputs the sound. The speaker 140 may be a bone conduction speaker or a cartilage conduction speaker. The speaker 140 may be a device independent of the AR glasses 10, such as a wireless earphone. In this case, the speaker 140 is connected to the AR glasses 10 via Bluetooth (registered trademark) or the like.
[0023] The communication module 150 is a device that complies with a protocol used for communication via the communication line 80. The communication module 150 may also be a device that complies with a protocol used for communication with other external devices. Protocols used for communication with external devices include, for example, Wifi (registered trademark) and Bluetooth (registered trademark).
[0024] Although not shown, the AR glasses 10 may be additionally provided with an inertial sensor, a positioning sensor, a vibrator, and the like.
[0025] 3 is a diagram illustrating a conceptual configuration example of AR module 120 according to this embodiment. AR module 120 shown in FIG. 3 corresponds to a method in which a blazed diffraction grating is placed in front of the user's eyes. AR module 120 shown in Fig. 3 includes a light guide plate 121 and a microdisplay 122. AR module 120 shown in Fig. 3 also includes a diffraction grating 123A to which image light L2 is input. AR module 120 shown in Fig. 3 also includes a diffraction grating 123B from which image light L2 is output.
[0026] The light guide plate 121 corresponds to a lens of glasses. The light guide plate 121 has a transmittance of, for example, 85% or more. Therefore, the user can directly view the scenery ahead through the light guide plate 121. The external light L1 travels straight through the light guide plate 121 and the diffraction grating 123B and enters the eye E of the user.
[0027] The microdisplay 122 is a display device that displays an AR image that is visually recognized by the user. Light of the AR image displayed on the microdisplay 122 is projected onto the light guide plate 121 as image light L2. The image light L2 is refracted by the diffraction grating 123A and reaches the diffraction grating 123B while reflecting inside the light guide plate 121. The diffraction grating 123B refracts the image light L2 toward the eye E of the user. As a result, external light L1 and image light L2 are simultaneously incident on the user's eye E. As a result, the user recognizes the presence of an AR image in front of the user's line of sight.
[0028] (AR server hardware configuration) 4 is a diagram illustrating an example of the hardware configuration of the AR server 30 according to this embodiment. As illustrated, the AR server 30 includes a data processing unit 300. The AR server 30 further includes an HDD (hard disk drive) 310 and a communication module 320.
[0029] The data processing unit 300 includes a processor 301. The data processing unit 300 further includes a ROM 302 and a RAM 303. The processor 301 is configured by, for example, a CPU. The processor 301 realizes various functions by executing programs. The ROM 302 and the RAM 303 are both semiconductor memories. The ROM 302 stores the BIOS and the like. The RAM 303 is used as a main storage device used for executing programs. The RAM 303 may be, for example, a DRAM.
[0030] The HDD 310 is an auxiliary storage device that uses a magnetic disk as a recording medium. In this embodiment, the HDD 310 is used as the auxiliary storage device. However, a non-volatile rewritable semiconductor memory may also be used as the auxiliary storage device. An operating system and application programs are installed on the HDD 310.
[0031] The communication module 320 is a device that complies with the protocol used for communication over the communication line 80 .
[0032] Although not shown, the AR server 30 may be additionally provided with a display, a keyboard, a mouse, and the like.
[0033] (Outline of AR system operation) ((First Aspect)) FIG. 5 is a diagram showing a schematic operation of the AR system 1 according to the first embodiment. In FIG. 5, it is assumed that a speaker U is speaking, and a listener L wearing AR glasses 10 is listening to the speech. At this time, a background image 200 including the speaker U is viewed through the AR glasses 10 (step 211). Audio information of the speaker U's speech is transmitted to the sign language interpreter S (step 212). The sign language interpreter S then provides sign language interpretation in real time based on the transmitted audio information (step 213). At this time, the camera 205 captures the actions of the sign language interpreter S to acquire a sign language video (step 214). As a result, the part of the sign language video below the face of the speaker U in the background image 200 is synthesized as a sign language image 201 below the face of the speaker U (step 215). In this case, the sign language image 201 is synthesized as if the sign language interpretation were being performed at the three-dimensional position of the speaker U. Note that text information 202 obtained by voice recognition of the audio information of the speaker U's speech may be further synthesized with the background image 200.
[0034] ((Second Aspect)) FIG. 6 is a diagram showing a schematic operation of the AR system 1 according to the second embodiment. 6, it is assumed that a speaker U is speaking and a listener L wearing AR glasses 10 is listening to the speech. At this time, a background image 220 including the speaker U is viewed through the AR glasses 10 (step 231). The AR server 30 also acquires audio information of the speaker U's speech, or acquires text information by performing voice recognition on the audio information (step 232). The AR server 30 then automatically generates a sign language animation A based on the audio information or text information (step 233). As a result, the part of the sign language animation A below the face of the speaker U in the background image 220 is synthesized as a sign language image 221 below the face of the speaker U (step 234). In this case, the sign language image 221 is synthesized as if sign language interpretation were being performed at the three-dimensional position of the speaker U. Note that text information 222 obtained by voice recognition of the audio information of the speaker U's speech may also be synthesized into the background image 220.
[0035] ((Third Aspect)) FIG. 7 is a diagram showing a schematic operation of the AR system 1 according to the third embodiment. 7, it is assumed that a speaker U is speaking and a listener L wearing AR glasses 10 is listening to the speech. At this time, a background image 240 including the speaker U is viewed through the AR glasses 10 (step 251). Audio information of the speaker U's speech is transmitted to the sign language interpreter S (step 252). The sign language interpreter S then provides sign language interpretation in real time based on the transmitted audio information (step 253). At this time, the camera 245 measures the three-dimensional movements of the sign language interpreter S (step 254). An avatar V reflecting this three-dimensional movement is then generated (step 255). As a result, the part below the face of the avatar V is synthesized as a sign language image 241 below the face of the speaker U in the background image 240. In this case, the sign language image 241 is synthesized as if the sign language interpretation were being performed at the three-dimensional position of the speaker U. Note that character information 242 obtained by voice recognition of the audio information of the speaker U's speech may be further synthesized into the background image 240.
[0036] ((Fourth Aspect)) FIG. 8 is a diagram showing the general operation of the AR system 1 according to the fourth embodiment. In FIG. 8, it is assumed that a speaker U is speaking about a bag B, and a listener L wearing the AR glasses 10 is listening to the speech. At this time, a background image 260 including the speaker U and the bag B is visible through the AR glasses 10 (step 271). Then, the AR glasses 10 transmit a captured image 265 corresponding to the background image 260 to the AR server 30 (step 272). The AR server 30 also acquires audio information of the speaker U's speech, and performs voice recognition on this audio information to acquire text information (step 273).
[0037] The AR server 30 then acquires the noun "bag" that appears in the utterance from the acquired text information. The AR server 30 also detects demonstratives and state words related to the noun from the text information. Here, demonstratives are words that indicate the object represented by the noun, such as "this." State words are words that express the state of the object represented by the noun, such as "new." The AR server 30 then recognizes objects from the captured image 265 and identifies the object represented by the noun acquired from the text information. That is, the AR server 30 determines which object the noun represents based on the text information and the object recognition result (step 274).
[0038] Thereafter, the AR server 30 generates an explanatory image 261 based on the demonstrative term and the state term acquired from the text information (step 275). Here, the explanatory image 261 is an illustration image that illustrates the instruction expressed by the demonstrative term and the state expressed by the state term. As a result, the explanatory image 261 is synthesized near the bag B in the background image 260 (step 276). In this case, the explanatory image 261 is synthesized as if it were located near the three-dimensional position of the bag B. Note that the background image 260 may further be synthesized with text information 262 obtained by voice recognition of the voice information of the speaker U's utterance.
[0039] ((Fifth Aspect)) FIG. 9 is a diagram showing a schematic operation of the AR system 1 according to the fifth embodiment. 9, it is assumed that a speaker U is speaking about a bag B, and a listener L wearing the AR glasses 10 is listening to the speech. At this time, a background image 280 including the speaker U and the bag B is visible through the AR glasses 10 (step 291). Then, the AR glasses 10 transmit a captured image 285 corresponding to the background image 280 to the AR server 30 (step 292). The AR server 30 also acquires audio information of the speaker U's speech, and performs voice recognition on this audio information to acquire text information (step 293).
[0040] The AR server 30 then acquires the noun "bag" that appears in the speech from the acquired text information. The AR server 30 also detects a state word related to the noun from the text information. Here, a state word is a word that represents the state of the object represented by the noun, as described above. The AR server 30 then recognizes the object from the captured image 285 and identifies the object represented by the noun acquired from the text information. That is, the AR server 30 determines which object the noun represents based on the text information and the object recognition result (step 294).
[0041] Thereafter, the AR server 30 generates an explanatory image 281 based on the state word acquired from the text information (step 295). Here, the explanatory image 281 is an onomatopoeia image that expresses the state represented by the state word with onomatopoeia. As a result, the explanatory image 281 is synthesized near the bag B in the background image 280 (step 296). In this case, the explanatory image 281 is synthesized as if it were also located near the three-dimensional position of the bag B. Note that the background image 280 may further be synthesized with text information 282 obtained by voice recognition of the voice information of the speaker U's utterance.
[0042] (AR server functional configuration) 10 is a block diagram showing an example of a functional configuration of the AR server 30 according to this embodiment. As shown in the figure, the AR server 30 includes a captured image acquisition unit 41, an audio information acquisition unit 42, and a display information acquisition unit 43. The AR server 30 also includes a three-dimensional position identification unit 44 and a display control unit 45.
[0043] The captured image acquisition unit 41 acquires a captured image including the speaker captured by the camera 11 of the AR glasses 10. Alternatively, the captured image acquisition unit 41 may acquire a captured image including the speaker captured by a camera provided other than the AR glasses 10. In the first to third aspects, the captured image acquisition unit 41 acquires a captured image including the speaker U. In this case, the captured image including the speaker U is an example of an utterance image including a speaker. Furthermore, the processing of the captured image acquisition unit 41 is an example of acquiring an utterance image. In the fourth and fifth aspects, the captured image acquisition unit 41 acquires a captured image including the speaker U and the bag B. In this case, the captured image including the speaker U and the bag B is an example of a speech image including a speaker. Furthermore, the processing of the captured image acquisition unit 41 is an example of acquiring a speech image.
[0044] The audio information acquisition unit 42 acquires audio information including the speaker's voice collected by the microphone 130 of the AR glasses 10. Alternatively, the audio information acquisition unit 42 may acquire audio information including the speaker's voice collected by a microphone provided outside the AR glasses 10.
[0045] The display information acquisition unit 43 acquires display information for displaying the content of the speaker's utterance based on the audio information acquired by the audio information acquisition unit 42. Here, the display information may be a supplemental image for supplementing the content of the speaker's utterance. However, the supplemental image is not a direct text representation of the content of the speaker's utterance. For example, the supplemental image is not the voice recognition result of the content of the speaker's utterance. In this case, the processing of the display information acquisition unit 43 is an example of acquiring a supplemental image for supplementing the content of the speaker's utterance, which is based on the content of the utterance but is not a text representation of the content of the utterance directly.
[0046] In the first to third aspects, the display information acquisition unit 43 acquires, as display information, a sign language image that represents the content of an utterance by a speaker in sign language. In this case, the processing of the display information acquisition unit 43 is an example of acquiring a sign language image that represents the sign language based on the content of an utterance.
[0047] In particular, in the first mode, the display information acquisition unit 43 transmits the audio information acquired by the audio information acquisition unit 42 to the sign language interpreter. As a result, the sign language interpreter performs sign language interpretation in real time. Then, the camera 205 (see FIG. 5) captures the actions of the sign language interpreter and transmits the video. As a result, the display information acquisition unit 43 receives the video. In this case, the processing of the display information acquisition unit 43 is an example of acquiring a video image capturing the sign language actions performed by the sign language interpreter based on the spoken content. In a second aspect, the display information acquisition unit 43 generates sign language animation based on the audio information acquired by the audio information acquisition unit 42. In this case, the display information acquisition unit 43 automatically generates the sign language animation without being based on the sign language interpretation actions of a sign language interpreter. The display information acquisition unit 43 may automatically generate the sign language animation using, for example, the method described in Patent Document 2. In this case, the processing by the display information acquisition unit 43 is an example of acquiring a moving image of sign language actions generated from the content of an utterance without requiring the sign language actions performed by a sign language interpreter based on the content of an utterance. Furthermore, in a third aspect, the display information acquisition unit 43 transmits the audio information acquired by the audio information acquisition unit 42 to the sign language interpreter. As a result, the sign language interpreter performs sign language interpretation in real time. Then, the camera 245 (see FIG. 7) measures the movements of the sign language interpreter and transmits the measurement results. As a result, the display information acquisition unit 43 receives the measurement results. Then, based on the measurement results, the display information acquisition unit 43 generates an avatar that reflects the movements of the sign language interpreter. In this case, the processing of the display information acquisition unit 43 is an example of acquiring a moving image of the sign language movements made by a virtual character generated from the sign language movements made by the sign language interpreter based on the spoken content.
[0048] In the fourth and fifth aspects, the display information acquisition unit 43 acquires, as display information, an explanatory image for explaining an object. Here, the object is included in the speech content of the speaker. Alternatively, the object may also be included in the captured image acquired by the captured image acquisition unit 41. In this case, the processing of the display information acquisition unit 43 is an example of acquiring an explanatory image that shows an explanation of the object included in the speech content. Specifically, the display information acquisition unit 43 may acquire, for example, an image indicating an object as the explanatory image. In this case, the display information acquisition unit 43 may generate the image indicating the object based on a demonstrative term included in the utterance content of the speaker. In this case, the processing of the display information acquisition unit 43 is an example of acquiring an image indicating the object generated based on a demonstrative term indicating the object included in the utterance content. Furthermore, the display information acquisition unit 43 acquires, for example, an image representing the state of an object as the explanatory image. In this case, the display information acquisition unit 43 may generate the image representing the state of the object based on a state word included in the utterance content of the speaker. In this case, the processing of the display information acquisition unit 43 is an example of acquiring an image representing the state of the object generated based on a state word representing the state of the object included in the utterance content.
[0049] In particular, in the fourth aspect, the display information acquisition unit 43 extracts a noun from the audio information acquired by the audio information acquisition unit 42. Then, the display information acquisition unit 43 extracts a demonstrative term and a state term related to the noun from the audio information acquired by the audio information acquisition unit 42. As a result, the display information acquisition unit 43 generates an illustration image based on the demonstrative term and the state term as an explanatory image. In a second mode, the display information acquisition unit 43 extracts a noun from the audio information acquired by the audio information acquisition unit 42. Then, the display information acquisition unit 43 extracts a state word related to the noun from the audio information acquired by the audio information acquisition unit 42. As a result, the display information acquisition unit 43 generates an onomatopoeia image based on the state word as an explanatory image.
[0050] The three-dimensional position specifying unit 44 specifies the three-dimensional position of an object included in the captured image acquired by the captured image acquisition unit 41. At this time, the three-dimensional position specifying unit 44 calculates the distance to the object using a stereo camera as the camera 11 of the AR glasses 10. In the first to third aspects, the three-dimensional position identifying unit 44 identifies the three-dimensional position of the speaker U included in the captured image. At that time, the three-dimensional position identifying unit 44 calculates the distance from the AR glasses 10 to the speaker U. In the fourth and fifth aspects, the three-dimensional position identifying unit 44 recognizes an object represented by the noun obtained from the display information from the captured image. The three-dimensional position identifying unit 44 also identifies the three-dimensional position of the object included in the captured image. At this time, the three-dimensional position identifying unit 44 calculates the distance from the AR glasses 10 to the object.
[0051] The display control unit 45 controls the display information acquired by the display information acquisition unit 43 to be displayed on the AR glasses 10. In this case, the display control unit 45 displays the display information so that it exists at the three-dimensional position of the object identified by the three-dimensional position identification unit 44. For example, the display control unit 45 controls the display information to be displayed at a two-dimensional position corresponding to the three-dimensional position of the object. Here, the two-dimensional position may be a two-dimensional position on a background image seen through the AR glasses 10. Furthermore, if the distance to the object is far, the display control unit 45 may display the display information smaller in accordance with the distance. If the distance to the object is close, the display control unit 45 may display the display information larger in accordance with the distance. In this case, the processing of the display control unit 45 is an example of control to display a supplementary image at a two-dimensional position in the utterance image corresponding to the three-dimensional position where the supplementary target exists in the utterance image.
[0052] In the first to third aspects, the display control unit 45 displays the sign language image so that it exists at the three-dimensional position of the speaker U. For example, the display control unit 45 controls the display of the sign language image at a two-dimensional position corresponding to the three-dimensional position of the speaker U's hand. Furthermore, if the distance to the speaker U is long, the display control unit 45 may display the sign language image smaller in accordance with the distance. If the distance to the speaker U is short, the display control unit 45 may display the sign language image larger in accordance with the distance. In this case, the processing of the display control unit 45 is an example of controlling the display of a supplementary image at a two-dimensional position in the speech image corresponding to the three-dimensional position where the speaker exists in the speech image. Furthermore, the processing of the display control unit 45 is an example of controlling the display of a supplementary image at a position where the speaker's hand is displayed in the speech image.
[0053] In the fourth and fifth aspects, the display control unit 45 displays the explanatory image so that it exists at the three-dimensional position of an object mentioned in the utterance of the speaker U. For example, the display control unit 45 controls the explanatory image to be displayed at a two-dimensional position corresponding to the three-dimensional position of the object. Furthermore, if the distance to the object is long, the display control unit 45 may display the explanatory image smaller in accordance with the distance. If the distance to the object is short, the display control unit 45 may display the explanatory image larger in accordance with the distance. In this case, the processing of the display control unit 45 is an example of controlling the display of a supplementary image at a two-dimensional position in the utterance image corresponding to the three-dimensional position of an object included in the utterance content in the utterance image. Furthermore, the processing of the display control unit 45 is an example of controlling the display of a supplementary image at a position within a predetermined range around the supplementary target in the utterance image.
[0054] (AR server operation) ((First Aspect)) FIG. 11 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the first aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 401). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 402).
[0055] Next, the display information acquisition unit 43 transmits this audio information to the sign language interpreter (step 403). The sign language interpreter then performs sign language actions based on the audio information, and the camera 205 (see FIG. 5) acquires a sign language image. As a result, the display information acquisition unit 43 receives this sign language image from the camera 205 (step 404).
[0056] Next, the three-dimensional position specifying unit 44 specifies the three-dimensional position of the speaker based on the captured image (step 405). Thereafter, the display control unit 45 controls the AR glasses 10 to display the sign language video at the three-dimensional position of the speaker (step 406). Specifically, the display control unit 45 controls the display so that the lower part of the sign language video is displayed at a three-dimensional position lower than the speaker's face. As a result, the sign language video can be viewed on the AR glasses 10 as if the speaker were using sign language.
[0057] ((Second Aspect)) FIG. 12 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the second embodiment. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 421). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 422).
[0058] Next, the display information acquisition unit 43 generates sign language animation based on this audio information (step 423).
[0059] Next, the three-dimensional position specifying unit 44 specifies the three-dimensional position of the speaker based on the captured image (step 424). Thereafter, the display control unit 45 controls the AR glasses 10 to display the sign language animation at the three-dimensional position of the speaker (step 425). Specifically, the display control unit 45 controls the display to display the lower part of the sign language animation at a three-dimensional position lower than the speaker's face. As a result, the sign language animation can be viewed on the AR glasses 10 as if the speaker were using sign language.
[0060] ((Third Aspect)) FIG. 13 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the third aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 441). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 442).
[0061] Next, the display information acquisition unit 43 transmits this audio information to the sign language interpreter (step 443). The sign language interpreter then performs sign language movements based on the audio information, and the camera 245 (see FIG. 7) measures the sign language movements. As a result, the display information acquisition unit 43 receives the measurement results of the sign language movements from the camera 245 (step 444). Then, the display information acquisition unit 43 generates an avatar that reflects the sign language movements based on the measurement results (step 445).
[0062] Next, the three-dimensional position specifying unit 44 specifies the three-dimensional position of the speaker based on the captured image (step 446). Thereafter, the display control unit 45 controls the AR glasses 10 to display the avatar at the three-dimensional position of the speaker (step 447). Specifically, the display control unit 45 controls the AR glasses 10 to display the lower part of the avatar's face at a three-dimensional position lower than the speaker's face. This allows the avatar to be seen through the AR glasses 10 as if the speaker were using sign language.
[0063] ((Fourth Aspect)) FIG. 14 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the fourth aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 461). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 462).
[0064] Next, the display information acquisition unit 43 performs voice recognition on the voice information to acquire character information (step 463). The display information acquisition unit 43 also acquires a noun from the character information (step 464). Furthermore, the display information acquisition unit 43 acquires a demonstrative term and a state term related to the noun from the character information (step 465). As a result, the display information acquisition unit 43 generates an illustration image corresponding to the demonstrative term and state term (step 466).
[0065] Next, the three-dimensional position specifying unit 44 recognizes the object represented by the noun acquired in step 464 from the captured image (step 467). The three-dimensional position specifying unit 44 also specifies the three-dimensional position of the object (step 468). Thereafter, the display control unit 45 controls the AR glasses 10 to display the illustration image at the three-dimensional position of the object (step 469). Specifically, the display control unit 45 controls the AR glasses 10 to display the illustration image at a three-dimensional position around the object. This allows the AR glasses 10 to see the illustration image as if it were located at the position of the object.
[0066] ((Fifth Aspect)) FIG. 15 is a flowchart showing an example of the operation of the AR server 30 in the AR system 1 of the fifth aspect. As shown in the figure, first, the captured image acquisition unit 41 acquires a captured image including a speaker from the AR glasses 10 (step 481). Next, the voice information acquisition unit 42 acquires voice information including the voice of the speaker from the AR glasses 10 (step 482).
[0067] Next, the display information acquisition unit 43 performs voice recognition on the voice information to acquire character information (step 483). The display information acquisition unit 43 also acquires a noun from the character information (step 484). Furthermore, the display information acquisition unit 43 acquires a state word related to the noun from the character information (step 485). As a result, the display information acquisition unit 43 generates an onomatopoeia image corresponding to the state word (step 486).
[0068] Next, the three-dimensional position specifying unit 44 recognizes the object represented by the noun acquired in step 484 from the captured image (step 487). The three-dimensional position specifying unit 44 also specifies the three-dimensional position of the object (step 488). Thereafter, the display control unit 45 controls the AR glasses 10 to display the onomatopoeia image at the three-dimensional position of the object (step 489). Specifically, the display control unit 45 controls the AR glasses 10 to display the onomatopoeia image at a three-dimensional position around the object. This allows the onomatopoeia image to be seen on the AR glasses 10 as if it were located at the position of the object.
[0069] (Processor) In this embodiment, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Furthermore, the operations of the processor in this embodiment may not only be performed by one processor, but may also be performed by multiple processors located at physically separate locations working together. Furthermore, the order of the operations of the processor is not limited to the order described in this embodiment, and may be changed.
[0070] (program) The present embodiment can be applied to a program and a program product. For example, a program to which this embodiment is applied can be understood as a program that enables a computer to realize the following functions: a function to acquire a speech image including the speaker; a function to acquire a supplementary image to supplement the speech content of the speaker, which is based on the speech content but is not text that directly represents the speech content; and a function to control the display of the supplementary image at a two-dimensional position within the speech image that corresponds to the three-dimensional position where the supplementary object exists within the speech image. The program to which this embodiment is applied can be provided not only by communication means but also by being stored in a recording medium such as a CD-ROM.
[0071] (Addendum) (((1))) one or more processors; the one or more processors: Acquire a speech image including the speaker; acquiring a supplemental image for supplementing the speech content of the speaker, the supplemental image being based on the speech content but not being a text directly representing the speech content; Control is performed so that the supplementary image is displayed at a two-dimensional position within the speech image corresponding to a three-dimensional position where a supplementary target exists within the speech image. Information processing system. (((2))) the target to be supplemented is the speaker, The information processing system according to (((1))), wherein the supplemental image is a sign language image representing sign language based on the speech content. (((3))) The information processing system according to (((2))), wherein the sign language image is a video image of a sign language interpreter performing sign language actions based on the speech content. (((4))) The information processing system described in (((2))), wherein the sign language image is a moving image of sign language actions generated from the speech content without requiring a sign language interpreter to perform sign language actions based on the speech content. (((5))) The information processing system according to (((2))), wherein the sign language image is a moving image of sign language actions made by a virtual character generated from sign language actions made by a sign language interpreter based on the speech content. (((6))) The information processing system according to any one of (((2))) to (((5))), wherein the two-dimensional position within the speech image is a position where the speaker's hand is displayed within the speech image. (((7))) the supplementation target is an object included in the utterance content, The information processing system according to any one of ((1))) to (((6))), wherein the supplemental image is an explanatory image that shows an explanation of the object included in the speech content. (((8))) The information processing system according to (((7))), wherein the explanatory image is an image indicating the object that is generated based on a demonstrative term that indicates the object and is included in the speech content. (((9))) The information processing system according to (((7))), wherein the explanatory image is an image representing the state of the object, generated based on a state word representing the state of the object, which is included in the utterance content. (((10))) An information processing system according to any one of ((7))) to ((9))), wherein the two-dimensional position within the speech image is a position within a predetermined range around the target to be supplemented within the speech image. (((11))) On the computer, A function of acquiring a speech image including the speaker; a function of acquiring a supplementary image for supplementing the speech content of the speaker, the supplementary image being based on the speech content but not a text directly representing the speech content; a function of controlling the display of the supplementary image at a two-dimensional position within the speech image corresponding to a three-dimensional position where a supplementary target exists within the speech image; A program to achieve this.
[0072] According to the invention of (((1))), it is more likely that a user will be able to intuitively recognize which supplementary subject a supplementary image intended to supplement the content of a speaker's speech is supplementing. According to the invention of (((2))), it becomes more likely that a user will be able to intuitively recognize which speaker's speech content a sign language image is supplementing. According to the invention of (((3))), sign language images can be displayed even without a function for generating moving images of sign language actions made by a virtual person from the sign language actions made by a sign language interpreter. According to the invention of (((4))), sign language images can be displayed without the sign language interpreter actually making sign language movements based on the content of the speech. According to the invention of (((5))), by using the video of the sign language actions performed by the sign language interpreter as is, it is possible to display the sign language image without placing a psychological burden on the sign language interpreter. According to the invention of (((6))), it is possible to make it appear as if the speaker is making sign language movements even if the speaker is not actually making sign language movements. According to the invention of (((7))), it is more likely that a speaker will be able to intuitively recognize which object contained in the speaker's speech content the explanatory image complements. According to the invention of (((8))), the possibility of intuitively recognizing which object the explanatory image complements is further increased. According to the invention of (((9))), the possibility of intuitively recognizing what state of an object an explanatory image represents increases. According to the invention of (((10))), the relationship between the explanatory image and the object can be easily seen. According to the invention of (((11))), it is more likely that a user will be able to intuitively recognize which supplementary subject a supplementary image intended to supplement the content of a speaker's speech is supplementing. [Explanation of symbols]
[0073] 1...AR system, 10...AR glasses, 30...AR server, 41...captured image acquisition unit, 42...audio information acquisition unit, 43...display information acquisition unit, 44...three-dimensional position identification unit, 45...display control unit
Claims
1. one or more processors; the one or more processors: Acquire a speech image including the speaker; acquiring a supplemental image for supplementing the speech content of the speaker, the supplemental image being based on the speech content but not being a text directly representing the speech content; Control is performed so that the supplementary image is displayed at a two-dimensional position within the utterance image corresponding to a three-dimensional position where a supplementary target exists within the utterance image. Information processing system.
2. the target to be supplemented is the speaker, The information processing system according to claim 1 , wherein the supplemental image is a sign language image representing a sign language based on the speech content.
3. The information processing system according to claim 2 , wherein the sign language image is a video image of a sign language action taken by a sign language interpreter based on the spoken content.
4. The information processing system according to claim 2 , wherein the sign language image is a moving image of sign language actions generated from the content of the utterance without requiring a sign language interpreter to perform sign language actions based on the content of the utterance.
5. The information processing system according to claim 2 , wherein the sign language images are moving images of sign language actions made by a virtual person generated from sign language actions made by a sign language interpreter based on the speech content.
6. The information processing system according to claim 2 , wherein the two-dimensional position within the speech image is a position where the speaker's hand is displayed within the speech image.
7. the supplementation target is an object included in the utterance content, The information processing system according to claim 1 , wherein the supplemental image is an explanatory image that shows an explanation of the object included in the speech content.
8. The information processing system according to claim 7 , wherein the explanatory image is an image indicating the object that is generated based on a demonstrative term that indicates the object and is included in the utterance content.
9. The information processing system according to claim 7 , wherein the explanatory image is an image representing a state of the object, which is generated based on a state word representing a state of the object included in the utterance content.
10. The information processing system according to claim 7 , wherein the two-dimensional position within the speech image is a position within a predetermined range around the target to be captured within the speech image.
11. On the computer, A function of acquiring a speech image including the speaker; a function of acquiring a supplementary image for supplementing the speech content of the speaker, the supplementary image being based on the speech content but not a text directly representing the speech content; a function of controlling the display of the supplementary image at a two-dimensional position within the speech image corresponding to a three-dimensional position where a supplementary target exists within the speech image; A program to achieve this.
Citation Information
Patent Citations
Video apparatus
JP1995191599A
Sensory Eyewear
JP2019535059A
Augmented reality system and method
JP2020077187A