Information processing apparatus, information processing method, and program

JP2025008721A5Pending Publication Date: 2026-05-29CANON KK

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2023-07-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies fail to accurately present environmental sounds to hearing-impaired individuals, such as ambulance sirens or car horns, as they rely solely on sound recognition, which can lead to misidentification.

Method used

An information processing device that captures images and sounds to identify sound sources, determines their priority, and displays corresponding character strings or speech bubbles on a display device, such as smartphones or AR glasses, based on a database that sets display levels.

Benefits of technology

Effectively visualizes important environmental sounds by identifying and prioritizing them through image and sound analysis, ensuring accurate notification to hearing-impaired users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To display a character string corresponding to a specific sound.SOLUTION: An information processing apparatus, on the basis of a picked-up image and a sound, specifies an object emitting the sound, and displays a character string corresponding to the sound emitted by the object.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] It is being considered to support the hearing impaired and hard of hearing by visually expressing sounds. For example, there is a technology that displays a character string obtained by recognizing a voice of a speaker in an image obtained by an imaging device capturing the speaker of an online conference, in accordance with the voice of the speaker. On the other hand, when going out, it is important to provide assistance such as notifying the hearing impaired and hard of hearing of high-priority environmental sounds such as ambulance sirens and car horns, in addition to human conversations.

[0003] For example, it is conceivable to display the environmental sounds using subtitles on a display device such as a smartphone or AR glasses. In this regard, Patent Document 1 discloses a technology for performing speech recognition processing on audio data of the environmental sounds, converting the environmental sounds into text, and visually presenting the text. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2011-250100 A Summary of the Invention [Problem to be solved by the invention]

[0005] However, when environmental sounds to be recognized are recognized only from the sound, there are cases where sound information cannot be properly presented to the user. For example, when walking down the street and accidentally picking up an ambulance siren broadcast on a television, the system cannot identify that it is an ambulance siren broadcast on the television and cannot inform the user of the fact.

[0006] The present disclosure provides a technique for displaying a character string corresponding to a specific sound. [Means for solving the problem]

[0007] An information processing device according to one aspect of the present disclosure that solves the above problem is characterized by having an acquisition means for acquiring an captured image and a sound, an identification means for identifying an object emitting the sound based on the captured image and the sound, and a display control means for displaying a character string corresponding to the sound emitted by the identified object. Effect of the Invention

[0008] According to the present disclosure, it is possible to display a character string corresponding to a specific sound. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram for explaining an example of use of an information processing device. [Diagram 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device. [Diagram 3] FIG. 2 is a block diagram showing an example of a software configuration of the information processing device. [Figure 4] 4 is a flowchart showing a flow of processing executed by the information processing device. [Diagram 5] FIG. 13 is a diagram showing an example of a voice type DB. [Figure 6] FIG. 13 is a diagram showing an example of visually notifying a sound. [Figure 7] FIG. 13 is a diagram showing an example of visually notifying a sound. [Figure 8] FIG. 2 is a block diagram showing an example of a software configuration of the information processing device. [Figure 9] 13 is a flowchart showing a process for changing a display level. [Figure 10] FIG. 13 is a diagram showing an example of a voice type DB after a change. [Figure 11] FIG. 2 is a block diagram showing an example of a software configuration of the information processing device. [Figure 12] FIG. 11 is a diagram showing an example of a change amount of a display level. [Figure 13] FIG. 13 is a diagram showing an example of a voice type DB. [Figure 14] FIG. 13 is a diagram showing an example of a voice type DB. [Figure 15] 4 is a flowchart showing a flow of processing executed by the information processing device. [Figure 16] FIG. 11 is a diagram showing an example of a change amount of a display level. [Figure 17] FIG. 13 is a diagram showing an example of a UI screen. [Figure 18] FIG. 13 is a diagram for explaining the specification of a voice generation target. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, the embodiments of the technology of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the technology of the present disclosure related to the claims, and not all of the combinations of features described in the present embodiments are essential to the solution of the technology of the present disclosure. Note that the same components are given the same reference numbers and descriptions are omitted.

[0011] <<Embodiment 1>> In this embodiment, an example of use of the information processing device when a user is walking in a city will be described. Fig. 1 is a diagram for explaining an example of use of the information processing device according to this embodiment. It is assumed that the terminal 110 is set to display on the display device 113 a captured image (hereinafter simply referred to as an image) captured by the front imaging device 111.

[0012] The user 101 has a terminal 110, which is an information processing device, and captures the surroundings with an imaging device provided in the terminal 110. The terminal 110 may be, for example, a mobile terminal such as a smartphone, a tablet, or AR glasses. The term "image" includes the concepts of both moving images and still images. That is, the terminal 110 can process both still images and moving images. The terminal 110 includes a front imaging device 111 that captures the front side of the terminal 110, and a rear imaging device 112 that captures the rear side of the terminal 110. The terminal 110 displays one of the two imaging devices 111 and 112 on a display device 113. In FIG. 1, since the front imaging device 111 and the user 101 face the same direction, it can be said that the terminal 110 obtains images of both the front and rear of the user 101. In addition, the terminal 110 obtains the sound of the surroundings of the user 101 based on a voice input device (not shown) built into the terminal 110. The terminal 110 identifies the object emitting the sound (hereinafter referred to as the sound-emitting object) from the captured image and the acquired sound. In Fig. 1, a dog 121 that is barking and emitting sound 131 and an ambulance 122 that is sounding its siren and emitting sound 132 are identified as the sound-emitting objects. The display level of the previously identified sound-emitting object is acquired based on a database (hereinafter referred to as the sound type DB) that holds a display level indicating the display priority level for each sound-emitting object stored in the terminal 110.

[0013] Terminal 110 compares a display level threshold value, which is held in terminal 110 and is used to determine whether to display to the user, with the display level of the identified voice emitting target, and determines whether to display. The display level threshold value used for the determination may be set by the user, or may be set automatically by terminal 110. If the determination result indicates that display should be performed, display device 113 built into terminal 110 displays a character string based on the voice uttered by the identified voice emitting target. The character string may be displayed at the edge of the display screen displayed by display device 113, or may be displayed using a speech bubble or the like.

[0014] In FIG. 1, the display device 113 of the terminal 110 displays the voice 131 of the dog 121 and the voice 132 of the ambulance 122 using a speech bubble 151 including the character "wanwan" (dog) and a speech bubble 152 including the character "peepo" (ambulance). That is, the terminal 110 performs display control so that the speech bubble 151 is displayed near the dog 141, which is near the object, within the display screen of the display device 113. The terminal 110 performs display control so that the horn of the speech bubble 151 is displayed toward the dog 141 on the display screen. The terminal 110 performs display control so that the horn of the speech bubble 152 is displayed toward the outside of the screen of the display device 113. In the present embodiment, the terminal 110 has been described as having the front imaging device 111, the rear imaging device 112, the voice input device, and the display device 113 built therein, but is not limited thereto. For example, the front imaging device 111, the rear imaging device 112, the audio input device, and the display device 113 may each be an independent device and connected to the system bus of the terminal 110 via a communication I / F.

[0015] (Hardware configuration of information processing device) 2 is a diagram showing an example of a hardware configuration of an information processing device according to the present embodiment. The information processing device 200 includes a CPU 201, a ROM 202, a RAM 203, an I / F 204, a storage device 205, and an I / F 206, and each unit is connected to a system bus 210. The CPU 201 is a central processing unit that controls the operation of each part in the information processing device 200 according to the contents of a program stored in the ROM 202, and executes a program loaded in the RAM 203. The ROM 202 is a read-only memory, and stores a boot program, firmware, various processing programs for implementing the processing described later, and various data. The RAM 203 is a work memory that temporarily stores programs and data for processing by the CPU 201, and various processing programs and data are loaded by the CPU 201. The I / F 204 is an interface for communicating with an external device such as a network device or a USB device, and performs data communication via a network and transmission and reception of data with the external device. The storage device 205 is a secondary storage area for saving various data, and may be configured to be secured inside the information processing device 200, or may be configured to be secured on the side of an external storage server connected via a network. The I / F 206 is an interface for transmitting and receiving data to and from internal devices, and is connected to an imaging device 207, an audio input device 208, and a display device 209. The imaging device 207 captures images to obtain still images and videos. The audio input device 208 acquires ambient sounds. The display device 209 displays images captured by the imaging device, various settings, and the like.

[0016] (Software configuration of information processing device) 3 is a block diagram showing an example of the software configuration of the information processing device 200. Details of each process and data will be described later with reference to a flowchart.

[0017] The imaging device 207 in the information processing device 200 has an image acquisition unit 301. The audio input device 208 in the information processing device 200 has an audio acquisition unit 302. The display device 209 in the information processing device 200 has a target object identification unit 303, an audio type identification unit 304, an audio type DB 305, a display determination unit 306, a display level setting unit 307, an operation unit 308, a subtitle generation unit 309, and a display unit 310.

[0018] The image acquisition unit 301 acquires an image captured by the imaging device 207. When the imaging device 207 is configured with an imaging device that captures both the front side and the rear side of the information processing device, the image acquisition unit 301 acquires an image of the front side and an image of the rear side. The audio acquisition unit 302 acquires audio picked up by an audio input device. The target object identification unit 303 detects and identifies the target object (also called the target object) that is generating the audio from the image acquired by the image acquisition unit 301 and the audio acquired by the audio acquisition unit 302.

[0019] The sound type identification unit 304 acquires a display level of the sound generation target for the user based on the sound generation target identified by the target object identification unit 303 and the sound type DB 305. The display determination unit 306 determines whether to display the sound to the user based on the display level acquired by the sound type identification unit 304 and a display level threshold set by the display level setting unit 307 via the operation unit 308.

[0020] When the display determination unit 306 determines that the subtitle should be displayed, the subtitle generation unit 309 generates a subtitle corresponding to the target of the voice generation. The generated subtitle may be a character string obtained by recognizing the acquired voice and converting it to text, may be the type of the identified voice generation target, may be a character string previously associated with the identified voice generation target, or may be a combination of these. The display unit 310 displays the subtitle generated by the subtitle generation unit 309. The display of the subtitle may display only the subtitle (character string), or may display the subtitle in the speech bubble together with the speech bubble. Each functional unit may perform part of the functions of other functional units.

[0021] (Flow of processing executed by information processing device) 4 is a flowchart showing the flow of processing executed by information processing device 200. In addition, the symbol "S" in the explanation of the flowchart represents a step. This also applies to the explanation of the following flowcharts. The flowchart of FIG. 4 may start when information processing device 200 is powered on, or when a dedicated application installed in information processing device 200 is started.

[0022] In S401, the image acquisition unit 301 and the audio acquisition unit 302 respectively acquire an image captured by the imaging device and an audio picked up by the audio input device. Note that the image and the audio each contain time information.

[0023] In S402, the target object identification unit 303 identifies the target generating the sound using the image and sound acquired in S401. A method for identifying the target generating the sound will be described. First, the sound input device acquires the distance from the user to the target generating the sound and the direction from the user to the target generating the sound based on the acquired sound. The distance and direction are acquired from the phase difference of the input sound using, for example, two or more microphones. The distance and direction may also be acquired from the input sound using one microphone having directivity for collecting sound. From the acquired distance and direction to the target generating the sound and the positional relationship between the imaging device and the sound input device, it is identified where the sound is generated in the image acquired by the imaging device. Note that the recognition of the positional relationship between the imaging device and the sound input device is input in advance to the information processing device in the case of a device in which the imaging device and the sound input device are integrated into one. If the imaging device and the audio input device are separate devices, for example, by capturing an image of the imaging device including the audio input device, the distance and position of the audio input device can be estimated from the size of the audio input device and the pixels of the image, and therefore the relative positional relationship between the imaging device and the audio input device can be estimated.

[0024] In S403, the voice type identification unit 304 identifies the type of voice by recognizing an object at the position of the voice generation target identified in S402 based on the position of the voice generation target in the image identified in S402. The recognition of the voice generation target object is performed using object recognition on the image using machine learning such as YOLO or CNN, for example, based on the position of the voice generation target in the image identified in S402. That is, the object may be identified by performing object recognition using a trained model on a partial image identified based on the voice in the image obtained by the information processing device. The type of voice may be identified, for example, by having a database that holds a group of character strings corresponding to the voice emitted by the object detected by object recognition. Then, it is determined whether the character string recognized by voice recognition matches the character string in the database, and if they match, it is identified as the type of voice. That is, the type of voice may be identified by performing voice recognition using a trained model on the voice. For example, if a database is stored that stores character strings that dogs utter, such as "woof," "bow," and "can," and a "dog" is detected by object recognition, when "woof" is recognized by voice recognition, the type of sound is identified as "dog barking." On the other hand, for sounds emitted from objects for which it is difficult to retain a group of character strings, the type of sound may be identified as "sound of (object name)" without matching character strings. For example, because objects emit a variety of sounds, such as radio sounds and television sounds, it is difficult to retain the character strings of the sounds they emit. Therefore, when it is recognized by object recognition that the source of the sound is a radio or television, the type of sound is identified as "radio sound" or "television sound."

[0025] In S404, the display determination unit 306 acquires a display level for the identified type of voice based on the type of voice identified in S403 and the voice type DB 305.

[0026] (About the voice type DB) FIG. 5 is a diagram showing an example of a voice type DB. The voice type DB 500 is an example of the voice type DB 305, and holds data showing the correspondence between a voice type 501 and a display level 502. For example, an ambulance siren, an emergency earthquake report, music in a store, a passing car sound, a dog barking, a radio sound, a television sound, and a rain sound are registered as the voice type 501. The display level 502 indicates a level of display for notifying a user. In FIG. 5, the display level is expressed in three stages, 0, 1, and 2, and the higher the number, the higher the priority of the display. In FIG. 5, the priority is set to be higher according to the level of importance or urgency for the user, but the priority standard may be changed according to the content to be notified by the information processing device. Alternatively, the user may manually set the display level for each voice type. In addition, in FIG. 5, the display level is expressed in three stages, but the present invention is not limited to this, and the resolution of the display level may be greater than or less than three stages.

[0027] In S405, the display determination unit 306 determines whether the display level acquired in S404 is equal to or greater than the threshold display level set by the display level setting unit 307. The threshold display level may be set by the user via the operation unit 308, or may be set in advance by the information processing device 200. If the determination result indicates that the display level acquired in S404 is equal to or greater than the threshold display level (YES in S405), the process proceeds to S406. If the determination result indicates that the display level acquired in S404 is less than the threshold display level (NO in S405), the flow shown in FIG. 4 is terminated.

[0028] In S406, the display unit 310 displays the subtitles of the target audio generated by the subtitle generation unit 309. When the process of S406 ends, the flow shown in Fig. 4 ends. However, in the flow shown in Fig. 4, the processes from S401 to S406 may be continuously executed until the power of the information processing device 104 is turned off or the dedicated application is ended.

[0029] (Audio visual notification) Fig. 6 is a diagram showing an example of visually notifying the sound in a scene where a dog is barking and an ambulance is sounding its siren. Fig. 6(a) shows a case where a dog's bark and an ambulance siren are notified, and Fig. 6(b) shows a case where an ambulance siren is notified. Figs. 6(a) and 6(b) show a case where a dog 141 and an ambulance are shown as sound generating targets in the images captured by the front imaging device and the rear imaging device 112, and a notification to the display device 113 is shown in a state where the display device 113 displays the images captured by the front imaging device.

[0030] In Fig. 6(a), the display level of the dog's bark and the ambulance siren are set to be equal to or higher than the display level threshold. With the above settings, both the dog's bark and the ambulance siren are determined to be audio display targets, and so the audio is displayed as subtitles, such as speech bubble 151 containing "woof woof (dog)" and speech bubble 152 containing "pee-poo (ambulance)."

[0031] 6(b), the display level of the dog's bark is set below the threshold, but the display level of the ambulance siren is set above the threshold. With the above settings, it is determined that only the ambulance siren is the audio display target, so the subtitle corresponding to the dog 141 is not displayed, and only the audio to be displayed is displayed as a subtitle, such as the speech bubble 152 including "Peepo (ambulance)". In other words, only the speech bubble 152 corresponding to the ambulance siren, whose display level is above the threshold, may be displayed.

[0032] When dog 141, which is the object of sound generation, is displayed on the display screen of display device 113 as in speech bubble 151, for example, the horns of speech bubble 151 may be displayed facing the object of sound generation. Also, when the object of sound generation is not displayed as in speech bubble 152, the direction in which the object of sound generation exists may be captured in a two-dimensional direction, and the horns of the speech bubble may be displayed facing in that direction. In other words, it may be displayed facing outward from the display screen.

[0033] The size of the speech bubble may be changed to express the depth direction. That is, a larger speech bubble may be displayed on the front side than on the back side, and a smaller speech bubble may be displayed on the back side than on the front side. In addition, when there are multiple voices displayed, for example, the speech bubbles may be displayed so that they do not overlap each other. In order to make it easier for the user to recognize the subtitles corresponding to the voices with high priority, the subtitles may be displayed in order from the front side to the back side of the display screen, from the high to the low display level. In addition, the speech bubble may be displayed continuously from the time the subtitle corresponding to the voice is displayed on the display device 113 until the voice is no longer generated. Alternatively, the speech bubble may be displayed for a certain period of time and then erased. In addition, the speech bubble may be displayed for a certain period of time regardless of the type of voice, or the display time may be changed depending on the display level.

[0034] FIG. 7 is a diagram showing an example of a notification assuming a device such as AR glasses. FIG. 7(a) shows an example of the use of AR glasses, FIG. 7(b) shows an example of the display of AR glasses, and FIG. 7(c) shows the relationship between a person's field of vision and the imaging range of the AR glasses. The AR glasses may use a known AR (Augmented Reality) technology to, for example, AR-display subtitles corresponding to a voice generation target in the vicinity of the actual position of the voice generation target. In FIG. 7, it is assumed that an imaging device that images the front and rear sides of a device such as AR glasses worn by a user 701 and an audio input device are provided. In FIG. 7, as in the case of FIG. 6(a), it is assumed that a dog 721 and an ambulance 722 are shown in the image of the imaging device, and the relationship between the display level of each sound and the threshold value of the display level is also the same. An image 711 is displayed on the AR glasses 710 so as to be visible to the user 701. As in FIG. 6(a), the horns may be displayed facing the target of the sound generation, such as speech bubble 751 including "woof woof (dog)" and speech bubble 752 including "pee-poo (ambulance)". Speech bubble 751 may be displayed near the target, dog 741. Also, when the target of the sound generation is not displayed, speech bubble 752 may include a character string indicating the direction of ambulance 722 from the user as a starting point, such as "direction: right rear". In FIG. 7(b), "direction: right rear" is displayed, but the present invention is not limited to this. In other words, it can be said that the lens of AR glasses 710 worn by the user displays a character string obtained from the sound emitted by the target.

[0035] On the other hand, when the dog 741, which is the source of the sound, is visible, the position of the source of the sound can be seen on the image, so the direction of the source of the sound does not need to be displayed. Also, the size of the text, the size of the speech bubble, the position at which the text is displayed, the position at which the speech bubble is displayed, the style of the text, the color of the text, the color of the speech bubble, etc. may be changed according to the display level of the sound. For example, when the display level is higher than the threshold, larger text or larger speech bubbles may be displayed compared to when the display level is lower than the threshold.

[0036] The character string or speech bubble may be displayed from the front side to the back side. In other words, the character string or speech bubble may be displayed so that the display level goes from a relatively high level to a relatively low level from the front side to the back side.

[0037] The style of the text may be changed from normal to bold, i.e., when the display level is lower than the threshold, the text may be displayed normal, and when the display level is higher than the threshold, the text may be displayed bold.

[0038] The color of the text or speech bubble may be changed from black to red or yellow. That is, when the display level is lower than the threshold, black text or a black speech bubble may be displayed, and when the display level is higher than the threshold, red or yellow text or a red or yellow speech bubble may be displayed.

[0039] <<Embodiment 2>> In this embodiment, a configuration for updating the voice type DB 305 is added to the information processing apparatus 200 of the first embodiment, and will be described with reference to the drawings. In this embodiment, the differences from the first embodiment will be mainly described.

[0040] (Software configuration of information processing device) 8 is a block diagram showing an example of the software configuration of an information processing device 800 according to this embodiment. The display device 209 in the information processing device 800 of this embodiment has the same functional units as the display device 209 in the information processing device 200, and also has a display level changing unit 801. The display level changing unit 801 sets a display level for each type of audio. The setting method may be a user setting via the operation unit 308 of the information processing device 800, or a setting may be made from outside the information processing device 800 via a communication I / F.

[0041] (Process flow for changing display level) Fig. 9 is a flowchart showing a process flow for changing the display level. The process of the flowchart in Fig. 9 may start when the user performs setting via the operation unit 308, or when the user performs setting via the communication I / F from outside the information processing device 800.

[0042] In S901, the display level change unit 801 acquires the type of audio whose display level is to be changed.

[0043] In S902, the display level change unit 801 acquires the display level after the change of the audio type acquired in S901. S901 and S902 may be acquired from the operation unit 308 of the information processing device 800, or may be acquired from outside the information processing device 800 via a communication I / F.

[0044] In S903, the display level change unit 801 updates the voice type DB 305 based on the voice type acquired in S901 and the changed display level acquired in S902. When the process of S903 ends, the flow shown in FIG. 9 ends.

[0045] (Example of updating the voice type DB) Fig. 10 is a diagram showing an example of the sound type DB after the display level change unit 801 has made a change. Fig. 10 shows a case where the display level of "car passing sound" has been changed among the sound types registered in the sound type DB 500 shown in Fig. 5. The sound type DB 1000 after the display level change shows the relationship between the sound type 1001 and the changed display level (3 stages) 1002. In Fig. 5, the display level of "car passing sound" was "1", but in Fig. 10, the display level of "car passing sound" has been changed to "2".

[0046] As described above, this embodiment is effective when there is a sound that is to be presented to the user of the information processing device 800 from the outside. For example, when the user of the information processing device 800 is walking on a road where there are many traffic accidents such as collisions between cars and pedestrians, the road maintenance side sets the display level of the sound type "sound of passing cars" to be high. This allows the user to be visually notified that the user has received the sound of passing cars more reliably than when the user is in another location, thereby reducing the possibility of the user colliding with a car.

[0047] <<Embodiment 3>> In this embodiment, a configuration for checking a user's behavior and state is added to the information processing device 800 of the second embodiment, and will be described with reference to the drawings. In this embodiment, the differences from the second embodiment will be mainly described.

[0048] (Software configuration of information processing device) FIG. 11 is a block diagram showing an example of the software configuration of the information processing device 1100 of this embodiment. The display device 209 in the information processing device 1100 of this embodiment has the same functional unit as the display device 209 in the information processing device 800, and also has a user behavior / state confirmation unit 1101. The user behavior / state confirmation unit 1101 confirms what behavior the user of the information processing device 1100 is taking, or the user's current location and the user's state including the user's age. For example, since it becomes difficult to hear sounds of a certain frequency due to aging, when it is recognized that the user's state is 60 years old or older, a change may be made such that the display level is increased by 1. That is, the user behavior / state confirmation unit 1101 confirms one or more items of the user's current location, the user's age, whether the user is standing, whether the user is sitting, whether the user is moving, and whether the user is stationary. Then, the display level corresponding to the type of voice may be changed according to the confirmation result confirmed by the user behavior / state confirmation unit 1101.

[0049] The user's behavior may be estimated by performing behavior recognition on an image acquired by the imaging device 207 of the information processing device 1100, or the user's movement may be recognized using a speed / acceleration sensor such as a gyro sensor in the information processing device 1100. Alternatively, behavior recognition may be performed on an image acquired by an external device such as a surveillance camera, and the obtained behavior recognition result may be received via the communication I / F.

[0050] The user's behavior is, for example, standing / sitting, moving / still, etc. The user's state may be estimated from an image captured by the imaging device 207 of the information processing device 1100, similar to the user's behavior, or may be input in advance by the user. Alternatively, the state may be recognized by an external device such as a surveillance camera and received via the communication I / F.

[0051] FIG. 12 is an example of a database used to update the voice type DB 500 shown in FIG. 5 when it is recognized that the user's state has changed from "outdoors" to "indoors". As shown in FIG. 12, a database 1200 is provided that defines how much the display level of each voice type should be changed in response to various changes in the user's behavior and state. The database 1200 shows the relationship between a voice type 1201 and a change amount 1202 in the display level. The voice type 1201 indicates the type of voice to be changed. The change amount 1202 in the display level indicates the change amount in the display level when the user's state changes.

[0052] 12, when the user is indoors, the ambulance siren has a lower priority in terms of importance and urgency than when the user is outside, so the display level change amount 1202 is set to -1. Similarly, the sound of a passing car has a lower priority in terms of importance and urgency when the user is indoors than when the user is outside, so the display level change amount 1202 is set to -1. On the other hand, for emergency earthquake alerts, in-store music, dog barking, radio sounds, television sounds, and rain sounds, it is determined that the priority does not change in terms of importance and urgency, and the display level change amount 1202 is set to 0.

[0053] When the sound type DB 500 shown in Fig. 5 is updated based on the database 1200 shown in Fig. 12 and the user state confirmed by the user action / state confirmation unit 1101, the updated database 1300 is obtained as shown in Fig. 13. Comparing Fig. 5 with Fig. 13, the display levels of ambulance sirens and car passing sounds are decreased by one, and the others remain unchanged. Also, the display levels when moving outside from the state of the music type DB 1300 shown in Fig. 13 are returned to the state of the sound type DB 500 shown in Fig. 5 by using a definition database of the amount of change in display level corresponding to this state change.

[0054] As described above, according to this embodiment, it is possible to present information necessary for the user in real time. For example, when the user of the information processing device 1100 is indoors, the user is not particularly interested in the sound of a passing car or the siren of an ambulance outside, so it is not necessary to present such information that is of no interest to the user.

[0055] In the above, voice input is used to recognize the position of the source of the voice or the character string of the voice, but the type of voice may be classified more specifically. For example, as shown in FIG. 14, the type of voice 1401 may be classified not only as music, but also more specifically as Beethoven's music, Chopin's music, etc. By classifying in this way, it is possible to visually present to the user what kind of music is currently being played, for example, as an aid for the hearing-impaired or hard-of-hearing when taking music lessons. In addition, character strings (character information) for the composer's name, song title, performer's name, musical score, etc. may be registered in a database, and the corresponding character information may be obtained from the database and displayed according to the identified type of voice and display level.

[0056] Although the above description is given using music as an example, the present invention is not limited to this. For example, the present invention can be applied to theater and the like. In this case, character strings (text information) for the author, work name, director, performer name, etc. may be registered in a database, and the corresponding text information may be obtained from the database and displayed according to the specified audio type and display level.

[0057] The information processing device 200 of the first embodiment may take into account a change in the display level corresponding to information on the traveling direction of a specific voice generating target. For example, if the voice generating target is an ambulance, the change in the display level corresponding to the classification "approaching the user", "moving away from the user", or "no change" obtained from information indicating the traveling direction of the voice generating target may be taken into account in the display level of the voice generating target.

[0058] (Flow of processing executed by information processing device) Fig. 15 is a flowchart showing the flow of processing executed by the information processing device 200. The explanation will focus on the differences from the flowchart shown in Fig. 4. In S403, it is assumed that a voice-emitting target from which information on the traveling direction is to be acquired is specified. For example, it is assumed that an emergency vehicle such as a police car, an ambulance, or a fire engine is specified as the voice-emitting target. Fig. 16 is a diagram showing an example of the relationship between the type of voice, the traveling direction, and the amount of change in the display level.

[0059] In S1501, the target object identification unit 303 acquires information indicating the traveling direction of the sound generating object based on the sound acquired in S401, or based on both the sound and the image acquired in S401. The target object identification unit 303 identifies which of the three categories, "approaching the user", "moving away from the user", and "no change", associated with the sound type, the sound corresponds to based on the acquired sound type and the acquired information on the traveling direction. Then, it determines the amount of change in the display level corresponding to the identified category. For example, when the siren of an ambulance is identified as the sound type, it determines the amount of change in the display level 1603 corresponding to the traveling direction 1602 based on the database 1600 shown in FIG. 16.

[0060] For example, distance information is acquired from the voice, and the corresponding category is identified from among the categories "approaching the user", "moving away from the user", and "no change" based on the acquired distance information. The sound source generation position is identified from the voice at time A and the voice at time B after a predetermined time t seconds have elapsed from the time A, using a sound source separation technology, respectively. Then, if the distance between the identified two positions is smaller than a predetermined threshold C, the category "approaching the user" is identified as the category to which the information on the traveling direction corresponds. The sound source generation position is identified from the voice at time A and the voice at time B, using a sound source separation technology, respectively, and if the distance between the identified two positions is greater than a predetermined threshold D, the category "moving away from the user" is identified as the category to which the information on the traveling direction corresponds. The sound source generation position is identified from the voice at time A and the voice at time B, using a sound source separation technology, respectively, and if the distance between the identified two positions is between thresholds C and D, the category "no change" is identified as the category to which the information on the traveling direction corresponds. Then, the amount of change in the display level corresponding to the identified category is determined.

[0061] In S1502, the display determination unit 306 changes the display level based on the amount of change in the display level determined in accordance with the information on the traveling direction acquired in S1501 and the display level acquired in S404, and acquires the changed display level.

[0062] In S1503, the display determination unit 306 determines whether the display level acquired in S1502 is equal to or greater than the threshold display level set by the display level setting unit 307. The threshold display level may be set by the user via the operation unit 308, as in S405 above, or may be set in advance by the information processing device 200. If the determination result indicates that the display level acquired in S1502 is equal to or greater than the threshold display level (YES in S1503), the process proceeds to S406. If the determination result indicates that the display level acquired in S1502 is less than the threshold display level (NO in S1503), the flow shown in FIG. 15 ends.

[0063] In this way, by changing the display level according to the information on the moving direction of the voice generation target, it is possible to visually present to the user the sounds around the user, such as an approaching emergency vehicle such as an ambulance.

[0064] In the above, a description has been given of an embodiment in which it is assumed that the type of voice to be recognized is set in advance in the information processing device, but the present invention is not limited to this, and the user may select and set the type of voice to be recognized. FIG. 17 is a diagram showing an example of a UI screen for a user to set the type of voice to be recognized. For example, as shown in FIG. 17, a UI screen 1700 may be displayed on the display unit 310 of the information processing device to receive a user operation for setting the type of voice to be recognized. The UI screen 1700 displays types of recognizable sounds 1701, such as an ambulance siren, a passing car sound, an emergency earthquake alert, music in a store, and the sound of rain, and a selection button 1702 for the user to select the type of voice that the user wants to recognize. By performing a user operation to select, such as clicking the selection button 1702 corresponding to the type of voice that the user wants to recognize, the corresponding type of voice can be set as the type of voice to be recognized. Note that FIG. 17 shows a case in which an ambulance siren and music in a store are selected and set as the type of voice to be recognized.

[0065] In the above, the embodiment of identifying the voice generation target using the image acquired by the imaging device and the voice acquired by the voice input device has been described, but is not limited thereto. For example, assume that data corresponding to an actual three-dimensional space is stored in an external database, and an ID indicating an ambulance is linked to a specific box-shaped space in the data. FIG. 18 is a diagram for explaining the identification of the voice generation target. When an image captured by a user 1801 using a terminal 1810 such as a smartphone includes an image corresponding to a specific box-shaped space 1822 in the database, the specific box-shaped space 1822 is recognized and an ID linked to this space 1822 is obtained. The voice generation target is recognized by this ID. For example, when the ID indicates an ambulance 1821, the ambulance is identified as the voice generation target. That is, a database is stored in advance, which corresponds to a three-dimensional space in which an ambulance, which is an object that generates a voice, exists, and in which an ID is linked to a space corresponding to the ambulance. The database may be stored in the information processing device or in an external device. When an image corresponding to the space linked to the ID is recognized in the images obtained by the information processing device, the target object is identified by the ID.

[0066] In this way, it is possible to recognize that an ambulance is an object without performing object recognition on the image captured by the terminal. Note that the accuracy of recognizing an object as an ambulance can be improved by combining the spatial ID of the object recognition from the image.

[0067] Although the above description has been given of a case where an image captured by one of the front image capturing device and the rear image capturing device is displayed, the present invention is not limited to this. Both images captured by the front image capturing device and the rear image capturing device may be displayed.

[0068] In the above, examples of environmental sounds have been described as earthquake early warnings, ambulance sirens, passing car sounds, music in a store, barking dogs, radio sounds, television sounds, and rain sounds, but the present invention is not limited to these. It can also be applied to sirens of emergency vehicles such as police car sirens and fire engine sirens, passing train sounds, and the like.

[0069] (Other embodiments) The present disclosure can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0070] The disclosure of this embodiment includes the following configuration examples. (Configuration 1) An acquisition means for acquiring a captured image and a sound; an identification unit that identifies an object that is emitting the sound based on the captured image and the sound; a display control means for displaying a character string corresponding to the sound emitted by the identified object; 13. An information processing device comprising:

[0071] (Configuration 2) a storage means for storing a display level indicating a priority level for displaying the identified object in association with the identified object; A setting means for setting a threshold value of the display level; and The display control means displays the character string when the display level corresponding to the specified object and stored in the storage means is equal to or higher than the threshold value set by the setting means. 2. The information processing device according to configuration 1.

[0072] (Configuration 3) The display control means Displaying the character string or a balloon containing the character string; Depending on the display level, any one of the size of the character string, the size of the speech bubble, the position at which the character string is displayed, the position at which the speech bubble is displayed, the style of the character string, the color of the character string, and the color of the speech bubble is changed and displayed. 3. The information processing device according to configuration 2.

[0073] (Configuration 4) The display control means When the display level corresponding to the identified object is higher than the threshold, the character string or the speech bubble is displayed larger than when the display level is lower than the threshold. 4. The information processing device according to configuration 3.

[0074] (Configuration 5) The display control means The character string or the speech bubble is displayed so that the display level changes from a relatively high level to a relatively low level from the front side to the back side. 5. The information processing device according to configuration 3 or 4.

[0075] (Configuration 6) The display control means When the display level corresponding to the identified object is lower than the threshold value, the character string is displayed as a standard. If the display level corresponding to the identified object is higher than the threshold, the character string is displayed in bold. 6. The information processing device according to any one of configurations 3 to 5.

[0076] (Configuration 7) The display control means When the display level corresponding to the identified object is lower than the threshold value, the character string or the speech bubble is displayed in black; When the display level corresponding to the identified object is higher than the threshold, the character string in red or yellow, or the speech bubble in red or yellow is displayed. 7. The information processing device according to any one of configurations 3 to 6.

[0077] (Configuration 8) The display level storing unit stores the display level of the display device. The display control means displays the character string when the display level changed by the change means and corresponding to the specified object is equal to or higher than the threshold value. 8. The information processing device according to any one of configurations 2 to 7.

[0078] (Configuration 9) The device further includes a confirmation means for confirming a state related to the user, The change means changes the display level in accordance with the state of the user confirmed by the confirmation means. 9. The information processing device according to configuration 8.

[0079] (Configuration 10) the confirmation means confirms one or more of the following: a current location of the user, an age of the user, whether the user is standing, whether the user is sitting, whether the user is moving, and whether the user is stationary; The change means changes the display level based on a result of the confirmation made by the confirmation means. 10. The information processing device according to configuration 9.

[0080] (Configuration 11) The identification means identifies a traveling direction of the object, the acquiring means acquires information indicating a traveling direction of the object identified by the identifying means, The change means changes the display level in accordance with the information on the traveling direction acquired by the acquisition means. 11. The information processing device according to configuration 10.

[0081] (Configuration 12) When the acquisition means acquires information on the traveling direction indicating that the object is moving away, the change means changes the display level to a lower level compared to before acquiring the information on the traveling direction. 12. The information processing device according to configuration 11.

[0082] (Configuration 13) When the acquisition means acquires information on the traveling direction indicating that the object is approaching, the change means changes the display level to a higher level than before the acquisition of the information on the traveling direction. 13. The information processing device according to configuration 11 or 12.

[0083] (Configuration 14) When the confirmation means confirms that the user is in a room, the acquisition means changes the display level corresponding to the specific object to a lower display level compared to the display level before obtaining the result of the confirmation. 14. The information processing device according to any one of configurations 10 to 13.

[0084] (Configuration 15) When the confirmation means confirms an object emitting a sound in a direction in which the user is moving, the acquisition means changes the display level corresponding to the confirmed object to a higher display level compared to the display level before the confirmation result was obtained. 15. The information processing device according to any one of configurations 10 to 14.

[0085] (Configuration 16) The apparatus further includes a setting means for setting the object to be recognized, The display control means displays a character string based on a sound emitted by the object set by the setting means among the specified objects. 16. The information processing device according to any one of configurations 1 to 15.

[0086] (Configuration 17) A generating unit is further provided for generating the character string to be displayed based on the sound, The display control means displays the character string generated by the generating means. 17. The information processing device according to any one of configurations 1 to 16.

[0087] (Configuration 18) The generating means generates a character string obtained by recognizing the sound, or generates a character string indicating a type of sound emitted by the identified object. 18. The information processing device according to configuration 17.

[0088] (Configuration 19) The display control means displays the character string on a display screen of a display means that displays the captured image, or displays the character string using a speech bubble on the display screen. 19. The information processing device according to any one of configurations 1 to 18.

[0089] (Configuration 20) a front side imaging means for imaging a front side of a user to obtain an image; a rear side imaging means for imaging a rear side of the user to obtain an image; having The acquisition means includes: a front side captured image acquired by the front side imaging means; a rear-side captured image captured by the rear-side imaging means; The display control means displays at least one of the captured image on the front side and the captured image on the rear side. 20. An information processing device according to any one of configurations 1 to 19.

[0090] (Configuration 21) The display control means displays the character string based on a sound emitted by the object included in the displayed captured image near the object. 21. The information processing device according to configuration 20.

[0091] (Configuration 22) 22. The information processing device according to configuration 21, wherein the display control means displays a tip of a speech bubble containing the character string therein toward the target object.

[0092] (Configuration 23) The display control means displays the character string based on the sound emitted by the object included in the captured image that is not being displayed. 21. The information processing device according to configuration 20.

[0093] (Configuration 24) 24. The information processing device according to configuration 23, wherein the display control means displays a tip of a speech bubble containing the character string therein toward the outside of the display.

[0094] (Configuration 25) The identification means identifies the type of the sound emitted by the identified object. 25. The information processing device according to any one of configurations 1 to 24.

[0095] (Configuration 26) The identification means performs voice recognition on the sound using a trained model to identify the type of the sound. 26. The information processing device according to configuration 25.

[0096] (Configuration 27) The identification means identifies the target object by performing object recognition using a trained model on a partial image identified based on the sound in the captured image. 27. An information processing device according to any one of configurations 1 to 26.

[0097] (Configuration 28) A database is stored in advance, the database corresponding to the three-dimensional space in which the object exists, and an ID is associated with the space corresponding to the object; When the identification unit recognizes an image corresponding to the space associated with the ID in the captured image, the identification unit identifies the object by the ID. 27. An information processing device according to any one of configurations 1 to 26.

[0098] (Configuration 29) 29. The information processing device according to any one of configurations 1 to 28, wherein the sounds include an emergency vehicle siren, an emergency earthquake alert, the sound of a passing car, music playing in a store, and a dog barking.

[0099] (Configuration 30) The display control means displays the character string on a display unit of an information processing device used by a user. 30. An information processing device according to any one of configurations 1 to 29.

[0100] (Configuration 31) 30. The information processing device according to any one of configurations 1 to 29, wherein the display control means displays the character string on a lens of AR glasses worn by a user.

[0101] (Configuration 32) an acquisition step of acquiring a captured image and a sound; an identifying step of identifying an object that emits the sound based on the captured image and the sound; a display control step of displaying a character string corresponding to the sound emitted by the identified object; 13. An information processing method comprising:

[0102] (Configuration 33) A program for causing a computer to execute the information processing method according to configuration 32.

Claims

1. An acquisition means for acquiring captured images and sound, A means for identifying the object emitting the sound based on the captured image and the sound, A display control means that displays a string of characters corresponding to the sound emitted by the identified object, An information processing device characterized by having the following:

2. A storage means for storing a display level indicating the priority of displaying the identified object, A setting means for setting the threshold of the display level, It further possesses, The display control means displays the string of characters if the display level stored in the storage means, corresponding to the identified object, is equal to or greater than the threshold set by the setting means. The information processing apparatus according to feature 1.

3. The display control means is Display the aforementioned string, or a speech bubble containing the aforementioned string. Depending on the size of the display level, the size of the text string, the size of the speech bubble, the position where the text string is displayed, the position where the speech bubble is displayed, the style of the text string, the color of the text string, or the color of the speech bubble will be changed for display. The information processing apparatus according to feature 2.

4. The display control means is If the display level corresponding to the identified object is higher than the threshold, a larger string or larger callout is displayed compared to when the display level is lower than the threshold. The information processing apparatus according to claim 3.

5. The display control means is The text string or the speech bubble is displayed such that the display level decreases from a relatively high level towards the back, and from a relatively low level towards the front. The information processing apparatus according to claim 3.

6. The display control means is If the display level corresponding to the identified object is lower than the threshold, the string is displayed by default. If the display level corresponding to the identified object is higher than the threshold, the string will be displayed in bold. The information processing apparatus according to claim 3.

7. The display control means is If the display level corresponding to the identified object is lower than the threshold, the black text string or the black speech bubble is displayed. If the display level corresponding to the identified object is higher than the threshold, the text in red or yellow, or the speech bubble in red or yellow, will be displayed. The information processing apparatus according to claim 3.

8. The storage means further comprises a means for changing the display level stored in the storage means, The display control means displays the string if the display level changed by the modification means, and the display level corresponding to the identified object, is equal to or greater than the threshold. The information processing apparatus according to feature 2.

9. The system further includes setting means for setting the object to be recognized, The display control means displays a string of characters based on the sound emitted by the object specified by the setting means among the identified objects. The information processing apparatus according to feature 1.

10. The system further includes a generation means for generating the string to be displayed based on the sound, The display control means displays the string generated by the generation means. The information processing apparatus according to feature 1.

11. The generation means generates a string obtained by recognizing the sound, or generates a string indicating the type of sound emitted by the identified object. The information processing apparatus according to feature 10.

12. The display control means displays the string within the display screen of the display means that displays the captured image, or displays it within the display screen using a callout. The information processing apparatus according to feature 1.

13. A front-facing imaging means that captures images of the area in front of the user and acquires the captured image, A rear-facing imaging means that captures an image of the area behind the user and acquires the captured image, It has, The acquisition means uses the front-facing image captured and acquired by the front-facing imaging means, The rear imaging means acquires the rear image captured by the rear imaging means, The display control means displays at least one of the front-facing captured image and the rear-facing captured image. The information processing apparatus according to feature 1.

14. The identification means identifies the type of sound emitted by the identified object. The information processing apparatus according to feature 1.

15. The aforementioned identification means identifies the target object by performing object recognition using a trained model on the partial image identified in the captured image based on the sound. The information processing apparatus according to feature 1.

16. A database is pre-maintained that corresponds to the three-dimensional space in which the object exists, and in which an ID is associated with the space corresponding to the object. When the identification means recognizes an image in the captured image that corresponds to the space associated with the ID, it identifies the object using the ID. The information processing apparatus according to feature 1.

17. The information processing apparatus according to claim 1, characterized in that the aforementioned sounds include emergency vehicle sirens, earthquake early warnings, passing car sounds, in-store music, and dog barking.

18. The display control means displays the string on the display unit of the information processing device used by the user. The information processing apparatus according to feature 1.

19. The information processing apparatus according to claim 1, characterized in that the display control means displays the string of characters on the lenses of AR glasses worn by the user.

20. The acquisition process involves acquiring both captured images and sound. A process of identifying the object emitting the sound based on the captured image and the sound, A display control step that displays a string of characters corresponding to the sound emitted by the identified object, An information processing method characterized by including

21. A program for causing a computer to perform each step in the information processing method described in claim 20.